Vestor
@kingdomvestor
journal
Did you know AI doesn’t actually “see” your images the way you do?
0 replies 2 views
Did you know AI doesn’t actually “see” your images the way you do?
I used to think that when I uploaded a photo to ChatGPT, Claude, Grok, or DeepSeek, the model was somehow looking at the same picture I was seeing.
A person. A laptop. A coffee mug.
Turns out… not exactly.
When you upload an image, the AI processes numerical information representing the image’s pixels. No “tired face.” No “laptop screen.” Just data.
From there, the model identifies patterns:
- Edges
- Shapes
- Textures
- Objects
- Positions and relationships
Step by step, these patterns become higher-level representations. Eventually, the model can arrive at something like: “there is a person sitting at a desk.”
Modern multimodal models then connect these visual representations to language, allowing them to describe what they infer from the image.
But here’s the part most people miss:
This process is approximate.
If an image is blurry, too small, partially hidden, unusual, or ambiguous, the model can still sound extremely confident while being completely wrong.
The pipeline looks simple:
Pixels → Representations → Concepts → Language
But every step can introduce errors. And the model doesn’t always know when it’s guessing.
That’s why AI vision can feel magical one moment and completely dumb the next.
It doesn’t see the way we do.
It reconstructs meaning from visual data — and sometimes, that reconstruction fails.
This is also why human verification still matters, especially when important decisions depend on what an image supposedly contains.
How has been your day frens?
Image failed to load.