Diwall

English
Download 1.24.4

Why a screenshot is not enough

An image is flat and it is honest about nothing. Four things a page contains that no capture will ever show you — and what happens when an agent assumes otherwise.

Giving a model a screenshot feels like giving it eyes. It is closer to handing someone a photograph of a room and asking them to open the third drawer.

The image is accurate. It is also flat, silent about its own limits, and it omits four things that decide whether an action will work.

One — what lies below the fold

A capture frames a window, not a document. On the recipes index of this site, in the default 1280 × 720 window, Set-of-Mark numbers the 17 interactive elements in view and reports 26 more outside it (measured on 26 September 2026):

"boussole": {"som_hors_viewport": 26}

Diwall says so rather than letting you find out. Without that count, an agent looking at the image concludes the page holds seventeen elements — it holds forty-three, and the agent is wrong about the page, confidently.

The fix is to scroll before acting, but the point here is earlier: the image does not know what it is missing. The JSON does.

Two — what is in the page but not on the screen

A <dialog> that has not been opened exists in the document. Its buttons are real, addressable, and invisible. A collapsed menu holds its links. A tab panel holds its content.

None of it appears on the capture, and all of it is one click away from appearing. An agent that reasons from the image alone believes those elements do not exist; an agent that reasons from the raw document believes they are available now. Both are wrong, in opposite directions.

What decides is whether the element is rendered, not whether it is present. That distinction has no visual form.

Three — what an element is, and where it goes

On a screenshot, a link and a button styled identically are identical. So are a disabled control and an enabled one, when the designer chose a subtle grey.

The page knows the difference. link "Dashboard" /url: "#a" states the role and the destination — before anything is clicked. option "production" [selected] states which choice is currently active. No amount of looking at pixels recovers that.

The accessibility tree →

Four — whether the page is still the same page

A capture is a moment. Between the moment it was taken and the moment an action runs, a banner can close, a modal can open, a list can gain four rows. Anything identified by its position in that image has silently moved.

This is not hypothetical: it is the most common way an automated click lands on the wrong element while reporting success. A screenshot has no way to warn you, because a screenshot has no notion of after.

It clicked, but not on the element you meant →

So what is the capture for

For you.

The image is the thing a person can check in one glance — the layout, the error banner, the button that is grey when it should be blue. It is what makes a report verifiable instead of believable.

The model needs the numbers and the tree; you need the picture. Diwall returns both from the same call because they answer different questions, and neither replaces the other.

In short

  • A capture frames a window; som_hors_viewport tells you what it cut off.
  • Present in the document is not the same as rendered on the screen.
  • Roles, destinations and states have no visual form.
  • An image has no notion of after — positions drift, silently.
  • The picture is for the human. The structure is for the model.