Generative AI Evaluation Intermediate

Prompt Adherence

The yardstick for whether a result matches what was requested

Key points
  • Prompt adherence is the yardstick for whether what's written in the request actually made it into the result.
  • It's a different yardstick from how well-made the picture is. A gorgeous picture that ignores the request still scores low.
  • There are two main ways to measure it: a person counting items one by one, and handing the grading to another model.
  • Items that are plainly visible — count, position, color — are easy to grade; mood or feeling is hard to grade.
  • Forcing adherence up makes the result stiff and unnatural, so where to stop is always a balancing act.
Contents

1The analogy

Stand in front of a dartboard aiming for one section. Land the dart in that section and the score goes up; land it in the next section over and no matter how nicely the throw looked, the score is a different number. It's not about how smooth the throw was — it's about whether it landed where it was aimed.

Looking at a result against the request calls for that same eye. The section aimed at is everything written in the request, and where the dart landed is what actually came out. Ask for three people holding red umbrellas and get two people with blue umbrellas, and no matter how gorgeous the picture is, it missed the section it was aimed at.

There are two ways to keep score. One is a person watching and counting; the other is an automated board that reads where the dart landed on its own. Either way, one throw doesn't tell you much about skill, so several throws get averaged.

2In detail

Well-made and as-requested are different yardsticks

Looking at a generated result, a person notices two things at once. One is craft: are the lines clean, do the colors look natural, are there no odd-looking fingers. The other is adherence: is what was written actually there.

The two move independently. A gorgeous landscape can turn up without the lighthouse that was actually requested, and a rough-looking picture can include every single element that was asked for. Picking a tool by craft alone is exactly why people end up rewriting the same request over and over while using it.

That's why evaluations that compare tools score the two yardsticks separately. Roll them into one combined number and there's no way to tell which side is falling short.

Counting by hand

The most reliable method is a person checking directly. It starts by breaking the request into items: a red umbrella, three people, a rainy night street — pieces that can each be checked. Then, looking at the result, each item gets marked present or not, and the fraction that made it in becomes the score.

Placing two results side by side and picking whichever is closer to the request is also common. Picking one of two wobbles less from person to person than assigning a number does.

The catch is that it's slow and expensive. Standards drift a bit from person to person too, so a single result is usually shown to several people, with how much they disagree recorded alongside the score.

Automated grading

Grading can also be handed to another model instead of a person. A common approach shows the result to a model that describes pictures in words, then checks how closely that description overlaps with the original request. Another approach asks a model that measures pictures and text on the same scale how close the two are.

Lately, a popular method breaks the request into yes-or-no questions and asks them one at a time: is the umbrella red, are there three people. The fraction of yes answers becomes the score — the same item-counting a person would do, just handed over to a machine.

Automated grading is cheap and can get through thousands of pictures in a day. In exchange, whatever the grading model can't see, it never catches — mangled lettering on a sign or an odd shadow can slip right past. So the usual approach is a rough automated pass first, with a person double-checking a sample afterward.

The fields that miss most often

The spots where adherence tends to break down are fairly consistent. First is count: ask for three and two or four often show up instead. Second is spatial relationships: words like left, above, or behind, which set where things sit relative to each other, tend not to hold.

Third is attributes bleeding together: ask for a red hat and a blue bag, and the hat can come out blue with the bag red. Fourth is text: lettering on a sign or a label tends to come out as a similar-looking but meaningless pattern. Fifth is exclusion: ask for something to be left out, and it often shows up anyway.

What gets lost when adherence goes up

There's a dial that pushes the result to follow the request more forcefully. Turn it up and the requested elements do get included, but colors blow out and the result turns stiff. Turn it down and things look natural and varied, but the request gets followed loosely.

Rather than cranking that dial all the way to nail one perfect result, it's often better to set it at a reasonable point, generate several pictures, and pick whichever lands closest to the request. Adherence is a value best judged as an average across several pictures, not a single result's grade.

3More precisely

Several automated metrics exist for measuring adherence. Placing a picture and text in the same space and measuring how close they sit is one; comparing a written description of the picture against the original request is another; breaking the request into yes-or-no questions and counting the answers is a third. None of them lines up perfectly with human judgment, so a new metric usually gets reported alongside how closely it tracks human evaluation. Trust a single metric too much and tune a model to push that number up, and the score can rise while the picture looks no better to a human eye.

The analogy breaks down in one place. A dartboard has sections drawn on it, so everyone reads a landed dart the same way, but a request has no lines drawn on it. Words like a grand mood or a warm feeling have boundaries that shift from person to person. And a dart aims at one section, while a request holds several elements at once, so deciding how much weight each item gets is itself part of the grading.

4Try it yourself

5Common misconceptions

  • It's easy to think a gorgeous result means the request was followed well, but actually craft and adherence move independently, so a pretty picture missing a requested item is common.

  • It's easy to think one adherence score is enough to pick a tool, but actually the score only becomes useful once it shows where it breaks down — count, position, text.

  • It's easy to think writing a longer, more detailed request raises adherence, but actually more items means more chances for something to get dropped.

7One-line summary

In shortPrompt adherence checks whether what was written actually made it into the result, item by item, rather than whether the result looks nice, and it's measured through a mix of human counting and automated grading.

Spotted an error or have a better analogy? Suggest an edit · Last updated2026-09-02