What the screen owes

Design quality stays an argument about taste until an organisation agrees what each moment owes the person using it.

Quality is the word design organisations use most and define least.

Everyone in a product review holds a definition. For one designer it is polish, for another clarity. The product manager means conversion, the engineer means stability, the brand lead means distinctiveness, the researcher means whatever hurt in the last round of interviews. Each definition is reasonable, and each is private. Reviews that sound like arguments about the work are often arguments between definitions, and they tend to be settled by the most senior one in the room. That is what taste usually amounts to inside an organisation: a private definition of quality with seniority attached.

Setting quality, as a leader, means replacing those private definitions with one shared object that people can point at. I have come to think the object has to do two things that pull against each other. It has to be small enough to use in a thirty-minute review. And it has to keep apart the questions a single score would average together, because the tensions between them are where the design decisions get made.

Four questions, asked from the person’s side

The standard I use has four questions, each written as something a person would ask of the product.

Fit asks whether this works for me. It covers whether the product delivers on its promise without unnecessary effort: what is available, how easily it is found, how little the person has to translate their need into the product’s terms.

Emotion asks how it makes me feel. It covers the response to craft, motion, tone and identity, and whether the product feels like itself and nobody else.

Legitimacy asks whether I can rely on them. It covers honest pricing, realistic promises and fair responses when things go wrong. It also covers craft that holds together, which surprises people until they see it. Small inconsistencies, like a misaligned grid or a badge in three sizes, tell people the company has stopped paying attention. Inconsistency registers as a trust signal before it registers as an aesthetic one.

Tenure asks whether this is worth staying with. It covers commitment over time, the one question no single session can answer.

Writing each answer from the person’s side changes what the review measures against. A team tends to judge a screen against the ticket that produced it, and the person using it never saw the ticket.

Why the readings stay apart

The same moment reads differently through each question, and the differences are the point. A one-tap reorder is strong on fit and builds tenure, and it is emotionally flat, which for a routine purchase may be exactly right. A recommendation that appears to come from a friend narrows the choice and warms the moment, and it spends legitimacy the instant a paid placement can pass for one. An AI prompt that feels warm on first use has to earn trust again with every answer it gives. Averaged into one score, each of these looks fine. Read four ways, each shows the decision that actually needs making.

The most expensive version of this failure is tenure read alone, because retention is easy to measure and easy to raise by the wrong means. In September 2025 Amazon agreed to pay $2.5 billion to settle the US Federal Trade Commission’s case over Prime, which alleged that people were steered into subscriptions they had not meant to take and that cancelling was made deliberately difficult. By the FTC’s account, the cancellation flow was known inside the company as Iliad. Amazon admitted no wrongdoing. Whatever a court would have found, the shape is instructive. A retention number can rise while the relationship it is supposed to measure decays, and only a separate legitimacy reading would show it.

Evidence for each question

A standard is only as good as the evidence that can answer it, and the four questions need different kinds.

What people do shows up in behavioural signals: completion, reuse, return, abandonment. Read alone, behaviour rewards compulsion, because a product people cannot leave looks exactly like a product people love. What people say, asked inside the product at the moment it matters, shows up in short attitudinal touchpoints. Read alone, attitude flatters, because the people who answer are disproportionately the ones who stayed. What trained eyes can see, and people feel without naming, shows up in scored craft audits of type, spacing, motion, copy and consistency. Read alone, an audit drifts back into taste. Scored against the four questions, it becomes accountable to something outside the auditor.

Tenure needs its own treatment, because a walkthrough or an audit can only hypothesise it. So a tenure reading is written as a hypothesis, with the signal to watch and the metric that would test it. That keeps the most consequential question honest about how little anyone knows on the day a screen ships.

What changes in the room

I use these four questions as the working quality standard in the design organisation I lead: in reviews, in craft audits, and in benchmarking against competitors. The clearest change is in the review itself. The conversation moves from whether a screen is good to what it owes the person using it, and which of the four debts it is paying or deferring. The question settles the argument, and seniority loses its casting vote. A junior designer can hold a director to the legitimacy of a pricing pattern without winning a debate about taste, and a director can explain a call in terms the whole team can check.

A standard earns its place when the people in a review can disagree about the work, and stop disagreeing about the word.


This thinking has ancestors. Quality has long been treated as more than one thing: Garvin’s eight dimensions of quality and the SERVQUAL model of Parasuraman, Zeithaml and Berry both divide it, and Google’s HEART framework, from Rodden, Hutchinson and Fu, pairs attitudinal and behavioural measures for user experience. Each question has its own ancestry: task-technology fit, in Goodhue and Thompson; emotional design, after Norman; organisational legitimacy, in Suchman; and the commitment-trust theory of relationships, in Morgan and Hunt. The study of deceptive interface design, from Brignull’s naming of dark patterns onward, documents what happens when tenure is pursued without legitimacy.

What is sharpened here is organisational. A quality standard written from the person’s side holds four readings apart, pairs each with the evidence able to answer it, and keeps tenure as a hypothesis until time can test it.