Research That Could Have Changed the Plan, or It Was Ceremony
Name the finding that would change the plan before you commission the study. If no finding would, you are buying ceremony, and the money is better spent measuring what people actually do.
On This Page

A team runs eight interviews across three weeks, builds a readout deck, presents it to a room that nods in the right places, and then ships the design that was signed off before the first session was booked. Nothing in the study was wrong. The recruiting was decent, the moderator was careful, the synthesis was tidy. It simply never had the power to change anything.
There is one test for this, and almost nobody applies it: name, in advance, the finding that would change the plan. Write it down before you recruit. If you cannot name one, or if the only named one is something nobody believes could happen, you are not running a study. You are running a ceremony that produces slides.
This is not an argument about rigour. Small, scrappy studies change decisions constantly, and expensive ones frequently change nothing. Arguments about UX research value tend to become arguments about method, and they should be arguments about authority: who could act on this result, and were they in the room when the question was set.
What follows is the version of this we apply to our own work: how to set the question, how honest to be about sample size, how to stop writing questions that answer themselves, and how to report a finding so that it survives contact with the most sceptical person in the room.
The question that settles it
Before anything is commissioned, write two sentences on one page. We are deciding between A and B. If we observe Y, we will choose B instead of A. That page takes ten minutes and does more work than any research plan we have ever read.
Three things happen when you try to write it. Sometimes you discover there is no decision, only a desire to know more, which is a legitimate thing to want and a poor thing to pay for on a project budget. Sometimes you discover the decision is already made and the study is being commissioned to make it look considered. And sometimes the sentence writes itself, at which point the method also chooses itself, because the observation you named tells you whether you need to watch people, count people, or read something you already own.
Pre-commit before you recruit
Record four things and circulate them: the decision, the options genuinely on the table, the observation that flips it, and the person who signs either way. That last one matters more than it sounds. Research that reports to nobody in particular gets filed, and the way a project is sequenced usually decides whether the findings arrive while the decision is still open or a fortnight after it closed.
Pre-committing has a second effect people underrate. It stops the study from being re-interpreted after the fact. When a result is uncomfortable, the temptation is to widen the frame until the result no longer applies. A written prediction, dated before fieldwork, makes that move visible to everybody.
Ceremony has tells
You can usually spot a ceremonial study from the calendar invite. These are the signals we look for, and any two of them together are enough to pause the spend and ask what the study is for.
- The research is scheduled after the design review. Sequence is intent made visible. Work that arrives after the decision can only validate it.
- Nobody can describe a bad result. Ask the team what outcome would count as bad news. Silence, or a vague answer, is the answer.
- The participants are the friendliest people available. Colleagues, existing champions and anyone who already understands the product will confirm almost anything.
- The readout has no "we were wrong about this" section. Every honest study overturns at least one internal assumption. A study that overturns none was not aimed at anything.
- Findings arrive as adjectives. Confusing, clean, intuitive, modern. Adjectives are how opinions travel when observations were not recorded.
- The recommendations were drafted before fieldwork closed. It happens more often than anybody admits, usually for good scheduling reasons.
Sample size honesty
Most disputes about sample size are really disputes about what the study is for. Two entirely different jobs get filed under the same word. One is discovery: does this problem exist, where exactly does it happen, and what were people thinking when it did. The other is measurement: how often, which of these is better, and by how much.
What a handful of sessions can prove
Small qualitative studies produce existence proofs, and an existence proof needs one instance. If a person cannot finish the form because the validation message renders below the fold, that is a fact about the interface, not a statistic about users. It stays true if you never test another person. This is the ground the five users or ship and measure argument is fought over, and the honest resolution is that small studies find broken things quickly and cannot tell you how common broken things are.
What they cannot say, ever
They cannot give you rates, preferences across a population, or the size of an effect. Turning five sessions into a percentage is arithmetic dressed as evidence, and once the number is in a slide it outlives every caveat attached to it. Write how many people did this, out of how many we watched, in words, every time. It reads as less impressive and it is the only version a reader can evaluate.
Here is the concession. Sometimes the decision genuinely needs a number and you cannot produce one at the budget available. The honest move is to say so and to name the smaller decision the study can support, rather than producing a number the method cannot carry. We have watched teams lose more credibility to one overstated figure than to any amount of admitted uncertainty.
Questions that answer themselves
The second way a study becomes ceremony is quieter: the questions are written so that only one answer is socially available. Participants are helpful people in an unfamiliar room, and they will tell an attentive stranger what that stranger appears to want.
| Question as asked | What it actually measures | Ask instead |
|---|---|---|
| Would you use a feature like this? | Politeness and imagination | Tell me about the last time you needed this. What did you do? |
| Do you find this page clean and modern? | Agreement with the interviewer | Nothing. Give a task, then stay quiet and watch |
| How important is security to you? | Whether security sounds important | Walk me through how you chose your current supplier |
| Which of these two designs do you prefer? | An aesthetic reaction nobody buys in | Give both a real task and time each attempt |
| Would you pay for this? | Hypothetical generosity | What do you pay for now, and what did you cancel? |
| Was anything confusing? | Recall, filtered by embarrassment | Replay the moment they paused and ask what they were thinking |
Past behaviour beats stated intention
People are reliable historians and unreliable forecasters. Ask what they did last month, what they had open on screen, what they searched for, what they gave up on. Those answers can be checked. Answers about the future cannot, and they are systematically kinder than reality. Sessions run this way also tend to produce fixes that are words rather than layouts, which is why microcopy is the cheapest conversion work available: the observation is specific, and so is the change it implies.
Explaining the design before the task destroys the session, because you have just given the participant the answer key. Set the task, then stop talking. Silence feels unbearable to the person running the study and is the single highest-yield technique available to them.
Measure what you can measure, ask about the rest
A surprising share of the questions teams pay to ask are already answered by data they own. Do not run a study to find out whether the site feels slow. Field measurements of loading, responsiveness and layout stability answer that for everybody who visited, not for the eight people you could recruit. Do not run a study to find out whether people abandon a form; the analytics you already have will show where they stop, and conversion work can test the fix on everyone rather than on a sample.
The same holds for the most common complaint in any organisation, which is that people cannot find things. Search logs, internal query terms and click paths confirm or refute that in a morning, and they point at which section is being looked for under the wrong name. That is the evidence that makes a menu restructure arguable rather than aesthetic.
Measurement has a hard limit, and it is the reason qualitative work exists. Numbers tell you what happened and never why. A drop-off at step three is a fact; whether people leave because they distrust the request, cannot find the information, or assumed they had already finished is a question only a person can answer. Use behaviour to locate the problem and conversation to explain it. Running it the other way around is how studies end up expensive and vague.
Reporting that survives a sceptic
Write the report for the most sceptical engineer in the room, the one who will ask how many people that was and whether the task was realistic. If the document answers those questions before they are asked, it gets acted on. If it does not, it gets thanked.
- The claim, about behaviour rather than feeling. "Four of six people scrolled past the delivery options" rather than "the delivery section felt unclear".
- The evidence, with its provenance. Where it came from, how many people it covers, what task they were given and what device they used.
- Confidence, stated plainly. High, moderate or speculative, plus the specific thing that would raise it. Never a decimal you cannot defend.
- What would falsify it. A finding that cannot be wrong is not a finding. Name the observation that would overturn it.
- Severity separately from priority. How bad it is when it happens is a research question. What gets fixed first is a business question, and conflating them lets whoever owns the roadmap quietly reclassify both.
- The recommendation, with the option you rejected. Showing the discarded alternative is what turns a report into a decision record.
That structure is not unique to research. It is the same skeleton any credible audit uses, which is why a tool like Prooflin resolves findings, severities, priorities and recommendations into reviewable reports rather than into a list of observations. Our parent company has written the long version of the finding itself in the anatomy of a finding, and the buyer-side companion, how to read an audit you were sold, is worth reading before you accept anybody else’s report as settled.
One discipline transfers directly from accessibility work. Where a failure maps to a documented rule with defined inputs and expected outcomes, cite the rule rather than describing the problem in your own words. A result somebody else can reproduce without having been in the room is worth several results that rest on your judgement, and reproducibility is exactly what published test rules exist to provide.
When ceremony is honest, and when it is not
Now the concession, because the argument so far has been one-sided. There are legitimate reasons to run a study that will not change the plan. Research is sometimes commissioned to give an already-correct decision enough social weight to survive a board meeting. Sometimes it exists to settle a standoff between two senior people who will each accept a stranger’s account over the other’s. Both are real uses of money.
What makes them dishonest is the label. Budget them as alignment work, tell the team that is what they are, and nobody wastes three weeks pretending the outcome is open. The damage is not the ceremony. The damage is researchers learning that their findings are decorative, which is how an organisation stops being able to hear bad news at all.
The reverse case matters just as much. Some decisions are cheap to reverse, and for those the fastest research is to ship behind a flag and watch. Others are expensive to undo and deserve real evidence gathered before the direction is set: a pricing model, a navigation restructure, or the decision to rebrand, where the study needs to happen while the answer can still be no.
The rule we would apply
Spend on research in proportion to how expensive the decision is to reverse, not to how important it feels in the meeting. That single rule kills most ceremonial studies and funds the ones that matter. It also settles the perennial argument about whether design work should be researched before or after: before, when undoing it means undoing everything built on top.
A first step that costs nothing. Take the last study your team ran, find the deck, and ask one question of the room: what finding would have changed the plan. If nobody can answer, you have a baseline and a reason to write the two sentences next time. If somebody can, ask whether that finding was actually reachable by the method used. Those two answers are more useful than most research audits.
And if you are about to commission something and cannot tell whether it will change anything, describe the decision to us rather than the study. Send us the decision, the options and the deadline, and the method will be obvious to both of us within a paragraph.
Common questions.
How many people do you need for a usability test?
It depends on the job the test is doing. A handful of sessions is enough to find broken things, because one person failing to complete a task proves the failure exists. It is never enough to measure how often that failure happens, or which of two designs performs better; for that you need behavioural data covering everybody rather than a recruited sample. Decide which job you are doing before arguing about numbers.
What is a leading question in user research?
A leading question is one where only one answer is socially comfortable, such as asking whether a page feels clean and modern. Participants want to be helpful, so they agree. Replace it with a request for history: what they did the last time they faced this situation, what they used, and what they abandoned. Past behaviour can be checked, while opinions about the future cannot.
Should research happen before or after the design is approved?
Before, if the research is meant to influence anything. Work scheduled after a design review can only confirm the decision, because unwinding an approved direction costs more than the study did. If a study genuinely must run late, define in advance what result would trigger a rollback and who has the authority to call it, otherwise the exercise is documentation rather than research.
How do you report qualitative findings without exaggerating them?
State the claim as an observation about behaviour, give the count in words rather than as a percentage, name the task and device, and attach a confidence level with the thing that would raise it. Add what would falsify the finding. Reports written this way survive sceptical readers because the reader can see exactly how much weight the evidence carries.
When is analytics better than talking to users?
Analytics is better at locating problems and worse at explaining them. Use behavioural data to find where people stop, which pages they never reach, and what they search for internally, because it covers everybody rather than a sample. Then use interviews or moderated sessions to find out why that behaviour happens. Running it the other way around produces expensive studies with vague conclusions.
Is it ever acceptable to run research that will not change the decision?
Yes, provided it is labelled honestly. Studies commissioned to build internal agreement or to settle a disagreement between senior stakeholders are a legitimate use of budget. The problem is presenting them as open enquiry. Call that work alignment, fund it accordingly, and keep the discovery budget for questions where the answer could still turn out to be inconvenient.
Facing this in your
own business?
Tell us where you’re headed — we’ll map the shortest honest route.