Skip to content
AI & Automation16 June 2026 · By the Intense Path Editorial Team

How Do You Review AI-Generated Work Before a Client Sees It?

Reviewing AI output is a search problem, not a proofread. The errors cluster in a few predictable places, so order the pass by what it would cost to be wrong and check the confident specifics first.

On This Page
Pass It On

Found this useful? Send it to someone who’s building.

How to Review AI Output Before a Client Sees It | Intense Path

The deliverable is done. It reads well, it arrived in an afternoon rather than a week, and somebody now has to put their name on it before a client opens the file. Most teams handle that moment the same way: read it from the top at a steady pace, fix a comma, send it. That pass reliably catches typos. It does not catch the one sentence that cites a standard which does not exist.

Reviewing AI output is not proofreading with extra steps. It is a search problem. Mistakes are not spread evenly through generated work. They gather in a small number of places, and attention is the thing in short supply. Spread it evenly and you will find the cheap errors while walking straight past the expensive ones.

Our position, plainly: most review processes are ordered by the document rather than by risk. The reviewer starts at the beginning because that is where documents begin. A better pass starts wherever being wrong would cost the most and only reads straight through once the checkable things have been checked. That ordering is the whole technique, and it holds whether the artefact is a blog draft, a migration plan or a component somebody generated inside an AI-assisted build and nobody has read since.

Why the uniform read fails

Human drafts leak. When a writer is unsure, the prose thins: a hedge appears, a sentence goes clumsy, a paragraph that should be there is missing and the join shows. Reviewers spend careers learning to read those tells, mostly without noticing they are doing it. The signal is not the claim itself. It is the texture around the claim.

Generated drafts do not leak. A sentence carrying a fabricated citation has the same cadence as the sentence beside it, because cadence and accuracy come from different parts of the process. The tell you have relied on for twenty years is simply absent. That is not a small adjustment to how you work; it is the reason careful reviewers miss things in a machine draft that they would never have missed in a colleague’s.

There is a second effect, and it belongs to the reviewer rather than the work. Uniform reading habituates you. By the fourth page you are reading for rhythm, and rhythm is precisely the thing the draft does well. Attention decays across a document while risk does not, which is why the last pages of a long deliverable are so often where the awkward thing was sitting.

Fluency and accuracy are produced by different mechanisms, so a draft reading well tells you almost nothing about whether it is right.

Where the errors actually sit

Three classes account for most of what reaches a client and should not have. They are worth naming separately, because each one is caught by a different kind of looking, and none of them is caught by reading harder.

The confident specific

Anything precise is a candidate. A precise claim is the kind a reader can check in under a minute, and a reader eventually will. Generated work produces these fluently and without any internal marker separating the ones drawn from something real from the ones assembled because the sentence needed a noun in that position.

  • Proper nouns. Product names, standards bodies, regulations, job titles, publications. Half-correct names are the most common single failure and the fastest to be spotted by whoever owns that field.
  • Numbers of any kind. Figures, dates, thresholds, limits, version numbers. If a number is not traceable to a document you can open, it does not go out.
  • Citations and links. Open every one. A link resolving to a plausible page about something adjacent is worse than a broken link, because nobody investigates it.
  • Anything phrased as a rule. "You must", "the specification requires", "this is not permitted". Rules invite reliance, and reliance is where advice quietly turns into liability.
  • Quotations. A quoted sentence attributed to a person or an organisation is either verifiable or deleted. There is no middle position available on that one.

The joins between sections

Local coherence in generated work is strong. Global coherence is weaker. Section two recommends the approach that section five quietly rules out. The same idea acquires two names on page three and page nine. The recommendation at the end does not follow from the analysis in the middle, but it does follow from the heading above it, which is enough to survive a linear read.

Reading straight through is the worst available way to find this, because you are holding one paragraph in your head at a time. The joins want a different motion entirely: headings and first sentences only, in order, at speed.

The omission that nothing points at

Absence has no cue. Nothing on the page announces "here is the exclusion you asked for and did not get". The caveat from the brief, the constraint the client repeated twice on the call, the failure mode you specifically wanted flagged: all of them can be missing from a document that reads as finished, and no amount of careful reading will surface them, because there is nothing there to read.

This is the one class where a written check genuinely beats a good reader. It is also the class that does most damage to trust, because a client who repeated something twice and cannot find it in the output draws a conclusion about attention rather than about tooling. Related territory, if you are still deciding what to hand over in the first place: what can safely be automated in a small team and what has to stay in someone’s hands.

Order the pass by what it costs to be wrong

Five passes, in this order. They are shorter than they sound, because each one looks for exactly one thing and ignores everything else. The instruction that matters is the sequence: check what can be checked before you improve what can be improved.

  1. The checkable pass. Every proper noun, number, date, citation, standard and method name gets confirmed against a source. Looking right is not a check. This is the pass that protects the relationship, and it is the one most often skipped, because it feels mechanical.
  2. The brief pass. Open the brief beside the draft and read the brief first. You are hunting for what is absent, and absence is only visible against the thing that asked for it.
  3. The join pass. Headings and opening sentences only. Contradictions and drifting vocabulary surface in a couple of minutes this way, and stay hidden for an hour in a full read.
  4. The consequence pass. Anything that instructs a machine or a lawyer goes to the person who would have to fix it at two in the morning. Code, configuration, a migration step, contract wording, generated markup.
  5. The read. Now go top to bottom for tone, sense and register. Last, deliberately: polishing a paragraph you are about to delete is the most reliable way to lose an afternoon.

The consequence pass deserves one specific note. Generated markup tends to be syntactically correct and structurally lazy: divs where a landmark belonged, a control that is really a styled span, a label that is placeholder text nobody replaced. Those are ordinary conformance failures against the WCAG success criteria, and they never appear in a read-through, because they are invisible to a person reading with a mouse and working eyes.

A budget for attention

Once you accept that errors are not evenly distributed, review stops being a duration and becomes an allocation. The useful question is not "how long should this take". It is "who or what catches this class, and what does missing it cost".

Error classWhere it hidesWhat catches itCost of missing it
Fabricated specificWherever the tone is most confidentChecking the claim against a sourceHigh, and public
Contradiction across sectionsThe joinsA headings-only skimMedium, and awkward in a meeting
Missing caveat or exclusionNothing points at itThe brief open beside the draftHigh on scope, invisible in the prose
Unsafe code or configurationPlausible, idiomatic snippetsThe engineer who would be pagedHigh, and slow to reverse
Inaccessible generated markupLabels, roles, focus orderA keyboard pass and a checkerMedium, and a stated commitment broken
Vocabulary driftTwo names for one thingA find across the documentLow alone, corrosive across a set
Wrong registerEverywhere at onceThe final readLow, and cheap to fix

The grid is not the point. The point is that almost every row has a different catcher, and one reader doing one pass is only ever one of them. Teams treating review as a single person’s job are choosing, without ever saying so, to catch a single row.

Constrain the input and the review shrinks

Everything above assumes the output arrives and then gets inspected. The cheaper move happens earlier. A model asked to write about your service from general knowledge produces work that has to be checked line by line. A model restricted to your own published pages produces work whose failure modes are narrower, and the review shifts from verifying every sentence to verifying the pages it drew on. That is a different and much smaller job, and it is the practical argument for grounding an assistant in your own content before you worry about how clever it sounds.

The same logic applies to anything answering in your name while nobody is watching. A widget like Flidu answers from the site’s own content and hands over to a contact action when it should, which turns "review every answer" into "keep the source pages correct". You still review. You review something finite.

Before you design the review, look at the input

If a reviewer keeps catching the same class of error, the fix is almost never a better reviewer. It is a narrower brief, a source the tool is allowed to use, or an instruction to say nothing rather than guess. And where an assistant speaks to the public, saying plainly that it is not a person changes how much a wrong answer costs you.

Reviewing past your own competence

A reviewer can only verify what they could have written. Obvious when stated, routinely ignored in practice, because the draft arrives finished and the reviewer arrives willing.

When you can judge it

Inside your own field, generated work is easy to review and slightly insulting to read. You will spot the missing qualification instantly, and you will also notice that the draft has flattened three years of hard-won distinction into one confident paragraph. Both reactions are useful. The first improves the work. The second tells you which parts a reader in that field will also find thin.

When you cannot

Outside it, be honest about what you are doing. Reading a paragraph on tax treatment, or clinical claims, or a jurisdiction you have never worked in, and thinking "that sounds right", is not review. It is agreement. Three responses are legitimate: route it to someone who can judge it, cut the claim, or mark it as unverified somewhere the client can see the label.

Here is the concession this piece owes you. Sometimes the correct answer is the third one, it makes the deliverable visibly weaker, and you send it anyway, because a document with two honest gaps beats a document with two invisible ones. Clients handle "we could not verify this" far better than most people expect. They handle discovering it for themselves considerably worse. How that gets stated is a question of how you use these tools responsibly rather than a question of tooling.

Write the pass down, then keep it short

A review that exists only inside one person’s habits leaves when that person does, and it drifts long before that. Writing it down is worth the twenty minutes. Writing thirty items down is not: long checklists get performed rather than used, which is how a tick-box culture arrives dressed as diligence.

Three rules keep a list honest. Every item names a failure that has actually happened to you, not one you imagine. Items that have caught nothing in six months get retired. And the list lives beside the work, in the same repository or workspace as the deliverable, because a checklist filed in a separate document is a checklist nobody opens.

The list is also the specification for what to automate next. A link checker, a spelling pass against your own product names, a script that flags every number in a draft for confirmation: each one removes a row from the table above and hands the human back the rows that need judgement. That is the good version of automation, and it is the same principle as building automation that removes a handoff rather than adding a dashboard nobody reads. Our parent company has written the reviewer’s side of this argument in the reviewer’s pass, which is worth reading beside this one.

Start with three items, not thirty

Pick the three failures that would most embarrass you in front of a client this quarter. Write those down as checks. Run them for a month, then add one. A list that grows out of real incidents keeps its credibility; a list assembled in a single sitting is being ignored by the second week.

Where we would start

For a small team shipping client work, run two passes rather than five: the checkable pass and the brief pass. Together they cost perhaps a quarter of the reading time and catch the failures that end engagements. Add the join pass once documents get long enough that nobody holds the whole thing in their head at once. Add the consequence pass the moment anything you generate touches production.

And here is the condition under which all of this reverses. If the artefact is internal, reversible and low-consequence, do not review it at all. Notes, first drafts, exploratory options, a throwaway script: review is theatre there, and the honest move is to ship it and fix whatever breaks. Save the pass for work that leaves the building. Deciding which is which is the real governance question underneath putting automated work into an operation, and it is answered per artefact, never once for the whole company.

One last thing, without hedging. The review step is not a tax on using these tools. It is the part that makes them usable at all. A team that reviews well can move quickly with generated drafts, because they know exactly where to look. A team that reviews uniformly moves quickly right up until the afternoon it does not. If you want a second pair of eyes on how your own pass is built, tell us what keeps getting through.

Take these with you
Errors in generated work cluster in the confident specifics, the joins between sections and the things quietly left out, so a review treating every paragraph alike finds the cheap mistakes and misses the costly ones.
Check the facts before you fix the tone, because polishing a section you are about to delete is the most reliable way to lose review time.
Absence is the hardest class of error to see, which is why the brief belongs open beside the draft rather than in somebody’s memory.
If you cannot evaluate a claim yourself, you are not reviewing it, and the honest options are to route it to someone who can, cut it, or label it as unverified.
Constraining what the tool is allowed to draw on shrinks the review before it starts, which is cheaper than reviewing harder.

Common questions.

How long should it take to review AI-generated work?

Budget review time by consequence rather than by length. A single paragraph naming a legal threshold deserves more attention than ten pages of internal notes. A workable rule is that anything a client will act on, publish or sign gets a full checking pass against sources, while reversible internal material gets a skim or nothing at all. If the work would be expensive to retract, the review is the cheap part.

What should you check first in an AI-generated draft?

Check the specifics first: names, numbers, dates, citations, standards, version numbers and anything phrased as a rule. These are the claims a reader can verify in under a minute, and they are where generated work fails most visibly. Tone, structure and flow can be repaired at any point afterwards. A wrong fact that has already reached a client cannot be quietly corrected once the file has been sent.

Why is AI-generated writing harder to proofread than human writing?

Because the usual warning signs are missing. When a person is unsure, the writing tends to thin out: hedges appear, sentences turn clumsy, paragraphs go missing and the seam shows. Reviewers learn to read those signals without consciously deciding to. Generated prose holds the same confident rhythm whether a sentence is correct or invented, so fluency stops being evidence of anything and checking has to become deliberate.

Should clients be told that AI was used in producing their work?

Yes, wherever it affects what they are buying or how they should use the result. Disclosure works best as a standing statement about how the practice operates rather than a note stapled to each file. Say which parts of the process involve automated assistance, who reviews the output, and who remains accountable for it. Silence becomes a problem later, when somebody asks and the answer sounds defensive.

How is reviewing AI-generated code different from reviewing text?

Generated code fails quietly rather than visibly. It usually runs, follows the idioms of the language and survives a casual read, so the defects sit in error handling, edge cases, permissions and newly introduced dependencies rather than in syntax. Review it by running it, reading the failure paths, and checking every dependency it added. The reviewer should be whoever would be called when the thing breaks.

What is the most common mistake teams make when reviewing AI output?

Giving every part of the work equal attention. A steady read from the first line to the last spends the same effort on a heading as on a claim about a regulation, and reviewers habituate as they go, so the later pages receive the least scrutiny. Sorting the work by what it would cost to be wrong, then checking in that order, catches more in less time.

Facing this in your
own business?

Tell us where you’re headed — we’ll map the shortest honest route.

Start a Project