Sales Psychology
Aug 29, 2025
What to Check Before an AI-Written Email Goes Out
AI drafts outreach well and cannot tell when it is wrong. Nine checks that catch the errors a proofread misses, and where they belong in the workflow.

AI drafts outreach well and has no idea when it is wrong. The review that matters is not proofreading, it is checking the small number of things a model cannot verify: whether the claim is true, whether the personalisation is actually about this person, whether the inferred pain is real, and whether the whole thing reads as written by someone who exists. Nine checks, and they take under a minute per email once they are habitual.
What the Model Cannot Know
A language model generating outreach is doing one thing well: producing text that resembles good outreach. It has no access to whether the company detail it used is current, whether the statistic it cited exists, or whether the problem it confidently attributed to this prospect is one they actually have.
That is the entire basis for review. Not that the writing is bad, because it usually is not, but that fluency and accuracy are unrelated. A wrong claim in confident prose is more dangerous than a clumsy sentence, because nothing about it looks wrong.
The failure mode is specific and worth naming: AI outreach fails on plausible-but-false detail. A funding round that closed two years ago described as recent. A named competitor that was acquired. A metric invented because the sentence needed a number. Each one is individually small and each one tells the recipient that nobody was paying attention.
This is the practical content of the human-in-the-loop argument, which we make in principle in human-in-the-loop AI for B2B outbound. This article is the checklist version.
The Nine Checks
In order of how much damage each one prevents:
1. Is every factual claim about the prospect current? Funding, headcount, role, tooling, recent news. Models work from stale training data and from whatever was in the enrichment field, which may itself be a year old. A detail that was true once reads worse than no detail, because it dates you precisely.
2. Is every number real? Any percentage, dollar figure, or multiple has to trace to something you can point at. Models generate numbers to fill the shape of a persuasive sentence; that is the single most common fabrication and the one with the longest tail, because a false statistic in a template goes out thousands of times.
3. Is any source attribution genuine? Phrasings of the form "according to [well-known research firm]" are produced readily and are frequently invented outright. If you do not have the report in front of you, cut the attribution and keep the point as your own observation, which is usually more persuasive anyway.
4. Does the personalisation actually distinguish this prospect? Test it by asking whether the line would be equally true of fifty other companies on the list. "I see you are scaling your sales team" passes as personalisation and fails this test.
5. Is the inferred pain plausible for this role at this company? Models are eager to assert a problem. A 12-person company does not have a lead routing problem; a VP of Engineering does not own pipeline. Mis-assigned pain is more off-putting than no pain at all.
6. Is there exactly one ask? Generated copy drifts toward offering options, which is generous and converts worse. One ask, at the end.
7. Does it sound like a person? The tells are consistent: em dashes, "I hope this finds you well", "in today's fast-paced landscape", three-item lists where two would do, and a closing paragraph that summarises what was just said. Read it aloud; the artificial rhythm is audible before it is visible.
8. Would you send it to someone you know? The most efficient single check. If you would be slightly embarrassed for a peer to receive it, the list should not receive it either.
9. Does it commit to anything you cannot deliver? Models write confident promises about outcomes and timelines. Whatever survives review becomes a commitment the moment someone replies.
Review the Template, Then Sample the Output
Reviewing every generated email individually does not scale and is not where the risk concentrates. Two-stage review works better.
Review the template exhaustively. Errors in the fixed copy are multiplied by your entire send volume, so this is where an hour of scrutiny returns the most. Check every claim, every number, every promise. Nothing goes out until the invariant text is clean.
Then sample the variable output. The personalised portions vary per prospect and are where fabrication and mis-assignment appear. Pull a random sample across your segments, not the first ten rows, which are often the cleanest records in the list. Read the merged result rather than the template with placeholders, because a line that reads fine as a template can be nonsense once a real value lands in it.
Pay disproportionate attention to the highest-value accounts. A generic error to a low-fit prospect costs you one reply. The same error to the account you most want costs you the account, and those are worth reading individually.
The Tells, Specifically
Worth internalising because they are consistent enough to grep for:
Em dashes. Models produce them constantly. Almost nobody uses them in a work email.
Tricolons. "Faster, cheaper, and more reliable." Three parallel items where a person would have written one or two.
Abstract nouns doing the work. Solutions, offerings, capabilities, landscape, ecosystem.
A closing that restates the opening. Correct essay structure, wrong for a five-line email.
Hedged enthusiasm. "I would love to explore how we might potentially help" says nothing while sounding eager.
Perfect parallel structure across bullets. Real people are inconsistent.
Several of these can be caught mechanically before a human reads anything, which is the right division of labour: let a script find the em dashes and the banned phrases, and spend the human attention on whether the claims are true.
What AI Is Genuinely Better At
Being fair about this, because the checklist reads as a case against the tool and is not.
Volume research is the clear win. Reading a hundred company pages and pulling out what each does is work a model does faster and about as accurately as a person, and it does not get bored by the sixtieth. Variant generation is another: producing eight framings of the same point gives a human something to select from, which is easier than writing eight.
It is also good at consistency. A model will not forget the terminology rules or drift in tone across a thousand emails the way a tired writer does.
The division that works: the model drafts and researches, the human decides what is true and what ships. Neither half is optional. Fully manual outreach does not reach useful volume, and fully autonomous outreach sends confident errors at scale, which is worse than sending less.
Where Review Fits in the Workflow
Before the template is approved. Full review of invariant copy: claims, numbers, attributions, promises.
After the first merge, before any send. Sample the personalised output across segments and read the merged version.
Individually for top accounts. Every email to a target account gets read.
Mechanically, on every send. Automated checks for banned phrases, em dashes, empty merge fields, and broken links. Empty merge fields are the most embarrassing and the easiest to catch.
Weekly on replies. Negative replies are the best available signal that something in the generated copy is landing wrong, and they are the only feedback that comes from the actual audience.
An empty merge field deserves special mention. "Hi {{FirstName}}" reaching a real inbox does more damage to a domain's credibility than any subtlety on this list, and it is entirely preventable with one validation rule.
FAQ
Can I skip review if the AI output looks good? Looking good is the problem. Fluency and accuracy are independent, and the errors that matter here are factual ones sitting inside well-formed sentences.
How long should reviewing one email take? Under a minute once the checks are habitual, and most of that is check one and check two. Template review takes much longer and only happens once per campaign.
Do I need to review every single email? No. Review the template exhaustively, sample the variable output, and read every email going to a high-value account individually.
What is the single most common AI error in cold outreach? An invented number, followed closely by a stale company detail presented as current. Both are invisible to a proofread and obvious to the recipient.
Can automated checks replace human review? They replace part of it. A script catches banned phrases, em dashes, empty merge fields, and broken links reliably. No script can tell you whether a claim about a prospect is true, which is the part that actually matters.
Does this slow outbound down enough to matter? Less than the alternative. A campaign paused to fix a fabricated statistic after it has gone out costs far more than the review would have, and the domain reputation damage from poor-quality sending outlasts the campaign.
Want AI-assisted outreach where a human actually signs off before anything sends? That is not a feature at Lidgen, it is the architecture. Book a demo.