What We Learned Fact-Checking 6,000 AI Newsletter Claims
We ran every claim our AI drafted through a verification layer and kept score. One in four could not be verified. Here is the honest breakdown.
Since May, every newsletter edition NLFactory drafts has been graded by its own fact-checker, and we have kept score the whole time. As of this week the scorecard covers 6,013 individual claims: statistics, dates, events, and attributions pulled out of AI-written drafts and compared against the sources retrieved for each edition. We just published the full running tally, and this post is the honest tour of what it says.
The numbers, without makeup
Of the 6,013 claims checked so far:
- 34% were fully supported. The claim matched what the sources said.
- 35% were partially supported. The core held up, but a detail drifted: a number rounded the wrong way, a date off, a qualifier stronger than the source justified.
- 27% could not be verified. The sources retrieved for the edition did not support the claim as written.
- 4% were unverifiable by nature. Opinions and predictions that no source could settle.
Put differently: roughly one claim in four failed verification, and only 43% of editions came through with zero flags. If those numbers surprise you, they surprised us less. The verification layer exists precisely because we assumed drafts would need it.
What happens to a claim that fails
Nothing that fails verification ships as written. Depending on how a newsletter is configured, a failed claim gets visibly labeled, corrected against the sources, or removed from the edition entirely. Along the way the checker has also rewritten 28 full sections and caught 108 editorial rule violations before delivery. The reader never has to do this work; it happens in the seconds between drafting and sending. You can see where verification sits in the pipeline on our How It Works page.
The finding that changed how we think
The biggest bucket was not outright failure. It was partial support: claims that were almost right. That pattern matters, because an almost-right claim is more dangerous than a wrong one. A reader can smell a wild error, but a statistic that is off by a little reads as authoritative and spreads. It is why we check every claim instead of spot-checking the suspicious ones, and why our curation pipeline treats source citations as mandatory rather than decorative.
Why we are showing our homework
AI-generated content has a trust problem, and it has earned it. Most accuracy numbers you see come from lab benchmarks, measured on test questions under ideal conditions. Ours come from a production system mailing real editions to real inboxes, and we would rather publish the ugly quarter of the pie chart than pretend it is not there. If a company selling AI-written newsletters will not tell you how often its drafts fail verification, it is fair to wonder why.
The scorecard stays public
The full dataset lives on our AI Fact-Check Data page: verdict definitions, month-by-month volume, and the methodology, refreshed daily straight from the production database. It is published under CC BY 4.0, so if you are writing about AI accuracy or automated publishing, the numbers are yours to cite with a link. The figures in this post are the snapshot from July 29, 2026; the page will always have the current ones. We will keep checking every claim, and the scorecard will keep telling on us either way. Since then we have added two companion reads: how the engine checks itself, and five checks you can run on any AI newsletter, ours included.
Quick answers
How many AI newsletter claims has NLFactory fact-checked?
As of July 2026, the verification layer has checked 6,013 individual claims across 1,541 production newsletter editions, and the count grows daily. The live totals, verdict breakdown, and full methodology are published on the AI Fact-Check Data page, which refreshes once a day straight from the production database.
What happens when an AI-written claim fails verification?
Depending on the newsletter's setting, a failed claim is visibly labeled, corrected against the retrieved sources, or removed before the edition is delivered. Nothing that fails verification ships as written. So far the checker has also rewritten 28 full sections and caught 108 editorial rule violations ahead of delivery.
Does a 27 percent failure rate mean AI newsletters are unreliable?
It means unverified AI drafts should not go straight to readers. The 27 percent covers claims the retrieved sources did not support as written, which is exactly why verification runs before delivery. With that layer in place, what actually reaches the inbox is grounded in the sources each edition cites.
Want a newsletter like this?
NLFactory turns any topic into a polished, AI-curated newsletter delivered on your schedule.
Start your free trial →