The right way to fact-check AI-generated articles
Language models fabricate numbers, quotes, and sources with perfect confidence, at measured rates high enough to guarantee your drafts contain some. Here is a fact-checking system built for business blogs: what to check, how to check it, and how to make it fast.
Somewhere in your last AI-drafted article, there is probably a claim that isn't true. Not because your tool is broken, but because inventing plausible specifics is what language models do when they don't know something, and they do it in the same confident voice they use for everything else. The draft won't flag it. The tool won't warn you. The first person to notice may be a customer, a competitor, or a lawyer.
The fix is not to stop using AI. It's to run a fact-checking pass that matches how these tools actually fail. This article covers the failure data, the checking workflow, and the shortcuts that make it fast enough to do every time. It pairs with the article on catching hallucinations, which covers what hallucinations are; this one is the operating manual for the verification step itself.
How often AI invents things, measured
The fabrication rates are not folklore. They've been measured repeatedly, and the numbers should recalibrate anyone who skims AI drafts and assumes they're fine.
A 2024 study in the Journal of Medical Internet Research asked models to generate medical literature reviews, then checked all 471 references they produced. GPT-3.5 had invented 39.6% of them. GPT-4 invented 28.6%. Google's Bard invented 91.4%. These were not garbled versions of real papers. Many simply did not exist.
A larger 2026 audit checked 69,557 citations produced by ten different models against three scholarly databases and found hallucination rates between 11.4% and 56.8% depending on the model and topic. The same audit found something useful, which we'll return to: when three or more models independently produced the same citation, it was real 95.6% of the time.
Even experts miss these. An analysis of the NeurIPS 2025 conference, one of the most selective venues in computer science, found over a hundred fabricated citations spread across 53 accepted papers, all of which had passed expert peer review. And the journal-publishing world is seeing the same curve: a 2026 analysis covered by Stat found fraudulent citations, blamed on AI hallucinations, becoming steadily more common in research papers. If reviewers with PhDs are letting fabrications through, "I read it and it seemed right" is not a quality bar.
One more number, because it defines the floor: even when a model is given the correct source material and asked only to summarize it, it still adds unsupported claims a measurable fraction of the time. Vectara's hallucination benchmark, which tests exactly this, puts today's best models at roughly 2-3% of summaries containing something the source doesn't say. There is no setting where checking becomes unnecessary.
The stakes are not hypothetical
Two public episodes made this concrete early. In 2023, a New York federal court sanctioned lawyers whose ChatGPT-assisted brief cited court cases that did not exist. The same year, the tech publisher CNET paused its AI-written finance articles after reviews found errors that forced corrections across the series. Both stories share the detail that matters for you: the output looked professional. Nobody involved thought they were publishing fiction.
For a business blog the exposure is smaller but real. A wrong price, a misstated regulation, an invented statistic in a post with your name on it. Google's quality framework asks whether content is accurate and trustworthy, and fabricated facts are exactly the failure mode that separates AI content that ranks from AI content that eventually gets flagged, by readers if not by algorithms.
The vendors themselves say so. ChatGPT's interface carries a permanent disclaimer, "ChatGPT can make mistakes. Check important info." Anthropic's release notes for its Claude models report progress by citing reductions in false statements, which tells you false statements are an acknowledged property of the product. Google's developer documentation for Gemini recommends "grounding" answers in search results specifically to reduce hallucination. When every maker of the tool tells you to verify the output, verify the output.
What actually needs checking: the triage
The good news is that fact-checking an AI draft is not proofreading. Most sentences in a draft are unfalsifiable connective tissue that needs no verification. The risk concentrates in a short list of claim types, and triage is what makes the job fast.
- Numbers. Statistics, percentages, prices, dates, populations, market sizes. The highest-risk category, because invented numbers look identical to real ones.
- Named sources. "A Harvard study found...", "according to Forbes...". Models routinely attach real institutions to invented findings. The claim and the attribution both need checking.
- Quotes. Anything in quotation marks attributed to a person. Models compose plausible quotes freely.
- URLs and citations. Links and references are fabricated often enough that every one needs a click. A link that resolves can still point somewhere that doesn't say what the draft claims.
- Legal, medical, and financial specifics. Regulations, thresholds, tax rules, dosages. Highest consequence per error, and the rules change by year and jurisdiction while the model's knowledge is frozen at its training cutoff.
- Claims about your own business. The strangest category: models will confidently state your founding year, service area, or pricing, and get them wrong. Readers will assume you, of all people, checked these.
The verification workflow
For each claim the triage catches, the process is five rules.
Rule 1: every specific claim gets a primary source or gets cut. A primary source is the origin of the fact: the study itself, the government agency, the company's own pricing page, not a blog summarizing them. If ten minutes of searching can't produce one, the claim doesn't run. This rule sounds brutal and is actually liberating, because it converts "is this true?" arguments into a mechanical procedure.
Rule 2: read what the source actually says. The subtle failure isn't the missing source, it's the real source that says something adjacent. A study exists but sampled 40 people, not "thousands". The statistic is real but from 2019, presented as current. The page says the opposite once you read past the headline. Check the number, the date, the sample, and whether the source's claim actually matches the draft's sentence.
Rule 3: never verify with the tool that wrote it. Asking the model "is this true?" or "give me a source for this" produces exactly the failure you're checking for: it will generate a confirmation, and often a fabricated citation to go with it. Verification happens in a search engine, a database, or the primary source's own site. The one legitimate model-based trick comes from that 2026 citation audit: agreement between several independent models is real evidence, since consensus citations were correct 95.6% of the time. Use that as a cheap first filter if you like, never as the final word.
Rule 4: when a claim can't be verified, weaken it honestly instead of deleting the paragraph. Often the draft says "studies show 73% of customers...". You can't find the study, but the underlying point is something you know from experience. So say what you actually know: "most of our customers tell us...". A specific claim you can stand behind beats a precise-sounding number you can't. What's never acceptable is keeping the fake precision because it sounds authoritative. That's the exact trade that got the spam sites demoted as untrustworthy content.
Rule 5: keep the receipts. As you verify, paste each source link next to its claim, and keep the list, ideally as links in the published post. This is what professional fact-checking organizations require of themselves: the International Fact-Checking Network's code commits its members to identifying their evidence sources so readers can replicate the check, and to a visible corrections policy. A business blog that links its sources and fixes its errors publicly is borrowing the trust practices of institutions whose entire product is being right, and readers can feel the difference.
The workflow on one real paragraph
Here's the system running on the kind of paragraph every AI draft produces. Suppose your draft about home insulation says: "According to a Department of Energy study, proper insulation reduces heating bills by 45%. Experts at Harvard estimate that 90% of homes are under-insulated, and as insulation specialist John Carver notes, 'most homeowners recoup the cost within two years.'"
Triage flags four items: a statistic with an institutional attribution, a second statistic with a vaguer one, a named-person quote, and an implied payback claim.
The DOE figure gets a search of the actual agency's site. What you'll typically find is something adjacent but not identical, perhaps a real page saying homeowners can save "up to 20%" on heating and cooling under specific assumptions. That's a CLOSE verdict: the source exists, the number doesn't match. The fix is to use the real figure with a link, and note the conditions. The "Harvard 90%" claim traces, if you chase it, to a decades-old industry estimate that's been laundered through years of blog posts with the university's name attached for gravitas; nothing on a Harvard property says it. UNVERIFIED, and it dies. The quote is the easy one: no insulation specialist named John Carver appears anywhere. He's a synthetic expert, invented to carry a plausible claim, and everything he "says" goes. The payback point survives in weakened, honest form if you have something real to anchor it: your own customers' typical numbers, clearly framed as your experience.
Total time: eight minutes. The paragraph that emerges is shorter, correct, sourced, and, not incidentally, more persuasive, because "the Department of Energy estimates up to 20% under these conditions, and here's what our customers actually see" reads like someone who knows the subject, while the original read like everyone else's AI draft, which is exactly what it was. Multiply by the four or five flagged paragraphs in a typical post and you have the whole practice.
Making it fast enough to actually do
A 1,500-word AI draft typically contains five to fifteen checkable claims. With triage, verifying them takes fifteen to thirty minutes, and two habits shrink it further.
First, reduce fabrication at the source. The biggest lever is how you brief the model: give it your facts, your numbers, and your source material, and instruct it to use only what you provided and to mark anything it adds from its own knowledge. Models still slip claims in, which is why the checking pass exists, but a well-fed draft might contain three foreign claims where a bare "write about X" prompt produces fifteen. Tools with retrieval or grounding, which pull from live sources and cite them, shift the job from "find a source" to "confirm the cited source says that", which is faster.
Second, front-load your own material. The claims you supply, your prices, your process, your customer stories, need no external verification, and they're also the substance that makes the post worth ranking. The higher the share of the post that comes from you, the smaller the checkable surface, and the better the post. The two goals point the same way.
The habit that compounds
Run this system for a few months and something changes: fact-checking stops being a chore appended to publishing and becomes the reason your blog is trustworthy in a sea of content that visibly isn't. Readers can't always articulate why one site feels reliable, but linked sources, precise claims, and the absence of that too-smooth fabricated certainty are what they're sensing.
The one-line version: triage every AI draft for numbers, names, quotes, links, and high-stakes specifics, verify each against a primary source you'd be willing to link, weaken what you can't verify into something honest, and never ask the model to grade its own homework. It's fifteen minutes per post, and it's the difference between AI making your blog faster and AI making your blog a liability.