Original data: the content moat AI can't copy
AI made ordinary content free to produce, which made it worthless to rank. The one content asset that got more valuable is data only you have. Here is the evidence that original numbers earn links and AI citations, and how a small business publishes them without a research department.
Every argument about content and AI eventually reaches the same uncomfortable question: if a model can write your article in eight seconds, why would anyone link to it, cite it, or rank it above the eight thousand identical articles generated from the same prompt?
For most content, there's no good answer. But one category of content a model cannot generate at any price: facts that don't exist outside your business. What your jobs actually cost last year. How long your projects really took. What your 400 customers said when you asked. A model can paraphrase those numbers after you publish them, but it can't produce them, and neither can your competitors. In an economy where the words became free, the data behind the words became the moat.
This isn't a slogan. There's now a measurable case that original data is the best-performing content a small business can publish, and the case has three parts: what Google says, what linkers do, and what AI assistants cite.
Google has been asking for this all along
The first question on Google's own self-assessment list for content is: "Does the content provide original information, reporting, research, or analysis?" Not the best-written version of existing information. Original information. Everything in Google's AI-content enforcement record is consistent with that question being the real test: the sites that got erased were pure restatement, and the content that survived contained something that wasn't already free.
There's even a paper trail suggesting how this might work mechanically. A Google patent describes an "information gain score": a measure of how much additional information a document gives a user beyond documents they've already seen. A patent isn't proof of what the ranking system does, so hold it loosely. But you don't need the patent. The logic stands on its own: a search engine drowning in near-duplicate content has every incentive to reward the page that adds something, and only pages with original material can.
The evidence from links
Links remain one of the strongest currencies in ranking, and original data is the most reliably linkable thing a small site can make, because writers need numbers to cite and there are never enough sources.
Ahrefs documented the mechanics in a case study: they built a page of SEO statistics, emailed 515 writers who cited such numbers, and earned 36 editorial links from 32 different websites, a 5.71% conversion rate that cold outreach for ordinary content never approaches. The asset wasn't charm. It was having numbers people needed.
Orbit Media, a web design agency, is the long-running small-business proof. Their annual survey of a thousand bloggers, which they've turned into an institution, has earned links from over 1,600 websites. One yearly research project outperforms hundreds of ordinary posts as a link asset. Their surveying also found that 47% of businesses now use original research in marketing, which cuts both ways: the tactic is proven, and the bar for "original" rises as more people take it up. A niche small-business dataset, though, has almost no competition, because nobody else can gather it.
The evidence from AI citations
The second audience for your data didn't exist five years ago: the AI assistants that now answer questions directly and cite sources while doing it.
The foundational academic work here, a Princeton-led study presented at KDD 2024, tested what makes generative engines more likely to feature a source and found that adding statistics, quotations, and citations to pages boosted their visibility in AI answers by up to 40% across a 10,000-query benchmark. Machines assembling an answer prefer sources that contain concrete, quotable facts. Your data page is exactly that.
The large-scale citation studies show who currently wins. Profound analyzed 680 million AI citations and found Wikipedia alone makes up nearly half of ChatGPT's most-used sources. Ahrefs' analysis of ChatGPT's top 1,000 cited pages found the same skew, Wikipedia at 29.7% and company homepages at 23.8%, with most cited pages sitting on very high-authority domains. Read one way, that's discouraging: the citation economy favors giants. Read correctly, it tells you what the giants have that you can copy in miniature: they are canonical sources of facts. Wikipedia gets cited because it's where facts about topics live. You can become where facts about your niche live, because Wikipedia has no page on Tauranga heat-pump installation costs and never will.
Two more findings sharpen the opportunity. Yext's study of 17.2 million citations found that verified, structured, directly distributed data accounted for more than half of distinct citation sources, machines prefer clean, factual, well-marked-up information. And Ahrefs found 28.3% of ChatGPT's most-cited pages have zero organic search visibility: pages that never won a Google ranking are still being cited by AI. The citation game has different rules than the ranking game, and it's young enough that getting your business cited is still an open field in most niches.
You already have the data
The phrase "original research" makes small business owners picture survey firms and statisticians. Drop the picture. Here's what original data actually looks like at small-business scale, in rough order of effort.
- Your operational numbers. You already know your average job cost, project duration, seasonal demand curve, most common repair, most common mistake customers make before calling. "What we learned from 312 bathroom renovations: the real costs, timelines, and surprises" is original research, and you produced the dataset by existing.
- Aggregate customer patterns. Anonymized and summed: what percentage of your inquiries mention price first, which month breaks the most pipes, what share of clients come from referrals. Ten minutes with your invoicing system produces numbers no one else on earth has.
- A before/after measurement. Track something for your own customers: energy bills before and after insulation, load times before and after a rebuild, callbacks before and after a process change. One honest measured result outweighs pages of "studies suggest".
- A small survey. Forty responses from your actual customers or local market is enough to say something real, provided you say how you got it. Orbit Media's institution started as exactly this.
- A priced index. Collect and publish the going rates in your market: what ten local suppliers charge, how quotes vary by suburb. Compilation of scattered public facts into one dataset is original work, and comparison pages are perennial citation magnets.
Note what every item shares: the effort is in the gathering, not the writing. Which means AI assistance fits perfectly here: the model can structure, phrase, and polish the write-up, because the value was never in the prose. This is the one content type where the machine can do most of the typing and the result is still un-copyable.
Packaging it so machines and writers can use it
A few mechanical choices determine whether your data actually gets cited.
State the numbers extractably. "43% of the 312 renovations we completed in 2025 went over the initial timeline, by a median of 9 days" is quotable by a journalist and liftable by an AI answer. A vague paragraph about how "many projects run late" is neither. Include a short methods note, what you counted, over what period, from what sample, because that's what separates citable data from marketing claims, and it costs three sentences. Put the key numbers in the page text, not only in images. Date the page, and refresh it yearly: Ahrefs found nearly 90% of ChatGPT's dated citations had been updated in the current year, and an annual update turns one asset into a compounding series. Content compounds in general, and a yearly dataset is the purest case: each edition inherits the links and reputation of the last.
Your first dataset in thirty days
To make this concrete, here's the minimum viable version, one month, a few hours total, using data you already possess.
Week one: pick the question. The best first dataset answers a question your customers already ask you constantly, where the honest answer is "it varies", because that's exactly where a table of real numbers is most valuable and least available. "What does a bathroom renovation actually cost here?" "How long does a website project really take?" "What do local suppliers charge for X?" One question, narrow enough that your records genuinely answer it.
Week two: pull and clean the numbers. An evening with your invoicing or job-management system: export the relevant jobs from the last one or two years, strip anything identifying, and compute the handful of figures that answer the question, the range, the median, the main cost drivers, the surprise (there's always a surprise, and it's usually the most quotable finding). Twenty to fifty data points is plenty for a first edition; the bar is honesty about sample size, not statistical grandeur.
Week three: write it up, stats page style. Lead with the key numbers stated extractably, follow with the method note (what you counted, over what period, how many jobs), then the interpretation only you can add: why the range is so wide, what the expensive cases had in common, what you'd tell a customer trying to land at the low end. AI assistance is safe and fast here, since you're supplying every fact.
Week four: put it to work. Link it from your relevant service pages and posts. Mention it where the question comes up, in quotes, in emails, in the community threads where your customers ask. If your niche has writers or newsletters that cover such topics, a short note pointing at the numbers is the entire outreach program, and the Ahrefs case study above is what that looks like when it works.
Then put a recurring reminder in next year's calendar to refresh it with the new year's jobs, and you've started the annual series that the compounding case is built on. The first edition is the hardest and the least rewarded; every subsequent one is cheaper and lands harder.
The honest limits
Original data is a moat, not a shortcut. The first dataset will earn less than you hope; the third, once writers in your niche know you publish real numbers, earns more than you expect, and the correlation studies are blunt about the fact that authority accumulates slowly: Semrush found domain authority and AI mentions correlate strongly (0.65), and authority is built, not announced. Publish on a topic where you genuinely have the best data available, at whatever scale, state your methods, and accept that this plays out over quarters. Meanwhile every competitor whose content strategy is "generate more articles" is building on land that AI floods a little further every month.
The one-line version: in a world where any words can be generated, the only defensible content is content that had to be *measured*, and your business measures things every day that no model and no competitor can fabricate. Count them, publish them with your method, refresh them yearly, and you own the one page in your niche that everyone else, human and machine, has to cite.