AI crawlers do not start with your best article. They start with the files that tell them what your site is and whether to trust it.
The first is llms.txt, a plain text file at your root that describes what your site is authoritative on and which pages matter most. Most sites do not have one. The ones that do often have a default file their platform generated, which says nothing useful. On the health site, we wrote it by hand: what the company does, which conditions it covers, which pages are reviewed by clinicians, and where the canonical answers live. If you have never written one, a free llms.txt generator by AEO GEO Labs will give you a starting structure; the useful work is filling it with what only you know about your site.
The second is schema. Not just Organization markup on the homepage, but Article, Person, and FAQPage on every content page, with the reviewer’s credentials made machine-readable. An assistant cannot tell that an article was medically reviewed unless the markup says so. We found this gap on nearly every site we audited afterward, including sites with excellent content.
The third is the internal link graph. Assistants follow links to decide which pages on a site are central. If your important pages are three clicks from anywhere, they are peripheral to a crawler no matter how good they are. We rebuilt the internal links so that every article pointed to the product or service it justified, and every product page pointed back to the articles that explained it.
None of this is writing. All of it shipped in the first three weeks, and the citation count started moving before a single new article was published.
Fix two: one page, one question
The second change was structural rather than editorial. Every page got rebuilt so that it answered exactly one question, stated the answer in the first sixty to ninety words, and used headings that were themselves questions.
Assistants quote the sentence that answers the question, not the paragraph that leads up to it. A page that takes four hundred words to get to the point is a page an assistant skims past. The direct-answer block at the top of each article was the single most quoted element across the whole site.
This is also where the medical reviewer went: a named clinician, credentials visible, on every health page. In a YMYL category, that byline is a ranking factor in all but name.
Fix three: content, at volume, on the foundation
Only then did the content pipeline start, at around 500 articles a month, each written to the one-question structure and each linked into the graph on the day it was published. By this point, every new page was being read by crawlers that already knew what the site was and already trusted it, so the pages earned citations within weeks rather than months.
We also shipped a few hundred free tools as indexable pages. Tools attract links without anyone asking for them, and assistants cite tools when a user’s question is “how do I calculate” rather than “what is.”
The part nobody tells you
Those 62,000 citations are not evenly spread across platforms. Most of them came from Perplexity and Google’s AI Mode. ChatGPT was the hardest surface to win and, on this account, citations there actually fell over the period. Anyone who tells you they can guarantee ChatGPT citations is overselling. What you can control is whether your site is structurally eligible to be cited at all, and most sites, even good ones, are not.
That is what the first two fixes do. They do not make your content better. They make it readable to the machines that decide whether anyone sees it. When we run AEO audits at AEO GEO Labs, the three files above are where we look first, and they are where the cheapest wins almost always sit.
Where to start
Check whether your llms.txt exists and whether a human wrote it. Open any article and look at the schema in the source: if you see Organization and nothing else, your reviewer and your FAQs are invisible to assistants. Pick your ten most important pages and count how many clicks they are from the homepage.




