The 30-second version
We wrote about what we changed on our own site for AI search and why. This is the companion piece: the actual checklist, in the order we run it. Six items: an llms.txt file that only claims what your pages corroborate, schema with proper entity linking via @id, robots.txt access for the AI crawlers you want reading you, honest dateModified values, answer blocks written to be extracted, and semantic article markup underneath it all. Nothing here requires a developer for more than an afternoon, and every item doubles as ordinary technical SEO. No hype, no guarantees, just the work.
Item 1: llms.txt, and the corroboration rule
llms.txt is a markdown file at your site root (yourdomain.com/llms.txt) that gives language models a curated map of your site: what the business is, what it does, where it operates, and links to the pages that matter, each with a one-line description. Structure it like a table of contents, not a brochure: an H1 with the business name, a short blockquote summary, then sections of links grouped by topic.
The rule that keeps you honest: every claim in llms.txt must be corroborated by a page on the site. If the file says you serve a city, a service page should say so too. If it names a specialty, the site should demonstrate it. You are handing a summary to systems that will also read the full site; the summary and the site need to agree. Keep it current the same way you’d keep a sitemap current, and treat it as claiming only, never inventing.
Item 2: schema entity linking with @id
Most sites have schema; almost none have connected schema. The upgrade is @id: give your Organization one canonical identifier (something like yourdomain.com/#organization), define it fully in one place, and have every other schema block reference that @id instead of re-describing the business from scratch. Same for the site owner or authors as Person entities.
Concretely: your Article schema’s publisher points at the Organization @id. The author points at a Person @id, defined once with name, jobTitle, and sameAs links to real profiles. Your LocalBusiness or ProfessionalService markup carries the same name, address, and phone, character for character, as your footer and your Google Business Profile. What you’re building is a small knowledge graph in which every page agrees about who you are. Machines reward that agreement with confidence, and confidence is the currency of being named in an answer. Run everything through a schema validator when you’re done; broken JSON-LD is worse than none.
Item 3: robots.txt access for AI crawlers
Decide, explicitly, which AI crawlers can read your site, because the default you inherited may not be the policy you want. The main user agents to know: GPTBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot, Google-Extended (Google’s AI training control), and OAI-SearchBot (OpenAI’s search crawler). Check your robots.txt today; some CDNs, security plugins, and boilerplate robots files block these wholesale, which means you may be invisible to AI search without ever having chosen that.
For a local service business, our default is to allow them: you want to exist where customers ask for recommendations. Blocking is a legitimate choice for publishers whose content is the product, but make it a choice. Also verify the crawlers aren’t being challenged by your firewall or bot protection at the CDN level; a robots.txt allow means nothing if every request gets a challenge page. Check your server logs or CDN dashboard for these user agents to confirm they’re actually getting through and getting 200s.
Item 4: dateModified honesty
Dates are a trust signal, and gaming them is a trap. The temptation is to bump dateModified sitewide so everything looks fresh. Don’t. Models and search engines both compare claimed dates against actual content changes, and a site whose every page was “updated” on the same day looks exactly like what it is.
The honest version: show a visible publish date and a visible updated date on articles, mirror both in Article schema (datePublished and dateModified), and only touch dateModified when the content substantively changed. The flip side of honesty is maintenance: actually revisit your important pages, actually update the stale parts, and then the fresh date is earned. A truthful date on genuinely maintained content beats a gamed date on abandoned content in every system that checks.
Item 5: extractable answer blocks
An answer engine builds responses from passages, so write passages built for it. The pattern: a heading phrased the way the question is asked, followed immediately by a two-to-four sentence direct answer that stands alone without the rest of the page. Definition first, nuance after. Then elaborate below for the human reader.
The test we use: copy the paragraph under any heading and read it in isolation. Does it answer the heading’s question, completely, with the subject named rather than pronoun-ed? “PGHDMA is a digital marketing agency in Pittsburgh, PA” extracts cleanly; “we’re located just outside the city” extracts as noise. Add a summary block at the top of long content (we use a “30-second version” section) so the most quotable paragraph on the page is one you wrote deliberately. FAQ sections with real questions and self-contained answers are the same trick in another shape.
Item 6: semantic article markup
The floor everything above stands on: HTML that describes its own structure. One h1 per page. Headings that nest logically instead of being chosen for font size. Real article and time elements, real lists instead of paragraphs pretending, a named byline in the markup. And as little wrapper junk as possible: a page that’s readable with a text-only parser is a page every crawler extracts correctly, old and new. This is where fast, clean, static sites quietly win; there’s nothing between the crawler and the content.
Run the checklist in order, because the items compound: crawler access gets you read, entities get you identified, answer blocks get you quoted, honest dates keep you trusted. If you’d rather have someone run it for you, this now lives inside our SEO service.
Common questions
Is llms.txt actually being used by the AI companies? Adoption is uneven and nobody outside those companies knows precisely how it’s weighted. It costs an hour, follows a public convention, and can’t hurt if it’s honest. We file it under cheap insurance, not magic.
Do I need all six items or is there a shortcut? If you only do two: answer blocks and entity-linked schema. Those align with how answers visibly get assembled today. But item 3 comes first logically; none of it matters if the crawlers can’t get in.
Will any of this hurt my regular Google rankings? No. Every item is either neutral or positive for classic SEO. Clean structure, consistent entities, honest dates, and direct answers were best practice before anyone said “AEO.”
How do I know if it’s working? Watch for AI referral traffic, check whether your pages get cited in AI Overviews for your queries, ask assistants your customers’ questions periodically, and ask every new lead how they found you. Measurement here is young; keep your expectations honest and your notes dated.
