
Large language models (LLMs) like ChatGPT, Gemini, and Claude now act like “meta-browsers” for your customers. People ask questions in a chat box, and the model pulls from the open web—sometimes with direct citations and links—to deliver answers. If your site is clear, crawlable, and trustworthy, it’s more likely to be used as a source, quoted accurately, and clicked.
This guide explains how to get your website ChatGPT-ready. It’s practical, UK-centric, and focused on things you can do without a huge budget. The goal isn’t just “letting AI read your site,” but shaping how your brand and pages are summarised, cited, and recommended.
Before you change settings, be clear about what you want:
Your choice can vary by section. For example, keep in-depth guides open (to attract citations) but block premium content, member areas, or sensitive resources.
LLMs summarise. If your page buries the lede, your key message may vanish. Reformat your most important pages so they’re easy to quote accurately:
Tip: Add a “Key facts” panel on service pages—price ranges, service area, turnaround times, inclusions/exclusions, contact options. These chunks are highly quotable.
Models increasingly prefer sources with obvious credibility. Bake E-E-A-T into your site:
When LLMs summarise, these cues help them present you as a reliable source—and help users click with confidence.
Even the smartest model can’t summarise what it can’t load.
noindex only where necessary.Schema won’t “force” LLMs to cite you, but it clarifies who you are and what a page represents.
sameAs social links).sameAs).offers, availability, aggregateRating and review where valid.Example (Organization JSON-LD):
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "Organization",
"name": "Firepages",
"url": "https://firepages.co.uk/",
"logo": "https://firepages.co.uk/path/to/logo.png",
"sameAs": [
"https://www.linkedin.com/company/your-handle",
"https://x.com/your-handle"
],
"contactPoint": [{
"@type": "ContactPoint",
"contactType": "customer support",
"telephone": "+44-xxx-xxxxxxx",
"email": "[email protected]",
"areaServed": "GB"
}]
}
</script>
Give models (and journalists) a clear, central source of truth:
This “credibility hub” helps LLMs describe your business accurately and drives more consistent summaries.
If you want chat tools to pull reliable facts, make them easy to fetch:
lastmod matters). Consider separate sitemaps for posts, products, video.Many AI crawlers respect robots.txt and X-Robots-Tag headers. Decide what you’ll allow, then implement it precisely.
Common user-agents to consider (examples):
GPTBot (OpenAI)ChatGPT-User (ChatGPT browsing agent for fetching pages)Google-Extended (Google’s opt-out for AI model training)CCBot (Common Crawl)ClaudeBot / anthropic-aiPerplexityBotGooglebot, BingbotAllow most, block training only (example):
# robots.txt
User-agent: *
Disallow:
# Block model-training crawlers (keep normal search open)
User-agent: Google-Extended
Disallow: /
User-agent: CCBot
Disallow: /
Sitemap: https://yourdomain.co.uk/sitemap_index.xml
Block a specific section (e.g., /members/):
User-agent: *
Disallow: /members/
Block AI for a single page via header:
X-Robots-Tag: noai
Notes:
noai/noimageai varies by crawler. Robots.txt rules for specific AI user-agents are the most widely honoured control today./ai-policy (or similar) explaining what’s allowed (e.g., “Quotation with attribution and link is permitted; wholesale training of full articles is not”). This is useful for compliance and goodwill.When LLMs summarise, they often include a source list. Increase your chances of being named and clicked:
If you’re a UK service business or retailer, shape content around the questions people actually ask:
Accessible sites are easier for both people and machines to parse:
Accessibility work tends to improve how your content is extracted and quoted.
You won’t get everything perfect on day one. Build a light measurement loop:
If you change robots/AI settings, document it and annotate analytics so you can see impact.
lastmod, remove render-blocking for core content./ai-policy and optional ai.txt/llms.txt).Getting “ChatGPT-ready” is not a separate channel from SEO—it’s the logical next step. You’re making content clearer, evidence-based, and easier to cite; tightening the technical plumbing so bots (and people) can access it; and setting sensible rules for how AI tools may use your work. Do that consistently, and you’ll show up more often in AI-generated answers with proper attribution—which is exactly where your next customers are looking.
What is llms.txt?
llms.txt (sometimes called ai.txt) is a simple, human- and machine-readable text file you place at the root of your site (e.g., https://yourdomain.co.uk/llms.txt). It states your preferences for how AI agents (large language model crawlers and chat assistants) may use your content — for example, whether they may crawl, quote, cache, or train on it, and any attribution or rate-limit requirements.
Important reality check: llms.txt is not a formal web standard. Some AI crawlers will look for and honour it; some won’t. Treat it as a clear, public policy statement and a courtesy signal. If you want hard enforcement, pair llms.txt with:
robots.txt rules for specific AI user-agents (e.g., GPTBot, ClaudeBot, PerplexityBot, CCBot, Google-Extended for training opt-out).noai, noimageai, noarchive) where appropriate.Think of llms.txt as your policy notice and robots/meta as your enforcement tools.
https://yourdomain.co.uk/llms.txt (root level).#.There is no single canonical schema, so keep it clear, consistent, and unambiguous. Use short key: value pairs and section headers.
You don’t need all of these, but they cover typical needs:
policy: short sentence summarising your stance.contact: email or URL for permissions/queries.updated: ISO date.license: brief note or URL (e.g., “Copyright © Firepages; quotation with attribution permitted”).training: allow | disallow | restrictedsummarisation: allow | disallowquotation: allow | allow-with-attribution | disallowattribution: your preferred name and link format.cache-ttl: guidance for how long models may cache content (e.g., 7d).crawl-delay: advisory seconds between requests (not widely honoured, but harmless).sitemap: absolute URL to your sitemap index.allow: path patterns or sections explicitly permitted.disallow: path patterns or sections to avoid.user-agent: start a block with a specific crawler’s name and override keys below it.# llms.txt — Open with attribution
policy: We allow reputable AI crawlers to access and use our public content for summarisation and model training, provided attribution and a link are given when content is quoted or materially relied upon.
contact: [email protected]
updated: 2025-08-24
license: Copyright © Firepages. Quotation permitted with attribution & link to https://firepages.co.uk/
training: allow
summarisation: allow
quotation: allow-with-attribution
attribution: "Firepages (https://firepages.co.uk/)"
cache-ttl: 7d
crawl-delay: 2
sitemap: https://firepages.co.uk/sitemap_index.xml
# Areas not meant for AI use
disallow: /wp-admin/
disallow: /cart/
disallow: /checkout/
disallow: /my-account/
# Per-agent notes (illustrative; robots.txt should be the enforcement source of truth)
user-agent: GPTBot
training: allow
quotation: allow-with-attribution
user-agent: ClaudeBot
training: allow
user-agent: PerplexityBot
training: allow
user-agent: CCBot
training: disallow # Common Crawl: prefer robots.txt enforcement too
# llms.txt — Controlled access (no training)
policy: You may crawl and generate summaries/quotes of our public content with attribution; model training on our content is not permitted.
contact: [email protected]
updated: 2025-08-24
license: Copyright © Your Company. No model training.
training: disallow
summarisation: allow
quotation: allow-with-attribution
cache-ttl: 3d
sitemap: https://yourdomain.co.uk/sitemap_index.xml
disallow: /members/
disallow: /internal/
# llms.txt — Closed to AI use
policy: We do not grant permission for AI crawling, summarisation, quotation, or model training.
contact: [email protected]
updated: 2025-08-24
training: disallow
summarisation: disallow
quotation: disallow
Note: Back these up in
robots.txt(e.g.,User-agent: GPTBot\nDisallow: /) and, where needed, withX-Robots-Tag: noaiheaders for specific pages.
X-Robots-Tag, <meta name="robots">, <meta name="ai"> patterns adopted by some crawlers) give page-level control.Keep the three consistent. If llms.txt says “training: disallow” but robots allows GPTBot everywhere, you’re sending mixed signals.
If there are parts of your site that shouldn’t be referenced (e.g., ephemeral inventory pages), llms.txt alone won’t prevent that. Use a mix of:
availability: "OutOfStock" so summaries reflect reality.noindex for thin/duplicate/temporary product variants.Example snippet for WooCommerce archives with stock filters:
# robots.txt (example, not llms.txt)
User-agent: *
Disallow: /*?stock_status=outofstock
(Adjust to your theme/plugin’s actual parameter names or use server rules to canonicalise away those URLs.)
/llms.txt.# See our AI use policy at /llms.txt
/ai-policy) that explains your stance in plain English and links to llms.txt. This helps journalists, partners, and compliance teams.GPTBot, ChatGPT-User, ClaudeBot, PerplexityBot, CCBot, and check the paths they request and response codes.robots.txt, llms.txt, and headers still match your intent after platform/theme updates.Given your preference to allow AI access and training while retaining copyright, a good first pass would be:
https://firepages.co.uk/llms.txt using the Open but attributed template above (swap in your contact and sitemap).robots.txt with the same intent (don’t block AI agents you wish to allow; explicitly block ones you don’t)./ai-policy in plain English, mirroring llms.txt.That gives AI tools a clear, consistent green light with attribution — and gives you a single source of truth to point to if questions arise.
Does llms.txt guarantee compliance?
No. It’s a courtesy signal. Combine it with robots and headers for practical control.
Should I use ai.txt or llms.txt?
Either is fine; some sites publish both with identical content. If you choose one, prefer llms.txt and reference it from robots.
Can I demand a specific citation format?
You can request it via attribution: (e.g., “Firepages (https://firepages.co.uk/)”). Some tools will honour this; others will just link your homepage or the specific page they used.
What about images and media?
Include lines like noimageai in HTTP headers where needed, and add a note in llms.txt (e.g., “Image training is not permitted”). Enforce with robots for /wp-content/uploads/ if you truly need to block.


Code 4 Systems Ltd
T/A Firepages
Registered Office
Oxford House
Arnold
Nottingham
NG5 8FB
Registered in England 03860755
