Step 1: Allow AI crawlers in robots.txt
Allow each AI crawler by its exact user-agent token, and treat training bots and search bots as separate decisions.
AI crawlers are the individual user-agents that AI vendors send to fetch web content: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User and PerplexityBot among them, alongside control tokens like Google-Extended. Each one is allowed or blocked by name in robots.txt. A single vendor runs several, doing different jobs.

The distinction costs visibility. OpenAI’s bot documentation separates GPTBot, which collects content for training foundation models, from OAI-SearchBot, which indexes for ChatGPT search. Sites opted out of OAI-SearchBot won’t be shown in ChatGPT search answers, by OpenAI’s own statement. ChatGPT-User handles live fetches when someone asks a question, and OpenAI documents that robots.txt rules may not apply to it, because a person requested that page.
Block GPTBot to keep your content out of training, and you haven’t touched ChatGPT search. Block the wrong token and you disappear from the answer.
Here are the AI agent lines from our own file at rampiq.agency/robots.txt:
# -------------------------------------------------------------
# AI & LLM Bot Access
# Explicitly allow crawling by AI search agents for discoverability
# -------------------------------------------------------------
# Allow OpenAI bots
User-agent: GPTBot
Allow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
# Allow Perplexity bots
User-agent: PerplexityBot
Allow: /
User-agent: Perplexity-User
Allow: /
# Allow Anthropic bots
User-agent: ClaudeBot
Allow: /
User-agent: Claude-Web
Allow: /
User-agent: anthropic-ai
Allow: /
User-agent: Claude-User
Allow: /
User-agent: Claude-SearchBot
Allow: /
# Allow Meta bots
User-agent: meta-externalagent
Allow: /
# Allow DeepSeek bots
User-agent: DeepSeekBot
Allow: /
# Allow Google Gemini bots
User-agent: Google-Extended
Allow: /
# Allow Bing bots
User-agent: Bingbot
Allow: /
# Allow Apple Intelligence
User-agent: Applebot-Extended
Allow: /
# Allow Common Crawl (widely used LLM training corpus)
User-agent: CCBot
Allow: /
The same file keeps rendering assets open for the engines that do render:
Allow: /wp-content/uploads/
Allow: /*.css$
Allow: /*.js$
Blocking CSS and JavaScript directories is a crawl-budget habit. It now hides the page from anything that builds a DOM before it reads.
Does allowing Google-Extended put you in AI Overviews?
No. Google-Extended controls something else.
Google’s crawler documentation describes it as the token publishers use to manage whether crawled content may train future Gemini models. The same page states that Google-Extended “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.”
Eligibility for AI Overviews and AI Mode runs through Googlebot and ordinary Search indexing instead. Google’s guidance: a page has to be indexed and eligible to appear with a snippet, with no additional technical requirements beyond that.
So the switch you’re looking for isn’t the one most teams reach for. nosnippet, data-nosnippet, max-snippet and noindex all restrict what Google may use in AI features. A page carrying an aggressive snippet limit for brand reasons is opting itself out of the answer. Google-Extended sits allowed, doing nothing for it.
Step 2: Fix JavaScript rendering gaps
Yes, JavaScript rendering blocks AI search engines, for most of them.
Vercel’s analysis of AI crawler traffic across its network found that none of the major AI crawlers render JavaScript. The list covers OAI-SearchBot, ChatGPT-User, GPTBot, ClaudeBot, Meta-ExternalAgent, Bytespider and PerplexityBot. They fetch JavaScript files without executing them: 11.50% of ChatGPT’s requests and 23.84% of Claude’s pulled JS assets that were never run.

Gemini is the exception, because it inherits Googlebot’s rendering. Applebot renders too. That study dates from December 2024 and excluded Microsoft Copilot, which had no unique user agent to measure.
Client-side rendering (CSR) ships a near-empty HTML shell and builds the page in the browser. Server-side rendering (SSR) ships finished HTML. For a non-rendering crawler, CSR content doesn’t exist.
This layer catches B2B SaaS hardest. The pages carrying the commercial answer are the componentised ones: feature detail, integrations, pricing tiers.
How to check what an AI crawler actually sees
Request the page the way a non-rendering bot does, and read what comes back:
curl -sL -A "GPTBot" https://example.com/pricing | grep -i "per seat"
If the string you expect isn’t in that output, no non-rendering engine can quote it. Use view-source rather than the browser inspector. The inspector shows you the rendered DOM, which is the one thing these crawlers never build.
Step 3: Add schema AI engines can parse
For B2B SaaS, four types carry almost all the weight: Organization with sameAs, SoftwareApplication, Article and FAQPage.
Each one resolves a different question. Organization establishes who publishes the site and, through sameAs, which external profiles are the same company. SoftwareApplication states what the product is and what category it sits in. Article attributes and dates editorial content.
FAQPage marks question and answer pairs as exactly that. Google stopped showing FAQ rich results in Search on 7 May 2026, by its own structured-data documentation, and it doesn’t ask anyone to remove the markup. The tags still describe the page to any system reading the source.
The sameAs array is the piece doing the entity work:
{
"@context": "https://schema.org",
"@type": "Organization",
"name": "Example Software",
"url": "https://example.com",
"sameAs": [
"https://www.linkedin.com/company/example-software",
"https://github.com/example-software",
"https://www.crunchbase.com/organization/example-software"
]
}
Every URL in that array should resolve, and the name on the other end should match the one you just declared.
Does adding schema actually increase AI citations?
Schema correlates strongly with citation. It hasn’t been shown to cause it.
Ahrefs ran the controlled version of this question. It tracked 1,885 pages that added JSON-LD against 4,000 matched control pages, between August 2025 and March 2026. Google AI Overviews citations fell 4.6%, a small change but statistically significant. AI Mode rose 2.4% and ChatGPT 2.2%, both indistinguishable from zero.
The same study reports the correlation everyone quotes: 53% of AI-cited pages carry JSON-LD, almost three times the rate of pages that aren’t cited.
Louise Linehan, who ran it, states the limit plainly. The treated pages were already heavily cited before the test, with 100 or more AI Overview citations each. The experiment says nothing about a site with no visibility, where schema is doing entity resolution rather than competing for a slot it already holds.
Ship the four types and budget them as infrastructure. Retro-fitting schema across a back catalogue that already gets cited isn’t where the next sprint goes.
Step 4: Build entity and authority signals
Consistency is the signal. An engine has to decide that the company on your homepage, the vendor in a review directory and the name in a podcast transcript are one entity. Until it does, it can’t cite any of them as you.
SEO entities, sometimes called Google entities, are the distinct machine-recognised concepts an engine tracks: a company, a product, a person. Engines connect them through structured data and repeated consistent mentions across the web, not through keyword matching.

Topical authority is the judgement an engine makes about whether a site has earned the right to be cited on a subject. It accrues across interlinked coverage of a topic, not from any single page.
Four things move both:
- Spell the company name identically everywhere, legal suffix included, or drop the suffix everywhere.
- Point sameAs at profiles that exist and agree with each other.
- Link pillar and cluster pages to each other, so the coverage reads as a body of work.
- Earn mentions on sites the engine can check independently of you.
llms.txt is a proposed plain-text file listing the pages a site wants LLMs to read. Treat it as optional. Gary Illyes of Google, speaking at a Search Central Deep Dive event in July 2025, said Google “doesn’t support LLMs.txt and isn’t planning to.”
Your technical AI-readiness checklist
Run all four layers against the page types a SaaS site actually has. This is the part of technical SEO for SaaS a general audit skips, because the failure lands on page types rather than on the whole domain.
| Page type |
Crawler access |
Rendering |
Schema |
Entity signals |
| Homepage |
Named Allow lines cover it |
Company description in raw HTML |
Organization present |
sameAs profiles resolve and agree |
| Product and feature pages |
No app-path Disallow catches them |
Feature copy in raw HTML |
SoftwareApplication present |
Product name identical across pages |
| Pricing |
No app-path Disallow catches it |
Tier names and gates in raw HTML |
SoftwareApplication presen |
Plan names match the ones in the app |
| Docs and help centre |
Subdomain serves its own robots.txt |
Content in raw HTML, not an app shell |
Article on each page |
Product naming matches the marketing site |
| Blog and editorial |
Named Allow lines cover it |
Body copy in raw HTML |
Article with dates |
Author and publisher named consistently |
| Case studies |
Named Allow lines cover it |
Quotes and figures in raw HTML |
Article present |
Client names spelled as those clients spell them |
Score each cell pass or fail rather than averaging the row into a number.
A fail in the crawler access column voids every cell to its right on that row, because an engine that can’t fetch the page never evaluates the rest. Fix left to right: access → rendering → schema → entities.
Set up tracking before you start fixing. On the Case IQ engagement, a dashboard watching ChatGPT and Perplexity referral traffic ran monthly with the client’s in-house team, alongside the robots.txt and schema work. Without that referral view, a pass on this table is an assumption.