Ready to grow?
Take action star 🚀
29
Sep

Technical AI Readiness: Making Your Site Crawlable, Renderable and Quotable for AI Engines

Technical AI Readiness

Four changes, in this order: allow the AI crawlers by name in robots.txt, serve your content as server-rendered HTML, mark up the entities the engines have to resolve, and keep those entities consistent everywhere you publish.

Most technical SEO for SaaS was built for one crawler that renders JavaScript and follows links patiently. AI engines run a different stack. Your site can pass every check Google sets. It’ll still be unreadable to the systems your buyers now ask first.

The short version: AI engines fail on B2B SaaS sites at four separate layers: crawler access, rendering, structured data, and entity consistency. Each layer has its own failure mode and its own check. Work through them in that order, because a blocked crawler makes the next three impossible to diagnose.

The worked example throughout is our own robots.txt. Its AI bot section names sixteen crawlers and allows every one of them. On the B2B SaaS sites where the app and the marketing site share a domain, layer one is where it breaks.

Why AI engines still can’t see most SaaS sites

Your organic performance may be fine. It just stopped predicting whether an AI engine will quote you.

The two are related, and less tightly than most teams assume. Ahrefs analysed 1.9 million citations across 1 million AI Overviews in July 2025. It found 76% of those citations came from pages already ranking in Google’s top 10. Another 14.4% came from pages that don’t rank in the SERPs at all.

Ranking gets you most of the way into one engine’s answers. It guarantees nothing in the others. They build their own indexes with their own crawlers.

That’s the gap this checklist closes. It’s infrastructure work, separate from the technical SEO work aimed at signup conversion most SaaS teams have already done. One makes the funnel convert. This one decides whether the engines can read the page at all.

Step 1: Allow AI crawlers in robots.txt

Allow each AI crawler by its exact user-agent token, and treat training bots and search bots as separate decisions.

AI crawlers are the individual user-agents that AI vendors send to fetch web content: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User and PerplexityBot among them, alongside control tokens like Google-Extended. Each one is allowed or blocked by name in robots.txt. A single vendor runs several, doing different jobs.

Ai Crawler

The distinction costs visibility. OpenAI’s bot documentation separates GPTBot, which collects content for training foundation models, from OAI-SearchBot, which indexes for ChatGPT search. Sites opted out of OAI-SearchBot won’t be shown in ChatGPT search answers, by OpenAI’s own statement. ChatGPT-User handles live fetches when someone asks a question, and OpenAI documents that robots.txt rules may not apply to it, because a person requested that page.

Block GPTBot to keep your content out of training, and you haven’t touched ChatGPT search. Block the wrong token and you disappear from the answer.

Here are the AI agent lines from our own file at rampiq.agency/robots.txt:

# -------------------------------------------------------------
# AI & LLM Bot Access
# Explicitly allow crawling by AI search agents for discoverability
# -------------------------------------------------------------

# Allow OpenAI bots
User-agent: GPTBot
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

# Allow Perplexity bots
User-agent: PerplexityBot
Allow: /

User-agent: Perplexity-User
Allow: /

# Allow Anthropic bots
User-agent: ClaudeBot
Allow: /

User-agent: Claude-Web
Allow: /

User-agent: anthropic-ai
Allow: /

User-agent: Claude-User
Allow: /

User-agent: Claude-SearchBot
Allow: /

# Allow Meta bots
User-agent: meta-externalagent
Allow: /

# Allow DeepSeek bots
User-agent: DeepSeekBot
Allow: /

# Allow Google Gemini bots
User-agent: Google-Extended
Allow: /

# Allow Bing bots
User-agent: Bingbot
Allow: /

# Allow Apple Intelligence
User-agent: Applebot-Extended
Allow: /

# Allow Common Crawl (widely used LLM training corpus)
User-agent: CCBot
Allow: /

The same file keeps rendering assets open for the engines that do render:
Allow: /wp-content/uploads/
Allow: /*.css$
Allow: /*.js$

Blocking CSS and JavaScript directories is a crawl-budget habit. It now hides the page from anything that builds a DOM before it reads.

Does allowing Google-Extended put you in AI Overviews?

No. Google-Extended controls something else.

Google’s crawler documentation describes it as the token publishers use to manage whether crawled content may train future Gemini models. The same page states that Google-Extended “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.”

Eligibility for AI Overviews and AI Mode runs through Googlebot and ordinary Search indexing instead. Google’s guidance: a page has to be indexed and eligible to appear with a snippet, with no additional technical requirements beyond that.

So the switch you’re looking for isn’t the one most teams reach for. nosnippet, data-nosnippet, max-snippet and noindex all restrict what Google may use in AI features. A page carrying an aggressive snippet limit for brand reasons is opting itself out of the answer. Google-Extended sits allowed, doing nothing for it.

Step 2: Fix JavaScript rendering gaps

Yes, JavaScript rendering blocks AI search engines, for most of them.

Vercel’s analysis of AI crawler traffic across its network found that none of the major AI crawlers render JavaScript. The list covers OAI-SearchBot, ChatGPT-User, GPTBot, ClaudeBot, Meta-ExternalAgent, Bytespider and PerplexityBot. They fetch JavaScript files without executing them: 11.50% of ChatGPT’s requests and 23.84% of Claude’s pulled JS assets that were never run.

Fixed Js Rendering

Gemini is the exception, because it inherits Googlebot’s rendering. Applebot renders too. That study dates from December 2024 and excluded Microsoft Copilot, which had no unique user agent to measure.

Client-side rendering (CSR) ships a near-empty HTML shell and builds the page in the browser. Server-side rendering (SSR) ships finished HTML. For a non-rendering crawler, CSR content doesn’t exist.

This layer catches B2B SaaS hardest. The pages carrying the commercial answer are the componentised ones: feature detail, integrations, pricing tiers.

How to check what an AI crawler actually sees

Request the page the way a non-rendering bot does, and read what comes back:

curl -sL -A "GPTBot" https://example.com/pricing | grep -i "per seat"

If the string you expect isn’t in that output, no non-rendering engine can quote it. Use view-source rather than the browser inspector. The inspector shows you the rendered DOM, which is the one thing these crawlers never build.

Step 3: Add schema AI engines can parse

For B2B SaaS, four types carry almost all the weight: Organization with sameAs, SoftwareApplication, Article and FAQPage.

Each one resolves a different question. Organization establishes who publishes the site and, through sameAs, which external profiles are the same company. SoftwareApplication states what the product is and what category it sits in. Article attributes and dates editorial content.

FAQPage marks question and answer pairs as exactly that. Google stopped showing FAQ rich results in Search on 7 May 2026, by its own structured-data documentation, and it doesn’t ask anyone to remove the markup. The tags still describe the page to any system reading the source.

The sameAs array is the piece doing the entity work:
{
"@context": "https://schema.org",
"@type": "Organization",
"name": "Example Software",
"url": "https://example.com",
"sameAs": [
"https://www.linkedin.com/company/example-software",
"https://github.com/example-software",
"https://www.crunchbase.com/organization/example-software"
]
}

Every URL in that array should resolve, and the name on the other end should match the one you just declared.

Does adding schema actually increase AI citations?

Schema correlates strongly with citation. It hasn’t been shown to cause it.

Ahrefs ran the controlled version of this question. It tracked 1,885 pages that added JSON-LD against 4,000 matched control pages, between August 2025 and March 2026. Google AI Overviews citations fell 4.6%, a small change but statistically significant. AI Mode rose 2.4% and ChatGPT 2.2%, both indistinguishable from zero.

The same study reports the correlation everyone quotes: 53% of AI-cited pages carry JSON-LD, almost three times the rate of pages that aren’t cited.

Louise Linehan, who ran it, states the limit plainly. The treated pages were already heavily cited before the test, with 100 or more AI Overview citations each. The experiment says nothing about a site with no visibility, where schema is doing entity resolution rather than competing for a slot it already holds.

Ship the four types and budget them as infrastructure. Retro-fitting schema across a back catalogue that already gets cited isn’t where the next sprint goes.

Step 4: Build entity and authority signals

Consistency is the signal. An engine has to decide that the company on your homepage, the vendor in a review directory and the name in a podcast transcript are one entity. Until it does, it can’t cite any of them as you.

SEO entities, sometimes called Google entities, are the distinct machine-recognised concepts an engine tracks: a company, a product, a person. Engines connect them through structured data and repeated consistent mentions across the web, not through keyword matching.

Authority Signals

Topical authority is the judgement an engine makes about whether a site has earned the right to be cited on a subject. It accrues across interlinked coverage of a topic, not from any single page.

Four things move both:

  • Spell the company name identically everywhere, legal suffix included, or drop the suffix everywhere.
  • Point sameAs at profiles that exist and agree with each other.
  • Link pillar and cluster pages to each other, so the coverage reads as a body of work.
  • Earn mentions on sites the engine can check independently of you.

llms.txt is a proposed plain-text file listing the pages a site wants LLMs to read. Treat it as optional. Gary Illyes of Google, speaking at a Search Central Deep Dive event in July 2025, said Google “doesn’t support LLMs.txt and isn’t planning to.”

Your technical AI-readiness checklist

Run all four layers against the page types a SaaS site actually has. This is the part of technical SEO for SaaS a general audit skips, because the failure lands on page types rather than on the whole domain.

Page type Crawler access Rendering Schema Entity signals
Homepage Named Allow lines cover it Company description in raw HTML Organization present sameAs profiles resolve and agree
Product and feature pages No app-path Disallow catches them Feature copy in raw HTML SoftwareApplication present Product name identical across pages
Pricing No app-path Disallow catches it Tier names and gates in raw HTML SoftwareApplication presen Plan names match the ones in the app
Docs and help centre Subdomain serves its own robots.txt Content in raw HTML, not an app shell Article on each page Product naming matches the marketing site
Blog and editorial Named Allow lines cover it Body copy in raw HTML Article with dates Author and publisher named consistently
Case studies Named Allow lines cover it Quotes and figures in raw HTML Article present Client names spelled as those clients spell them

Score each cell pass or fail rather than averaging the row into a number.

A fail in the crawler access column voids every cell to its right on that row, because an engine that can’t fetch the page never evaluates the rest. Fix left to right: access → rendering → schema → entities.

Set up tracking before you start fixing. On the Case IQ engagement, a dashboard watching ChatGPT and Perplexity referral traffic ran monthly with the client’s in-house team, alongside the robots.txt and schema work. Without that referral view, a pass on this table is an assumption.

See Where Competitors Beat You in ChatGPT

Are AI tools recommending your competitors instead of you? Understand how AI platforms describe your brand vs. your competitors and the exact steps to flip those results.

Check My AI Visibility
l.kiseleva_small

Frequently asked questions

Can a CDN or firewall block a crawler that my robots.txt allows?

Yes. Bot-management rules at the edge can challenge or block an agent your robots.txt explicitly allows, because those rules run on your server whatever the file says. Check your server logs for the user-agent token, not just the contents of the file.

How long does a robots.txt change take to register?

About a day. OpenAI documents that it can take roughly 24 hours for its systems to adjust after a robots.txt update, and other vendors publish no figure at all. Confirm the change in your server logs rather than assuming it landed the moment you saved the file.

What if my docs or help centre sit on a subdomain?

Each hostname serves its own robots.txt, and an engine reading your root file never sees the subdomain’s. A docs subdomain stood up years ago still carries whatever defaults it shipped with. Check the file at every hostname you publish on.

How do I tell whether an AI engine has actually crawled my site

Filter your server logs by user-agent token and look for the agents you care about. Analytics packages miss them, because non-rendering crawlers never execute the tracking script. Log data is the only reliable record that a fetch happened.

Where this goes next:

Access, rendering, schema and entity consistency are the preconditions. They make a page readable and quotable. They don’t make it worth quoting, which is content work. Crawlable ≠ citable.

Our own engagements have moved on both together. Avassa’s AI Overview visibility grew 2900%, with AI referral traffic up 408%. Case IQ reached 850+ keywords ranking inside AI Overviews, alongside 145% growth in Top-3 rankings.

Both companies came to us after a site migration had cut their visibility. In both, the technical fixes ran alongside GEO content work, so neither number isolates the infrastructure layer.

If you want the diagnosis run across every page rather than the six types above, that’s the AI Search + GEO Audit.

About the author
Liudmila Kiseleva

Liudmila is one of the best-in-class digital marketers and a data-driven, very hands-on agency owner. With top-level education and experience, Liudmila is a true expert when it comes to digital marketing strategies and execution.

Related Posts

B2B SaaS GEO Case Studies: What Actually Changed in AI Citations

Two B2B SaaS clients, real before/after numbers: how Rampiq grew AI Overview visibility and AI referral traffic for Avassa and Case IQ.
09.29.2026
Liudmila Kiseleva
Read more

Technical AI Readiness: Making Your Site Crawlable, Renderable and Quotable for AI Engines

Allow AI crawlers in robots.txt, fix JS rendering gaps, add LLM-ready schema: a checklist so ChatGPT, Perplexity and Gemini can quote your SaaS site.
09.29.2026
Liudmila Kiseleva
Read more

Warm-Audience LinkedIn Ads: The Sequence That Books Meetings Instead of Burning Budget

Turn CRM lists and site visitors into a warm LinkedIn Ads sequence: matched audiences, retargeting windows, and Thought Leader Ads that book sales meetings.
09.28.2026
Liudmila Kiseleva
Read more

How to choose an AI visibility agency

Learn how to choose an AI visibility agency by checking proof, reporting, entity strategy, technical GEO execution, and red flags before signing a retainer.
09.03.2026
Liudmila Kiseleva
Read more

6 Best Generative Engine Optimization (GEO) Courses & Training Programs in 2026

We reviewed the best GEO courses and AI search training programs for 2026, from free certifications to team training built on your own company data.
08.04.2026
Liudmila Kiseleva
Read more

Is your marketing team ready for AI search?

A scored self-assessment to find out whether your marketing team is ready for AI search, covering knowledge, content, measurement, and ownership.
07.31.2026
Liudmila Kiseleva
Read more

Let’s talk!

We would love to hear about your digital marketing challenges and come up with a plan to reach your greatest business goals.

POP-UP