AI Search Optimization Guide
Best Schema for AI Citations: What Actually Helps in 2026
Schema markup is one of the most oversold parts of the AI-visibility stack. Here’s the honest, data-backed version: which schema types are worth adding, what controlled 2026 studies actually found, and why FAQ and HowTo schema aren’t the tactic they used to be.
By Rahul Saini, Author at Search Counsel Co. Last updated August 2026.
The short answer
Schema markup won’t magically get you cited by AI. Controlled 2026 studies found that adding JSON-LD produced little to no lift in AI citations, and when AI engines fetch a page they read your visible text, not your hidden schema. What schema does do is make your brand and content legible: it helps Bing and Copilot, feeds accurate facts into model training, powers the rich results Google still supports, and keeps your data machine-readable. Implement a focused set (Organization, Person, Article, Product, and Dataset for original research), keep it accurate and matched to what’s on the page, and treat it as a foundation, not a citation lever.
Which schema earns its keep in 2026
| Schema type | Best used for | Google rich result? | Why keep it for AI |
|---|---|---|---|
| Organization | Brand identity (+ sameAs) | Knowledge panel signals | Tells engines who you are |
| Person | Author identity | Author details in some results | Machine-readable E-E-A-T |
| Article | Blog and news content | Yes | Attribution and freshness (dateModified) |
| Product / Offer | Ecommerce pages | Yes | Price, availability, rating facts |
| BreadcrumbList | Site structure | Yes | Navigation and context |
| FAQPage | Genuine Q&A content | No (removed May 7, 2026) | Still parsed by Google and other engines |
| HowTo | Step-by-step instructions | No (removed 2023) | Clean step structure for readers and AI |
| Dataset | Original research and data | No (Dataset Search only) | Describes your numbers as data |
Jump to what you need
Article note: Written by Rahul Saini at Search Counsel Co. The claims here are checked against controlled studies from Ahrefs and searchVIU, statements from Google Search Central and Microsoft Bing, and Google’s own structured-data documentation. Rich-result support and study findings change, so verify against the current docs before you rely on any single point.
1) Does schema actually help AI cite you?
Short version: not much, at least not directly. Schema markup is code you add to a page to describe your content in a way machines can read. It’s useful for exactly that. But it is not the thing that gets you cited in an AI answer, and half the agencies selling AI-optimization packages treat it like it is.
Google’s own guidance says the quiet part out loud. Its AI-features documentation states there’s no special schema.org structured data you need to add for AI Overviews or AI Mode, only that any structured data you use should match the visible text on the page. Google’s John Mueller also confirmed in 2025 that structured data isn’t a direct ranking factor. So the honest framing is that schema helps machines understand your entity and structure, and the citation itself is earned by clear, extractable content, corroboration, and authority, as our guide to how AI engines choose sources sets out.
Simple rule: schema tells machines what your content is. It doesn’t make them cite it. Use it for clarity, not as a shortcut.
2) What the 2026 studies found
This is where the honest version parts ways with the marketing. Two kinds of tests have looked at schema and AI citations, and both point the same direction.
Controlled before-and-after studies show little to no lift. In May 2026, Ahrefs tracked 1,885 pages that added JSON-LD schema between August 2025 and March 2026, matched them against 4,000 similar control pages, and measured citation changes across Google AI Overviews, AI Mode, and ChatGPT. Adding schema produced no meaningful citation lift on any platform. AI Mode and ChatGPT sat close enough to zero to count as noise, and AI Overviews showed a small dip the authors don’t attribute to schema. Their explanation is the important part: schema tends to live on better-maintained, higher-authority sites that already publish strong content and earn links, so cited pages have schema by correlation, not because schema caused the citation.
Retrieval tests show AI engines skip your schema at fetch time. A searchVIU experiment tested whether ChatGPT, Claude, Perplexity, Gemini, and Google AI Mode used schema when fetching a page live. None did. During direct retrieval, every system pulled only the visible HTML and ignored hidden JSON-LD, Microdata, and RDFa. In a separate OtterlyAI test, a unique fact was hidden only inside FAQ schema, and no AI platform used it to answer the question, even when pointed at the exact page.
To be fair to the other side, schema isn’t worthless here. Microsoft’s Fabrice Canel confirmed in March 2025 that schema markup helps Microsoft’s LLMs understand content for Copilot, and Bing powers part of ChatGPT search. Google’s Search team said in April 2025 that structured data gives an advantage in classical search results. And there’s a training-time mechanism: when a model is trained on a page, the data pipeline parses your JSON-LD and turns it into plain-language facts that get folded into what the model “knows” about you. So schema can shape the brand facts a model carries across sessions, even though it’s ignored at the moment of retrieval.
Watch out: ignore the marketing numbers. Claims like “2.5x more AI answers” or “300% accuracy from schema” come from vendor pages, not controlled tests, and the controlled tests point the other way. Add schema for the real reasons below, not for a citation multiplier that hasn’t held up.
3) The schema that still matters (and what it does)
None of this means skip schema. It means use a focused set for the right reasons. Here’s the order that earns its keep.
- Organization is your brand’s identity record: name, URL, logo, and sameAs links to your social and directory profiles. Without it, AI systems and knowledge graphs have to guess who you are, and that ambiguity lowers their confidence in citing you. The sameAs links let engines verify your identity across multiple sources. Get this one right first.
- Person ties each article to a named author with their own schema and a real bio page. This is your experience and expertise in machine-readable form, and it pairs with the visible author credentials that actually build trust.
- Article tells engines what a piece is, who wrote it, and when it was published and updated. The dateModified property is the one people skip, and it matters because AI answers favor recent, well-attributed content, which is a good reason to run a real content refresh cycle.
- Product and Offer carry price, availability, and rating facts on ecommerce pages, which is where structured data still produces reliable rich results. Our guide to product schema and rich results covers the detail.
- BreadcrumbList communicates your site structure and is still a supported rich result.
These are the types that survived Google’s pruning of rich results across 2023 to 2026. Product, Review, Article, Organization, LocalBusiness, and BreadcrumbList still produce rich results in Search. Most of the rest do not, which brings us to the two everyone asks about. For the identity side of this, see our guide on entity optimization: making AI understand your brand.
4) FAQPage and HowTo: what changed
If you added FAQ or HowTo schema to grab more space in Google’s results, that era is over. But the content underneath still earns its keep, so don’t rip it out.
The FAQ timeline is worth knowing because the industry keeps garbling it. Google launched FAQ rich results in 2019. In August 2023 it restricted them to well-known, authoritative government and health sites, so most commercial pages lost the feature then. The March 2026 core update cut impressions further. Then on May 7, 2026, Google stopped showing FAQ rich results in Search entirely, with the Search Console FAQ report and Rich Results Test support going away in June 2026 and the Search Console API support in August 2026.
Here’s the line that gets lost in the coverage: Google explicitly said it will keep using FAQ structured data to understand pages, even though the visual result is gone. FAQPage is still a valid Schema.org type, the markup can stay in place without causing problems, and other engines like Bing and Perplexity can still read it. HowTo followed the same arc earlier: its rich results were deprecated on desktop in September 2023 and the documentation was removed, so HowTo no longer shows in Search either.
The takeaway is a clean split between two things that got blurred for years. Rich results were a display feature. Structured data is comprehension. Google ended the display feature. The comprehension layer is still there, and every signal we have says comprehension is where the value sits now. One Ahrefs analysis found only 38% of pages cited in Google AI Overviews rank in the top 10 of regular search, so the signals that win citations aren’t the same ones that win rankings, and clear, question-led answers are among the strongest.
From experience: don’t strip your FAQ sections because the rich result is gone. That Q&A content is doing the work now, in AI answers and featured snippets, and it covers the long tail of questions people actually ask. Removing it quietly costs you coverage.
5) Dataset schema and original research
Dataset is the one type on the list tied to a real AI-visibility play: publishing original data. It describes a dataset, its name, description, creator, and where to access it, so machines can understand your numbers as data rather than prose.
One accuracy note, because it’s widely misstated: Dataset markup feeds Google Dataset Search, a separate vertical, not regular Google Search results, and Google is phasing Dataset out of standard Search displays. So it isn’t a Search rich result. Its value is making your original research legible.
Why that fits AI search: models can’t invent facts, so they reach for pages that supply specific statistics. If you publish benchmarks, a survey, or first-hand findings, a clear on-page data table plus Dataset schema gives both machines and readers something concrete to quote. That’s the same original-data strategy behind our AI Citation Index, and it’s a stronger citation bet than any markup on its own.
6) How to implement schema the right way
The mechanics are simple once you drop the idea that more schema is better.
- Use JSON-LD. It’s the format Google and other engines prefer, and it sits in a script tag separate from your HTML, so it’s clean to maintain.
- Keep it lean. Schema.org has more than 800 types. A normal blog post needs only a couple: Organization plus Article, with a Person author and a BreadcrumbList. Delete the extra types your CMS or plugin added that don’t describe the page.
- Match your visible content. Google’s AI guidance is explicit that structured data should match the text on the page. Marking up things that aren’t visible is the kind of mismatch that gets ignored or flagged.
- Connect entities with @id. Define one Organization and reference it from each Article instead of repeating the block.
- Put it in the raw HTML. Schema injected by JavaScript can be missed entirely by crawlers that don’t render, a point our guide to machine-first architecture expands on.
- Validate. Run the Rich Results Test and the Schema.org validator, and watch Search Console for structured-data errors. Note the Rich Results Test drops FAQ validation in June 2026, so use general structured-data validation for those.
A lean setup for a blog post looks like this. One script block, two entities, a shared Organization reference:
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@graph": [
{
"@type": "Organization",
"@id": "https://yoursite.com/#org",
"name": "Your Company",
"url": "https://yoursite.com",
"logo": "https://yoursite.com/logo.png",
"sameAs": [
"https://www.linkedin.com/company/yourcompany",
"https://x.com/yourcompany"
]
},
{
"@type": "Article",
"headline": "Your article title",
"author": {
"@type": "Person",
"name": "Author Name",
"url": "https://yoursite.com/author/name/"
},
"publisher": { "@id": "https://yoursite.com/#org" },
"datePublished": "2026-01-01",
"dateModified": "2026-01-01",
"mainEntityOfPage": "https://yoursite.com/your-article"
}
]
}
</script>
That’s the whole thing for most pages. Add Product on store pages, FAQPage where you have genuine questions and answers, and Dataset on original research. Every post in this guide runs Article, FAQPage, and BreadcrumbList, so you can view the source of this page to see it in practice.
7) Common schema mistakes
Most schema problems come from a short list of errors.
- Using schema to compensate for weak content or low authority. It won’t. Fix the content and the corroboration first.
- Marking up content that isn’t on the page. A mismatch between your schema and your visible text gets ignored or looks deceptive.
- Boilerplate with only the minimum fields. The optional properties, things like author, image, dateModified, sameAs, and description, are what give machines the context to use your content confidently.
- Duplicate or conflicting blocks. Two Organization entities, or plugin schema fighting your manual schema, creates parsing problems.
- Chasing FAQ and HowTo rich results that no longer exist. Keep the content, drop the expectation of a SERP dropdown.
- Treating schema as the whole strategy. It’s plumbing, not the product.
8) What actually drives AI citations
If schema is the foundation, here’s what sits on top of it and does the real work. It’s the same short list from the rest of this guide: an extractable answer in the first line of each section, corroboration so your claims show up across trusted third-party sources, a clear entity so the engine knows exactly who you are, and genuine authority. Schema supports the entity and structure layers. It doesn’t replace the content and consensus layers.
So the sequence is: set up a lean, accurate schema layer once, then spend your real effort on the things that move citations. Start with how to write answer-first content AI will quote and using statistics and FAQs to boost your AI citation rate, then build the off-site consensus covered in how to build multi-source consensus for AI citations.
How we do it: At Search Counsel Co. we set schema once as a foundation through our [FRAMEWORK NAME] process, then put the effort into extractable answers, original data, and third-party consensus, because that’s where the citations actually come from. If you’d rather hand it off, see our AI SEO and GEO services.
Sources used for this guide
Because this topic is full of marketing claims, this guide leans on controlled studies and primary statements rather than vendor assertions.
| Source | What it supports |
|---|---|
| Ahrefs schema study (Linehan & Guan, May 2026) | 1,885 pages vs 4,000 controls; no meaningful AI-citation lift from adding schema; the correlation-not-causation point. |
| searchVIU and OtterlyAI retrieval tests (2026) | AI engines read visible HTML at retrieval and ignore hidden JSON-LD, Microdata, and RDFa. |
| Google Search Central documentation | FAQ rich results removed May 7, 2026; HowTo removed 2023; no special schema needed for AI Overviews; Dataset used only by Dataset Search. |
| Microsoft Bing (Fabrice Canel, March 2025) | Schema markup helps Microsoft’s LLMs understand content for Copilot. |
| Google Search team statement (April 2025) | Structured data gives an advantage in classical search results. |
| John Mueller, Google (2025) | Structured data is not a direct ranking factor. |
FAQ: schema and AI citations
Does schema markup get you cited by AI?
Not directly. Controlled 2026 studies found little to no citation lift from adding schema, and AI engines read your visible page text at retrieval, not your hidden markup. Use schema for entity clarity and the rich results Google still supports, and earn citations with extractable content and corroboration.
Is FAQ schema still worth adding?
The rich result is gone as of May 7, 2026, so it no longer wins you a SERP dropdown. But keep genuine FAQ content, because Google still uses the markup to understand pages, other engines can read it, and the Q&A answers real questions in AI results and snippets.
Do ChatGPT and Perplexity read schema?
At retrieval, they read the visible HTML and ignore hidden JSON-LD. Schema can still matter at training time, when data pipelines convert it into facts the model retains, but it isn’t the mechanism behind a live citation.
What’s the most important schema type?
Organization. It gives engines your name, URL, logo, and sameAs links so they know who you are and can verify you across sources. Get it right first, then add Person and Article.
JSON-LD or Microdata?
JSON-LD, in a script tag. It’s the format Google and other engines prefer, and it keeps your structured data separate from your HTML, which is cleaner to maintain.
Does schema help my Google rankings?
It isn’t a direct ranking factor, per John Mueller in 2025. It does power the rich results Google still supports and helps engines understand your content, which is reason enough to keep a lean, accurate set. See Google ranking factors for what does move rankings.
Should I remove my FAQ or HowTo schema now?
No need. Google has said unused structured data doesn’t cause problems, both are still valid Schema.org types, and other engines may use them. Just stop expecting a rich result from them.
What is Dataset schema for?
Use it if you publish original data. It describes your dataset for Google Dataset Search (a separate vertical, not regular Search) and makes your numbers legible as data, which supports the kind of original research AI likes to cite.
Conclusion: schema is the foundation, not the lever
Schema markup deserves a spot in your setup, just not the starring role most guides give it. Add a lean, accurate set once, Organization and Person and Article for content, Product for stores, Dataset for original research, keep it matched to what’s visible on the page, and validate it. Then stop expecting it to move AI citations, because the controlled studies say it won’t on its own.
Put the real effort where the evidence points. Move on to how to write answer-first content AI will quote and how to make sure AI crawlers can access and read your site, which do more for citations than any markup. Both sit inside our AI search optimization guide.
Editorial note: This guide is for general marketing and technical education. Structured-data support and study findings change quickly, so verify each point against Google’s current structured-data documentation and the original studies before you rely on it.
