How AI engines decide you're "you": entity resolution and citations
Before an AI engine cites your brand, it has to recognize that scattered references across the web all point at one subject — you. That step, called entity resolution, decides whether your content work compounds into citations or gets attributed to a fragmented identity no engine trusts. Four levers control it: identity consistency across your own pages, third-party corroboration of the same facts, structured data that states them explicitly, and a single canonical home the engines treat as authoritative. Get the resolution right and every citable passage you write has a stable identity to attach to. Get it wrong and even great content loses the slot to a competitor whose brand the engine resolved cleanly.
If you have done the obvious GEO work — allowed the crawlers, written answer-first content, monitored a prompt set — and you are still losing citations to competitors with thinner pages, the bottleneck is almost never the writing. It is upstream, at a step most guides skip: whether the engine has resolved your brand into a single, trusted entity it is willing to name.
This is the mechanism layer underneath how to get cited by ChatGPT. That guide covers eligibility and content shape. This one covers the identity problem that decides whether either of those can land.
The citation is the last step, not the first
A citation looks like the engine picking your page out of a list. That framing hides three stages that happen first:
- Resolve — the engine groups every reference to "your brand" it has seen across crawl, training data, and structured data into one record.
- Retrieve — when a prompt fires, it pulls candidate passages from that resolved entity's content and from the broader web.
- Cite — it attributes the synthesized answer to one or more of those sources.
Most GEO advice jumps to step 3. But if step 1 fails, step 2 returns a weaker candidate set and step 3 attributes the answer to someone else — or to no one. A brand that has written excellent citable atoms can still lose because the engine never consolidated those atoms under one confident identity.
The practical implication: you can measure content and still be measuring the wrong layer. Citation rate moves when retrieval or writing improves, but it has a ceiling set by resolution quality. Raise that ceiling first.
What an "entity" actually is to an engine
An entity is not a homepage. It is a resolved record with attributes: a canonical name, known aliases, a primary domain, a category, and relationships to other entities (founder, parent company, competitors, products). When the engine writes an answer, it is working from that record, not from any single page.
The closest analogy is a knowledge graph node. The engine asks, in effect: given everything I have seen about the subject "PilotCite", what do I confidently know to be true? The answer is only as complete and correct as the facts that were successfully reconciled into that node.
This is why two brands with near-identical content can have wildly different citation outcomes. The first has a clean, well-corroborated entity record. The second has a record full of gaps, contradictions, or competing aliases — so the engine falls back on a safer, better-resolved competitor.
The four levers that decide resolution
Entity consistency is the umbrella term, but it is not one thing. It is four distinct levers, and a weakness in any one of them caps the whole record.
1. Identity consistency across your own surface
The facts on your own pages are the engine's primary source. If your homepage calls you "PilotCite", your about page says "PilotCite, Inc.", your footer says "PilotCite Inc", and your press kit uses "Pilot Cite", the engine has to guess whether those are one entity or three. Guessing is expensive and error-prone, so engines hedge by trusting the brand less.
The same applies to category, founder names, founding date, headquarters, and product names. Pick one canonical form of each fact, state it the same way everywhere you control, and treat any variation as a bug.
2. Third-party corroboration
Engines weight independent repetition heavily. If your homepage claims a category and three independent sources — a directory, a review site, a press article — repeat the same category, that fact is treated as well-resolved. If only your own pages say it, the engine keeps lower confidence.
This is the unglamorous foundation under most "authority" advice. The work is not link-building in the SEO sense; it is making sure the same facts about your brand appear, consistently, on sources the engine already trusts. AI brand hallucinations almost always trace back to a fact that was never corroborated outside your own site.
3. Structured data that states facts explicitly
Schema.org markup lets you hand the engine the attributes in a form it cannot misread. Organization with name, url, logo, sameAs (linking to your official profiles), and foundingDate removes the guesswork from fields that prose renders ambiguously.
sameAs is the lever most teams underuse. Each entry is an explicit statement that "this profile is the same entity as my site" — which helps the engine consolidate your Wikipedia, LinkedIn, Crunchbase, and GitHub presences into one node instead of treating them as unrelated.
4. A single canonical home
Every entity needs one domain the engines treat as authoritative for it. Split your presence across three domains — a marketing site, a docs subdomain on a different root, an app on a third — and you have asked the engine to decide which one speaks for the brand. Consolidation through canonicalization, consistent internal linking to one root, and avoiding duplicate entity records across domains is what lets the engine commit.

How resolution fails
Each of the four levers has a characteristic failure mode. Recognizing which one you have tells you where to spend the next hour.
Fragmented identity. The engine holds two or three partial records for what is actually one brand. Symptoms: you appear in some answers under your full legal name and in others under your product name, and the two never consolidate. Fix: align every surface to one canonical name and wire them together with sameAs.
Alias drift. The engine resolved an alias (a former name, a product nickname, a common misspelling) to the wrong entity — often a competitor or an unrelated company. Symptoms: you get cited for prompts that should name you, but under the wrong label, or not at all. Fix: reclaim the alias explicitly on your own pages and in structured data.
Wrong canonical. The engine resolved you correctly but points the authoritative source at the wrong domain (a regional site, an old domain, a reseller). Symptoms: your content is read, but the entity record is anchored somewhere you do not control. Fix: canonicalize aggressively to one root and sunset redundant domains.
Contradictory facts. Different surfaces state different categories, founders, or dates. Symptoms: the engine's answers describe you imprecisely or conflate you with an adjacent brand. Fix: pick one truth per field, propagate it everywhere, and treat any drift as a regression to fix on the next cycle.
A 30-day entity cleanup
Resolution problems are fixable, but only if you treat the fix as a project rather than a one-off. The sequence below is ordered so each step makes the next one easier to measure.
- Week 1 — Audit your own surface. List every canonical fact (name, legal name, category, founder, founding date, HQ, product names) and every page that states each one. Flag every mismatch.
- Week 2 — Pick one canonical form per fact and propagate it across homepage, about, footer, press kit, and product pages. Treat variants as bugs, not alternatives.
- Week 3 — Add or repair Organization schema with sameAs pointing at every official profile (Wikipedia, LinkedIn, Crunchbase, GitHub, official socials). Each sameAs is a consolidation instruction.
- Week 4 — Close third-party gaps. For each core fact, ensure at least two independent sources repeat it. Prioritize directories and review sites the engines already crawl.
After the 30 days, the entity record is cleaner but the payoff is on the answer side, not on your site. That is where measurement has to move.
How to measure whether resolution is working
You cannot inspect an engine's internal entity record directly. You infer resolution quality from the answers it produces, tracked as rates over a stable prompt set rather than from any single check.
Three signals tell you resolution is improving:
- Mention rate rises without content changes. If you appear in more answers after the cleanup than before, and you did not ship new content, the gain is coming from better resolution — the engine is now confident enough to name you where it previously stayed silent.
- Citation rate catches up to mention rate. A wide gap between the two — you are named but not linked — often means the engine knows the brand but cannot confidently attribute a source. As resolution firms up, citations follow mentions.
- Answer framing stabilizes. Early on, answers describe you inconsistently (wrong category, outdated founder, misspelled name). As the record consolidates, the framing converges on the canonical facts you propagated.
All three are rates, not snapshots. Run them on a schedule against the same prompt set, trend each platform against itself, and treat any meaningful change as a prompt to open the underlying answers and read what shifted. That loop — measure, read, fix, re-measure — is what turns entity work from a one-time project into a maintained asset, and it is the same discipline that keeps AI visibility metrics honest.
Frequently asked questions
No. Keyword targeting is about which prompts your content should answer. Entity resolution is about whether the engine recognizes your brand as the subject of those answers in the first place. You can target the right prompts perfectly and still lose citations if the entity record is fragmented or misattributed.
No. llms.txt is a hint about which pages to read; it does not state entity attributes or consolidate records. Organization schema with sameAs does the consolidation work that llms.txt cannot. They solve different problems and are not substitutes.
The retrieval layer can reflect structured-data and on-page fixes within days of recrawl. Third-party corroboration and training-data effects take weeks to months. Expect a fast bump from your own-site fixes and a slower, compounding gain as independent sources update.
Not directly. Resolution is driven by consistent, corroborated facts, not by link volume. A handful of authoritative sources repeating your canonical facts does more for the entity record than many low-quality links pointing at your homepage.
One canonical block on your primary domain, with sameAs linking out to every official profile. Multiple competing blocks across domains invite the engine to treat them as separate entities, which is the exact problem you are trying to avoid.
