TL;DR: AI models like ChatGPT or Perplexity don't recommend keywords — they recommend entities they know. A brand that isn't anchored in knowledge graphs, Wikidata, and consistent third-party coverage doesn't exist for these channels. This post explains why and gives a concrete 5-step checklist.
When you ask ChatGPT today, "Which tool is right for X?", you get a list of brand names. No search results, no blue links — names. And those names don't come from an index that's updated daily. They come from what the model learned during training as reliable, consistent, and prominent. If your brand isn't anchored there, you won't be mentioned. That's the core problem Generative Engine Optimization (GEO) addresses — and entity optimization is the most important lever within it.
What is an entity — and why AI models think differently than traditional search engines
An entity is a uniquely identifiable concept: a person, an organization, a product, a place. In traditional SEO, the matching principle was simple: you optimize a page for a keyword, the search engine matches queries against its index and ranks the hits.
Language models work differently. They don't have an updatable index. They have internalized knowledge — patterns, associations, and relationships distilled from billions of texts. When GPT answers a question about "project management tools," it doesn't pull from a list of hits but from learned associations: Notion is a tool for notes and project management, founded in 2013, popular with indie makers and teams.
The crucial difference: a search engine can rank a page that went live yesterday. A language model can't. It only knows what was described often and consistently enough during training. Your brand has to be anchored in the model as an entity — not as a keyword on a page.
This means traditional on-page SEO is largely ineffective for LLMs. Keyword density, meta tags, internal linking — none of it directly affects whether a model knows your name. What counts is semantic presence in the data sources the model learned from.
Where AI models get their knowledge: understanding knowledge graphs, Wikidata, and training data
Language models are trained on huge text corpora: web crawls (above all Common Crawl), Wikipedia, books, and structured data sources. For entity optimization, the latter are especially relevant.
Knowledge graphs are structured databases that store entities and their properties in machine-readable form. Google's Knowledge Graph is the best known — it powers the info boxes in search results and is an important reference source for various models. A brand that exists there as an entity has structurally higher odds of showing up in model answers too.
Wikidata is the open, editable counterpart — run by the Wikimedia Foundation, freely accessible, and directly linked to Wikipedia. Wikidata entries flow into training corpora and into Google's Knowledge Graph, and serve as a reference for models like Gemini or GPT. A Wikidata entry for your brand is thus one of the most direct and cheapest levers there is.
Wikipedia itself is weighted heavily and is often the primary entity description for models. The problem: strict notability criteria make a dedicated article difficult to impossible for small brands without significant third-party coverage. For most founders, Wikidata is the more realistic entry point.
External coverage and Schema.org: When trade media or relevant publications consistently describe your brand with the same attributes, training corpora treat that as a reliable signal. Consistency is decisive here — contradictory descriptions weaken the entity profile. Schema.org markup (Organization, Product, Person) on your own website helps crawlers capture properties in machine-readable form. Not a differentiator, but a hygiene factor.
Why brands without entity presence stay systematically invisible in AI answers
This is not a marginal problem. The share of information searches answered through LLM interfaces or AI Overviews is growing. For many questions — "What's the best tool for X?", "How do I solve Y?" — an AI answer is now the first point of contact, and often the only one.
In such answers, models don't give a complete market overview. They name entities they know for certain. The principle resembles availability bias in humans: what appears prominently and consistently in the training corpus gets recommended. What is described rarely or inconsistently goes unmentioned.
Structurally disadvantaged are:
- New brands without historical web presence in relevant corpora
- Niche products that are underrepresented in generalized data sets
- Founders without a personal profile who haven't linked their brand to known anchors (people, institutions, industries)
On top of that comes a crowding-out effect: models don't invent consistent brand names — they either know one, or they name one they do know. A model that doesn't know your product name will suggest an alternative it does know. That's the direct opportunity cost of missing entity presence.
The effect is self-reinforcing: whoever gets named in AI answers gets traffic, mentions, and links — which further strengthens entity presence in future training runs. The right time to start was yesterday. The second-best is now.
Entity optimization in 5 steps: a checklist for founders and indie makers
This checklist deliberately focuses on what's doable — no budget for PR agencies, no 6-month roadmap.
-
Create a Wikidata entry
Create an entry at
wikidata.orgfor your brand or for yourself as founder. Required fields: name, description (1 sentence), official website (P856), founding date (P571), industry (P452). Link to existing Wikipedia articles where possible (e.g., the industry or technology). A complete entry is indexable — an empty one does nothing. -
Define a canonical entity description
Write a single sentence that describes your brand unambiguously: what is it, for whom, and what sets it apart? Example: "FORGE is an AI-powered operating system for solopreneurs that manages tasks, finances, and projects in a local environment." Use exactly this wording everywhere: website homepage, about page, Wikidata, GitHub bio, social profiles, guest posts. Consistency is the signal.
-
Implement Schema.org markup
Add
Organizationmarkup on the homepage: name, URL, logo, founding date, description, social links (sameAs). For personal brands:Personschema withjobTitle,worksFor,sameAs. It's no silver bullet, but it makes crawlers' work easier — and takes 30 minutes. -
Build external mentions with the correct description
Guest posts, podcast mentions, interviews, tool directories — anything that places your brand with the canonical description (step 2) in external sources. Quality over quantity: one article in a relevant trade outlet counts for more than ten entries in generic startup directories. Goal: at least 5–10 external sources that describe your name consistently.
-
Build the founder profile as an anchor entity
People are often easier to anchor as an entity than new brands. Complete LinkedIn profile, consistent GitHub bio, own website with Person schema. When you're described in interviews or articles as the founder of [brand], that carries over to the brand as an entity signal. Especially effective for early products without a standalone public profile.
What you don't need to do
- Don't force a Wikipedia article — without genuine third-party coverage it will be deleted
- Don't build keyword-stuffing pages for LLMs — that doesn't work
- Don't wait for quick results — entity presence builds over months, not weeks
Entity optimization is not a sprint. It's the work that ensures you still show up in AI answers two years from now — even after models, interfaces, and ranking algorithms have changed several times by then.