If you've ever wondered why ChatGPT confidently recommends Notion for notes, Figma for design, and Linear for project management while your product sits completely ignored - you're asking exactly the right question.
AI search engines don't pick tools randomly. But they also don't follow a clean, inspectable algorithm you can game with a few tweaks. Understanding what's actually happening under the hood is the difference between a useful GEO strategy and a wasted quarter chasing the wrong things.
This post breaks down how ChatGPT and Perplexity actually work (they're different), which signals research shows matter most, what you can realistically influence, and where most founders waste effort.
ChatGPT and Perplexity work differently - and that changes everything
Before talking about signals, you need to understand that these two engines have fundamentally different architectures. Treating them as the same thing will send you in the wrong direction.
ChatGPT (in base mode without search enabled) draws primarily from its training data. That data has a cutoff date. If your product didn't have editorial coverage, backlinks, or independent mentions before that cutoff, ChatGPT's internal model simply doesn't know you exist. When users ask commercial or comparison questions, ChatGPT does trigger live web search - but the base layer of brand familiarity still comes from training.
Perplexity works differently. It retrieves live web content on virtually every query. Fresh, indexed content can influence Perplexity's responses within days or weeks rather than months.
The practical implication: if you're a new product trying to appear in ChatGPT's base answers, there's a real ceiling on how fast you can move. Perplexity is more accessible in the short term because it's reading the live web right now.
The signals that actually matter
In May 2026, Cyrus Shepard (Zyppy) synthesized 54 studies on AI citations and scored 23 ranking factors from 0 to 10. ProCloser.ai then field-tested that framework across five client GEO programs over 12 months. (Source: ProCloser.ai, May 2026) Here's what actually moved citation rates.
1. Google ranking - still the foundation (score: 9.4/10)
The single highest-impact factor for AI citation visibility is still ranking in Google's top 10 for the target query. Ahrefs research confirms that 38% of Google AI Overviews citations come from results already in the top 10.
This surprises founders who assumed GEO would replace SEO. It won't. AI systems search the live web to generate their answers. Pages that rank well get surfaced first. Solid on-page SEO and quality do-follow backlinks feed GEO directly. These aren't separate tracks - they're the same track.
2. Fan-out coverage (score: 9.3/10)
"Fan-out coverage" means mentions across many independent sources - not just your own blog or social accounts. Review platforms, tech newsletters, community threads, comparison articles, podcast transcripts, GitHub discussions. The more independent places reference you, the more signal AI systems have that you're a real, legitimate option in your category.
A brand that exists only on its own website is a brand AI systems treat as thin or unknown. Training data is built from the broader web. If the broader web barely mentions you, training data barely knows you.
3. Explicit, clear phrasing (score: 8.1/10)
This is the easiest fast win. ProCloser found that stripping hedge language from content moved citation rates within two weeks across all five of their test clients.
What that means in practice: instead of writing "some teams prefer [Product] for X use cases," write "[Product] does X." AI systems synthesize confident, citable claims into answers. Vague, hedged writing gets filtered out. Clear statements get picked up.
Check your own site. Homepage copy, feature pages, comparison content. If every sentence has "some," "many," "often," or "typically" padding it, that's immediate work to do. It costs nothing and the research suggests it genuinely moves the needle.
4. Structured data and schema markup
Schema.org markup helps AI systems parse your content reliably. Product schema, FAQ schema, Organization schema. It's not a dramatic factor - but it reduces friction between your content and the retrieval system trying to understand it. Think of it as speaking the machine's language more clearly.
5. Wikidata and entity presence
Research from OppAlerts, analyzing 105,000+ ChatGPT prompts across 145 industries, found Wikidata was the dominant signal in multiple verticals: ERP software (R² of 42.9%), Furniture (R² of 49.9%), Hotels (R² of 42.3%). (Source: TheDigitalBloom.com)
Most SaaS products have no Wikidata entry. This is a genuine gap - though Wikidata has real notability requirements. You need independent sources to back up the entry, not just self-referential content from your own site.
6. Community discussion
OppAlerts found Reddit highly predictive for categories where users ask for real opinions: Enterprise AI platforms (R² 27.9%), live entertainment (25.3%). AI systems trained on human conversation absorb community discussions at scale.
Founders who participate genuinely in relevant communities - answering questions, sharing honest comparisons, being present - build real signal here over time. Spam-posting product announcements into communities does the opposite.
The honest reality: most of this is inside the model
Here's what most GEO content won't tell you.
OppAlerts' dataset found that the best single external predictor - search engine appearances - explains only 5.8% of variance in LLM recommendation scores. The other 80-85% comes from inside the model itself: patterns baked in at training time, which you cannot directly access, edit, or petition.
There's no equivalent of Google Search Console for ChatGPT's parametric memory. You can't submit a reconsideration request. You can't pay to be included. You can't write a blog post that directly rewrites what a model learned during training.
What this means practically: GEO is a longer game than most founders hope. The right work - building coverage, credibility, structured data, clear writing - will percolate into live retrieval and eventually training data. But "eventually" is measured in months, not days.
One more honest note: LLMs.txt, the format some are promoting as an AI equivalent of robots.txt, scored 2.0/10 in Cyrus's framework and showed no measurable citation lift in ProCloser's field tests. Publishing it doesn't hurt, but prioritizing it over the factors above would be a mistake.
Where directree fits into this picture
One of the signals AI systems use is your presence across structured, third-party sources. Not just your own website.
directree listings are structured and labelled. Every data point carries one of three labels: Observed (verified fact), AI-inferred (explicitly flagged, never disguised as fact), or Founder-edited (claimed by a verified owner). That labelling is machine-readable and honest - exactly the kind of source an AI system can reason about cleanly, without having to guess what's claimed versus confirmed.
A listing also gives you a do-follow backlink and presence in a structured software registry indexed by search engines and AI crawlers alike. It's one signal in the broader entity-building picture, not a shortcut. We won't pretend otherwise.
For the full picture of what founders can do to build AI search visibility, read our GEO guide for founders and the AI search optimization overview.
One concrete step today: a free, honest, do-follow listing on directree. Paste a URL, get a structured listing in about 30 seconds.
Submit a toolWhat not to waste time on
Some things circulating in the GEO space right now are noise:
- LLMs.txt: Free to add, but field research shows no citation lift. Do it in five minutes if you want, but don't treat it as a priority.
- "Prompt engineering" your homepage: Writing your pages to "sound like a chatbot prompt" is overthinking it. Clear, confident, well-structured prose wins - it always has.
- Buying low-quality mentions: Link spam gets filtered out in regular SEO and won't build AI citation signal either. Quality and independence matter.
- Optimizing for only one AI platform: ChatGPT and Perplexity have different architectures. A strategy built around just one platform will miss the others, and more engines are emerging constantly.
Checking one answer by hand is useful, but patterns need a record. An AI visibility tool can retain the answers and citations behind a change, while a free AI visibility check gives you a quick first read.
GEO FAQs
For direct answers on the terminology and practical limits, read our FAQ on increasing AI visibility, ChatGPT visibility, doing GEO, and optimizing content for LLMs.
FAQ
Does ranking on Google automatically mean I'll show up in AI answers?
Not automatically - but it's the strongest external signal researchers have found. Ahrefs found that 38% of Google AI Overview citations come from pages already in the top 10. Solid SEO is the foundation that feeds everything else.
How quickly can I influence Perplexity vs. ChatGPT?
Perplexity retrieves live content on every query, so fresh, indexed content can influence it within days or weeks. ChatGPT's base model depends on training data, which updates on a slower cycle and has a cutoff you can't control. For near-term GEO work, Perplexity responds faster.
Does social media presence help with AI recommendations?
In some categories, yes. Community discussion on Reddit is measurably predictive for queries where people ask for genuine opinions. A Twitter following alone probably doesn't move citation rates much. What matters is whether independent communities are talking about your product in places that generate text used in training data.
Can I force an AI to recommend my product?
No. There's no direct submission path, no paid placement, no optimization trick that guarantees inclusion. The work is building genuine authority - coverage, credibility, clear content, structured data - and letting that build over time. Anyone claiming otherwise is selling something.
What's the one thing a founder should do this week?
Audit your site copy for hedge language. Replace "many teams find [Product] helpful for X" with "[Product] does X." It's free, it takes an afternoon, and it's the highest-scoring fast win in the research. Start there.
