AI visibility is easy to make vague. A dashboard can show a number, a green arrow, and a competitor chart without telling you what was actually asked or what an AI engine actually said.
That is not useful enough for a founder deciding where to spend time. At directree, we measure a narrower thing: whether an AI answer mentions a product when a buyer asks a relevant, brand-free category question. We keep the answer behind the result, so a score is never meant to stand on its own.
This is how GEO Monitor currently works. It is deliberately a measurement product. It cannot place a company in ChatGPT, change a directree listing, or guarantee a recommendation.
Start with a prompt and engine cell
A scan begins with active prompts for one tracked product. A prompt is a buyer question such as “What are the best project management tools for a remote team?” The important rule is that it should not name the company being measured. If the prompt says “Is Acme good?”, the brand has already been injected into the answer. That can be useful research, but it is not an honest discovery test.
GEO Monitor combines each active prompt with every engine the customer's plan permits. That pair is a measurement cell. The current engine set can include ChatGPT, Perplexity, Gemini, and Google AI Overviews.
Starter runs a search-mode cell for each allowed prompt and engine. Pro also runs two cells for supported engines: search mode and memory mode. Google AI Overviews currently runs as search mode only.
That distinction matters. Search-connected answers can depend on what the engine retrieves at the moment of the request. Memory-mode answers ask without web search. They are not interchangeable, and we do not combine them silently.
We store the raw answer first
For every completed cell, GEO Monitor stores the raw response text, the prompt, engine, mode, scan, and provider cost. It also stores the citations the provider returns when they are available.
The raw response is the receipt. It lets a user answer basic but crucial questions:
- Was the product actually recommended, or merely mentioned in passing?
- Which competitors were named first?
- Was the answer describing the product accurately?
- Did the engine cite a page that contains useful evidence, or an irrelevant page?
This is why we are cautious about treating an aggregate score as the product. A score is a summary. The answer is the evidence.
If you want a small real-world baseline before setting up monitoring, run the free AI visibility check. It uses the same core idea: category buyer questions, not a prompt that hands the model the brand name.
How mention detection works
We maintain brand names, the monitored domain, and any configured aliases for a tracked product. The scanner checks the raw answer for those terms case-insensitively. It records whether there was a match, how many matches occurred, the first paragraph where a match appears, and a short excerpt around the first match.
The scanner also looks for configured competitors and a small set of known brands. Those detections are stored with the answer, ordered by their first paragraph position. That gives a user a concrete competitor view without pretending the result is a market-wide ranking.
There are limits. Name matching can miss an unusual spelling or find an ambiguous one. An answer can recommend a product without using the exact tracked alias. A mention can also be negative or irrelevant. That is another reason the receipt matters: users can open the text instead of relying on a parser's interpretation.
The current visibility score is simple
The dashboard's score is calculated from the stored result cells. A result gets:
- 1.0 when the tracked brand is found in the first paragraph
- 0.7 when the brand is found later in the answer
- 0 when the brand is not found
We add those earned values and divide by the number of result cells, then show the result as a percentage.
For example, across four cells, a brand mentioned in the first paragraph twice and later once earns 2.7 points. Dividing by four gives 67.5%, which the dashboard rounds to 68%.
This weighting is intentionally modest. It recognizes that an early mention is usually more visible than a later one, but it does not claim that the difference is scientifically universal or that a later mention has no value. It is a practical summary of the evidence on the screen.
It is also not a share-of-voice calculation, a probability of recommendation, or a statement about all buyer prompts. It only describes the configured prompt and engine cells in the selected scan history.
Citations are research leads, not a causal claim
When a provider returns citations, GEO Monitor saves them with the result and keeps a citation rollup by domain and engine. For each cited page, the system may check whether the page contains one of the tracked brand terms. It records that presence check and updates the citation count over time.
That can help reveal patterns. Perhaps a competitor appears alongside a comparison page repeatedly cited by one engine. Perhaps a product is named but none of its own pages are among the cited sources. Those are useful leads for research.
They are not proof that adding a link, changing a page, or getting a citation will cause the next answer to change. AI-answer systems retrieve and synthesize information in ways that are not fully transparent. We do not turn a correlation into a ranking promise.
For the broader site and content foundations, read our AI search engine optimization guide. For a focused buyer-tool view, see the AI search optimization tool page.
Accuracy Watch is separate evidence
On GEO Pro, Accuracy Watch compares brand-containing responses against product facts a customer has entered. A compact model judgement labels each fact as consistent, contradicts, or not mentioned and stores a short evidence note.
This is intentionally additive. If that step fails, it does not fail the scan. It is also not presented as an objective truth engine. It is a fast way to flag an answer that deserves a human read, especially when a product description, price, or capability is stated incorrectly.
The factual source remains the customer's maintained facts and the raw AI answer. A model-generated check is a review aid, not a final verdict.
What our method does not measure
There are several things this method does not claim to measure:
- Every prompt a real buyer might ask
- A universal rank inside an AI engine
- Personalised answers from a user's own account or location
- Sentiment, unless a human reads the answer in context
- Causation between a site change and a later AI mention
- Whether a directree listing changes an answer
GEO Monitor tracks AI engines. It has zero effect on your directree listing or ranking. That separation is important to us because a directory should not sell a hidden path into a recommendation.
A sensible way to use the data
Start with 10 to 15 questions that reflect a real buying decision. Keep them stable for a few scans. Read the answers, not just the chart. Record the competitors and sources that keep appearing. Then improve genuine gaps: unclear product pages, inconsistent factual claims, missing documentation, or weak third-party corroboration.
After that, scan the same cells again and compare the receipts. If the answers change, you have something to investigate. If they do not, you still have an honest baseline rather than an invented success story.
Where directree fits
A directree listing is a structured, public profile with fields labelled as Observed, AI-inferred, or Founder-edited. It is not a paid way to influence an AI answer. If your product needs a clear, accurate public footprint, submit it to directree. Then use an evidence-first monitor to see what AI systems actually do with the wider web.
GEO FAQs
For plain-language answers about measurement and answer engines, see our AI visibility FAQ, how to measure answer engine optimization, AEO vs SEO, and Google search vs LLM search.
FAQ
Why do you weight a first-paragraph mention at 1.0 and a later one at 0.7?
The current dashboard treats an early mention as more visible while still assigning meaningful value to a later mention. It is a practical summary, not a universal ranking formula.
Does GEO Monitor save the actual AI answers?
Yes. Each result stores the raw answer text, the relevant prompt, engine, mode, detected mentions, and citations returned by the provider.
What is the difference between search and memory mode?
Search mode permits web search for the answer. Memory mode does not. Pro runs both for supported engines because they can produce materially different evidence.
Can the score prove that my AI visibility improved?
No. It can flag a change across the same configured cells. Read the stored answers and citations before concluding why the change happened.
