Glossary

GEO & AI visibility glossary

30 terms from generative engine optimization and AI search, defined in plain language — with links to the docs, tools and guides where each concept shows up in practice.

GEO (Generative Engine Optimization)

Generative Engine Optimization is the practice of making a business, brand, or website more likely to appear — and be recommended — in answers produced by generative AI systems such as ChatGPT, Gemini, Perplexity, and Google AI Overviews. Where classic SEO optimizes for ranked lists of links, GEO optimizes for synthesized answers: the model reads, summarizes, and recommends rather than merely linking. Practical GEO work includes making content crawlable by AI crawlers, publishing clear factual pages about what the business does and where, adding structured data, earning citations on sources the engines trust, and monitoring how often the brand is actually mentioned. GEO overlaps heavily with AEO (Answer Engine Optimization); many teams use the terms interchangeably. The discipline is young and measurement conventions are still settling, which is why honest tools disclose how they sample prompts and which platforms they query.

Related: What is AI visibility? · Methodology

AEO (Answer Engine Optimization)

Answer Engine Optimization is a close sibling of GEO: optimizing so that answer engines — systems that respond with a single synthesized answer instead of a list of links — describe and recommend your business correctly. The term is popular among enterprise platforms (Profound, for example, brands itself around AEO), while GEO is more common in the small-business and content-marketing world. In practice the work is the same: ensure AI crawlers can read your site, publish unambiguous facts (services, locations, pricing, proof), maintain consistent entity information across the web, and measure whether the engines mention or cite you when buyers ask relevant questions. The main difference from classic SEO is that there is no position #4 in an answer — you are either part of the answer or you are invisible, which raises the stakes of measurement.

Related: Profound alternative · How a scan works

AI visibility

AI visibility is the degree to which a brand or business appears in AI-generated answers when people ask relevant questions — for example, whether a plumber in Austin is named when someone asks an assistant for plumber recommendations in Austin. It is the AI-era analogue of search rankings, but harder to observe: answers are personalized, probabilistic, and change between runs, so visibility must be measured by sampling many prompts and counting mentions, positions, and sentiment. Tools in this category (RankedByAI among them, alongside Peec AI, Otterly.AI, and others) automate that sampling and report metrics such as mention rate and share of voice. Because sampling methods differ between tools, absolute numbers are less meaningful than trends over time measured consistently by the same method — a reason to prefer tools that publish their methodology.

Related: What is AI visibility? · Visibility grade & mention rate

Answer engine

An answer engine is a system that responds to a query with a directly composed answer rather than a list of documents. ChatGPT, Perplexity, Gemini, Copilot, and Google's AI Overviews all behave as answer engines: they gather or recall information, synthesize it, and present a recommendation or explanation, sometimes with citations. For businesses the crucial property is selection: an answer engine typically names only a handful of options, so inclusion is binary and competitive. The term is used to distinguish these systems from search engines (which rank links) and from pure chatbots (which may not retrieve information at all). Optimizing for answer engines — AEO — means becoming one of the few options the engine considers safe and relevant to recommend, and being described accurately when it does.

Related: Which AI platforms are covered · Tool alternatives

LLM (Large Language Model)

A large language model is the neural network at the core of modern AI assistants — GPT-4o and its successors, Gemini, Claude, Llama, DeepSeek, and others. LLMs are trained on vast text corpora to predict language, which gives them broad knowledge with two caveats that matter for visibility work: their built-in knowledge has a cutoff date, and they can generate plausible but wrong statements (hallucinations). Assistants compensate by grounding: retrieving fresh web content at answer time and citing it. For a business this means two distinct surfaces to influence — what the model already believes about you from training data, and what it reads about you live from your site and third-party sources. GEO addresses both, but the live-retrieval surface is the one you can change fastest.

Related: How a scan works · Methodology

Prompt

A prompt is the text a user sends to an AI assistant — in visibility work, specifically the buyer-style questions whose answers you care about, such as “best CRM for a small law firm” or “affordable wedding photographer near Lisbon”. Visibility tools track a set of prompts per brand: they submit each prompt to AI platforms on a schedule, record whether the brand is mentioned, in what position, with what sentiment, and which competitors appear. The choice of prompts is the biggest methodological lever in the category: hand-picked prompts can flatter a brand, while automatically generated buyer questions aim to represent what real customers ask. Because answers are probabilistic, a single prompt run proves little; meaningful measurement samples many prompts repeatedly and reports rates rather than one-off results.

Related: How prompts work · Add, edit & disable prompts

Prompt sampling

Prompt sampling is the measurement technique behind most AI visibility metrics: instead of asking one question once, a tool submits many relevant prompts (and often repeats them over time) and aggregates the results into rates — mention rate, average position, share of voice. Sampling is necessary because AI answers are non-deterministic: the same question can yield different recommendations across runs, users, and phrasings. Sound sampling requires a representative prompt set (real buyer questions, not cherry-picked ones), sufficient volume, consistent scheduling, and honest disclosure of sample size — small samples deserve explicit small-sample caveats. When comparing tools or reading any visibility score, the first question to ask is how the prompts were chosen and how many samples the number rests on; without that context, a percentage is marketing, not measurement.

Related: Methodology · How a scan works

Buyer question

A buyer question is a prompt phrased the way a real potential customer would ask it — “who's the best emergency dentist open on weekends in Leeds?” rather than “dentist Leeds”. Buyer questions matter in AI visibility because assistants answer conversational queries, and because purchase-intent questions are where being recommended translates into revenue. Tools differ in how they source them: some ask the user to type in prompts by hand (which risks tracking flattering queries), while others generate buyer questions automatically from the business's category, services, and location, aiming for a representative sample of real demand. A good buyer-question set covers the full journey — discovery (“options for X”), comparison (“X vs Y”), and validation (“is X trustworthy?”) — so the resulting visibility metrics reflect how the business actually appears throughout a buying decision.

Related: How prompts work · Prompt suggestions & export

Share of voice

Share of voice (SOV) in AI visibility is the fraction of sampled AI answers, within a defined prompt set, in which a given brand appears relative to all brand appearances — a competitive measure of how much of the conversation a brand owns. If assistants are asked one hundred buyer questions in your category and your brand appears in answers thirty times while all tracked brands together appear one hundred fifty times, your SOV is 20%. SOV is most useful for tracking relative movement: whether you are gaining or losing ground against named competitors on the same prompts over time. Like all sampled metrics it depends heavily on the prompt set and sample size, so compare SOV numbers only within one consistent methodology, never across different tools.

Related: Share of voice, sentiment & trends · Competitor analysis

Mention rate

Mention rate is the percentage of sampled AI answers that mention a given brand at all — the most basic AI visibility metric. If a tool asks assistants fifty buyer questions relevant to your business and your brand is named in ten of the answers, your mention rate is 20%. Mention rate is usually complemented by position (are you the first recommendation or an afterthought?), sentiment (how are you described?), and share of voice (how do you compare to competitors on the same prompts?). Because answers vary between runs, mention rate is a statistical estimate: it needs adequate sample sizes and consistent methodology to be meaningful, and honest reports disclose both. Tracked over time on a stable prompt set, it is the cleanest single signal of whether your GEO work is paying off.

Related: Visibility grade & mention rate · AI visibility benchmark

Citation

In AI search, a citation is a source link that an assistant attaches to part of its answer — the pages Perplexity lists as references, or the links under a Google AI Overview. Citations matter twice over. For users they are the trust trail; for businesses they are the mechanism by which content earns presence in answers: if your page is cited, you are part of the answer even when your brand name is not spoken. Being citable requires crawlable, factual, well-structured pages that directly answer common questions. Analyzing which domains an engine cites for your category reveals where to earn coverage — directories, review sites, local news, industry publications. A page that never gets cited despite ranking well in classic search often signals crawler blocking or content that models cannot easily extract facts from.

Related: Citation gap · How a scan works

Citation gap

A citation gap is the set of sources that AI assistants cite when answering buyer questions in your category — but that don't mention your business. If assistants answering “best boutique hotels in Porto” repeatedly cite two travel guides, a review platform, and a local blog, and your hotel appears on none of them, those four sources are your citation gap. It is one of the most actionable findings in AI visibility work because it converts an abstract score into a concrete outreach list: get listed, reviewed, or covered on the specific sources the engines already trust for your niche. Closing a citation gap tends to move visibility faster than on-site changes alone, since assistants lean on those third-party sources to decide which options are safe to recommend.

Related: Citation gap · Website diagnosis

Grounding

Grounding is the process by which an AI assistant bases its answer on retrieved, verifiable content — typically live web results — rather than relying only on what it memorized during training. When Perplexity searches the web and cites pages, or Gemini pulls in Google results before answering, the answer is grounded. Grounding matters for businesses because it is the fast path to visibility: model retraining takes months and is out of your control, but grounded answers reflect what is on the web right now. If your site is crawlable, your facts are current, and trusted third-party sources describe you accurately, grounded answers can start recommending you within days of a fix. It is also why blocking AI crawlers is usually self-defeating: an assistant cannot ground an answer in content it cannot read.

Related: Which AI platforms are covered · Methodology

RAG (Retrieval-Augmented Generation)

Retrieval-Augmented Generation is the architecture behind grounded AI answers: before generating a response, the system retrieves relevant documents (from a search index, a database, or the live web) and feeds them into the model's context so the answer can be based on current, specific information rather than training memory alone. RAG is what allows an assistant to know about a restaurant that opened last month or a price that changed yesterday. For visibility work, RAG defines the playing field: the retrieval step behaves much like a search engine (so retrievability and relevance matter), and the generation step decides what gets said (so clarity and extractability of your content matter). Content that is easy to retrieve and easy to quote — direct answers, clean structure, unambiguous facts — wins in RAG pipelines.

Related: How a scan works · Get recommended by ChatGPT

Hallucination

A hallucination is a confident but false statement produced by a language model — a wrong opening hour, a service you don't offer, an award you never won, or a business that doesn't exist. Hallucinations occur because LLMs generate plausible text rather than look up verified facts, and they matter commercially: an assistant that hallucinates your prices or misstates your service area is actively misleading your customers. The main defenses are on your side of the equation: publish clear, current, unambiguous facts on pages AI crawlers can read, keep third-party listings consistent, and monitor what assistants actually say about you so you catch errors early. Grounded answers hallucinate less than pure-memory answers, which is another reason to make your site easy for retrieval systems to use.

Related: Data privacy & retention · Methodology

Knowledge cutoff

The knowledge cutoff is the date after which a language model's training data ends — anything that happened later is invisible to the model's built-in memory. If a model's cutoff predates your rebrand, your new location, or your business entirely, its unassisted answers will be outdated or omit you. Cutoffs are why grounding matters so much for businesses: assistants that retrieve live web content can know about you regardless of when the model was trained, while memory-only answers cannot. Practically, the cutoff splits GEO into two horizons: influencing training data (slow, indirect — being widely and consistently described on the web so future model versions learn about you) and influencing retrieval (fast — making current facts crawlable and citable today). Monitoring both surfaces tells you which problem you actually have.

Related: How a scan works · Set up weekly monitoring

Training data

Training data is the corpus of text a language model learns from — web pages, books, code, forums, and licensed sources collected up to the model's knowledge cutoff. What the web said about your business during collection shapes what the model believes about you by default: a business described consistently across many reputable pages tends to be recalled accurately; one with sparse or contradictory coverage may be forgotten or garbled. You cannot edit training data retroactively, but you can influence future rounds: sustained, factual, consistent coverage on crawlable pages is what tomorrow's models will learn from. In the meantime, retrieval (RAG) lets current assistants read today's web, which is why GEO work targets both horizons — long-term reputation in the corpus and short-term accuracy at retrieval time.

Related: Get recommended by ChatGPT · AI Crawler Checker

AI crawler

AI crawlers are the bots that AI companies use to fetch web content — OpenAI's GPTBot and OAI-SearchBot, Anthropic's ClaudeBot, Google-Extended, PerplexityBot, Meta's crawlers, and others. They serve two distinct purposes: collecting training data for future models, and fetching pages live to ground current answers. Whether these bots can read your site directly affects your AI visibility: a site that blocks them (deliberately or through misconfigured robots.txt, firewalls, or bot protection) cannot be cited in grounded answers and contributes nothing to future training data. Auditing crawler access — which bots are allowed, which are blocked, and whether that matches your intent — is one of the first checks in any AI visibility diagnosis, and one of the cheapest fixes when it is wrong.

Related: AI Crawler Checker · Website diagnosis

robots.txt

robots.txt is the plain-text file at the root of a website that tells crawlers which paths they may fetch. It has become the primary control surface for AI access: rules targeting user agents like GPTBot, ClaudeBot, Google-Extended, or PerplexityBot decide whether those systems can read your content for training or live answers. Two failure modes are common. Some sites block AI crawlers unintentionally — a blanket disallow inherited from a template, or bot protection layered on top — and then wonder why assistants never cite them. Others assume robots.txt guarantees compliance; it is a convention, not an enforcement mechanism, though the major AI crawlers document that they honor it. Reviewing your robots.txt against your actual AI visibility goals is a five-minute check with outsized impact either way.

Related: AI Crawler Checker · Website diagnosis

llms.txt

llms.txt is a proposed convention — a Markdown file at /llms.txt — that gives large language models a curated, machine-friendly summary of a website: what the business is, what its key pages contain, and where to find the most important information. The idea is analogous to robots.txt (a standard location that bots know to check) but oriented toward comprehension rather than access control: instead of parsing your whole site, a model can read one concise document that you control. Adoption by AI providers is still uneven and evolving, so an llms.txt file is best treated as a low-cost complement to — never a substitute for — crawlable pages, structured data, and clear on-page facts. It is cheap to generate, easy to keep current, and harmless where it is ignored.

Related: llms.txt Generator · Website diagnosis

Structured data

Structured data is machine-readable markup embedded in web pages — most commonly JSON-LD using schema.org vocabulary — that states facts explicitly: this is a LocalBusiness named X, at address Y, open these hours, offering these services, with this rating. Search engines have used it for years to power rich results; for AI systems it reduces ambiguity in exactly the places ambiguity hurts, helping retrieval pipelines and answer engines extract correct names, prices, locations, and relationships instead of inferring them from prose. Well-structured pages are easier to cite accurately and less likely to be garbled in a synthesized answer. The usual priorities for a business site are Organization or LocalBusiness markup, Product or Service where relevant, FAQPage for question content, and consistency between the markup and the visible text.

Related: Website diagnosis · Fix prescriptions & verification

JSON-LD

JSON-LD (JavaScript Object Notation for Linked Data) is the recommended format for embedding structured data in web pages: a script block containing a JSON object that describes the page's entities using schema.org types. Unlike older formats that interleave markup with visible HTML (microdata, RDFa), JSON-LD lives in one self-contained block, which makes it easier to generate, validate, and keep in sync with page content. For AI visibility, JSON-LD is the practical vehicle for most structured-data work: LocalBusiness details, product and pricing facts, FAQs, article metadata, and organizational identity all travel in it. Search engines parse it directly, and clean JSON-LD correlates with pages that answer engines can quote accurately. Validation tools catch syntax errors; the harder discipline is keeping the claims in the markup truthful and current.

Related: Website diagnosis · llms.txt Generator

Entity

An entity is a distinct thing the machine can identify — a specific business, person, product, or place — as opposed to a mere string of characters. Search and AI systems increasingly reason about entities: “Rosa's Bakery on Main Street” should resolve to one consistent business with one address, one set of hours, and one reputation, no matter how the name is phrased. Entity clarity is foundational for AI visibility because assistants recommend entities, not keywords. It is built through consistency: the same name, address, and key facts on your site, your Google Business Profile, directories, and social profiles; structured data that declares the entity explicitly; and third-party coverage that reinforces it. Inconsistent entity signals — old addresses, name variants, duplicate listings — make models hedge, confuse you with others, or omit you.

Related: Website diagnosis · What is AI visibility?

E-E-A-T

E-E-A-T stands for Experience, Expertise, Authoritativeness, and Trustworthiness — the framework from Google's Search Quality Rater Guidelines for judging content quality. It is not a direct ranking factor but a description of what quality systems try to reward, and the concept transfers naturally to AI answers: assistants tend to recommend businesses that look credible across the open web — real authorship, verifiable credentials, consistent facts, genuine reviews, coverage on independent trusted sources. For GEO purposes, E-E-A-T translates into practical work: publish content demonstrating first-hand expertise, make the people behind the business visible, earn reviews and mentions you don't control, and keep claims verifiable. Content that fakes authority tends to fail at exactly the moment that matters — when an engine cross-checks it against the rest of the web.

Related: Get recommended by ChatGPT · Fix prescriptions & verification

Sentiment analysis

Sentiment analysis in AI visibility measures how a brand is described in AI answers, not just whether it appears. Being mentioned as “a popular choice with consistently strong reviews” and being mentioned as “an option some customers report billing issues with” are both mentions — with opposite commercial value. Visibility tools classify the language around each brand mention (positive, neutral, negative, often with finer categories) and aggregate it across sampled prompts, turning tone into a trackable metric alongside mention rate and share of voice. Sentiment findings are actionable in a specific way: assistants usually echo the sentiment of their sources, so negative AI sentiment typically traces back to identifiable review platforms, forum threads, or articles — which tells you exactly where reputation work is needed.

Related: Share of voice, sentiment & trends · How a scan works

AI Overviews

AI Overviews are the AI-generated summaries Google shows above classic results for many queries, composed by Gemini models from retrieved web content, with source links. Together with the fuller AI Mode, they represent the largest-scale deployment of AI answers, because they reach mainstream Google users who never chose an AI product. Their practical impact is twofold. Inclusion becomes binary: an Overview names a few options, and businesses outside it lose attention even when they rank organically below. And clicks decline: users who get their answer in the Overview often do not visit any site (see zero-click search). Appearing in AI Overviews follows the grounded-answer playbook — crawlable content, clear facts, strong entity signals, and presence on the sources Google's systems already trust for the topic.

Related: Which AI platforms are covered · Semrush AI Toolkit alternative

Visibility grade

A visibility grade is a summary score — RankedByAI uses letter grades from A to F — that condenses a business's AI visibility measurements into one readable verdict. Behind the letter sit sampled metrics: how often assistants mention the business on relevant buyer questions, its position among recommendations, sentiment, and site-readiness factors such as crawler access and structured data. A grade is a communication device, not a precision instrument: its honest purpose is to make relative status and progress legible (a D that becomes a B after fixes tells a clear story), not to claim decimal accuracy. Any trustworthy grade discloses its inputs, sample sizes, and methodology, and flags small samples explicitly. Compare grades only within one tool's consistent method — a “B” from two different tools means two different things.

Related: Visibility grade & mention rate · Methodology

Monitoring cadence

Monitoring cadence is how frequently AI visibility measurements are repeated — daily, weekly, or monthly re-runs of the same prompt set across the same platforms. Cadence matters because AI answers drift: models are updated, retrieval indexes refresh, competitors publish content, and reviews accumulate, so any single measurement is only a snapshot. Weekly cadence suits most small businesses: frequent enough to catch meaningful shifts and verify whether fixes worked (the diagnose → fix → re-scan loop), without drowning in the noise of run-to-run randomness that daily sampling amplifies. Consistency beats frequency — the same prompts, platforms, and method each cycle is what makes trend lines meaningful. Cadence is also a pricing axis in this category: many tools reserve daily updates for higher tiers, so match the cadence you pay for to how fast you actually act on findings.

Related: Set up weekly monitoring · Email notifications

Comparing tools? See the alternatives pages →

See these concepts on your own data — free scan

Free scan