SEO / AEO / GEO7 min read
How to Measure SEO, AI Referrals, Citations, and the Business They Create
A technical audit can tell you whether a page is eligible to be found and reused. It cannot tell you how often people see the business, whether an answer engine cites it, or whether that discovery creates a qualified conversation.
By Dan Grams, Founder & Principal Consultant
My take
Measure modern discovery in layers: implementation readiness, indexed search visibility, observable AI referrals, documented citation samples, on-site behavior, and qualified outcomes. Keep each layer separate so a score or anecdote never becomes evidence for a result it did not measure.
What matters most
- Treat readiness, visibility, referral traffic, citation, and conversion as different measures.
- Preserve known AI referral parameters and ask customers how they found you.
- Sample answer-engine citations with a documented, repeatable method.
- Connect discovery to qualified outcomes without pretending attribution is complete.
Begin with a measurement map, not one AI visibility score
Modern discovery produces several kinds of evidence. A crawler can access a page. A search engine can index it. A page can receive impressions and clicks. An AI answer can mention or cite the business. A person can arrive through a referral, return directly later, book a call, or name an answer engine in conversation. Each event answers a different question.
Collapsing them into one score makes the dashboard simpler and the conclusion weaker. I prefer a layered model because it keeps the evidence honest: readiness describes implementation, visibility describes observed exposure, referral describes a detectable visit, and commercial reporting describes what happened after discovery.
Layer one: implementation readiness
Readiness covers observable conditions that support discovery and comprehension: successful responses, intentional crawl and index directives, canonical consistency, useful titles and headings, meaningful primary content, internal links, accurate structured data, and accessible rendering. These checks are valuable because they identify work the site owner can change.
Readiness is not a ranking, citation, or recommendation. Passing a crawler check does not prove the system indexed the page. Adding schema does not force an answer engine to select it. Optional files and experimental browser-agent interfaces should be reported as such, not promoted into universal ranking factors.
Layer two: search visibility and indexed demand
Use Search Console or the relevant search-platform reporting to observe queries, pages, countries, devices, impressions, clicks, click-through rate, and changes over time. Group queries by the buyer problem they represent rather than reviewing only a list of individual keywords. Separate branded and nonbranded demand.
Annotate material site changes, launches, migrations, seasonality, and measurement breaks. A rising average position can coexist with fewer qualified visits if the query mix changed. A lower click-through rate can coexist with stronger business results if the page earns broader visibility. The commercial interpretation belongs beside the search metrics.
Layer three: observable AI referrals
Some answer systems send identifiable referral traffic. Preserve the referrer and campaign parameters your analytics platform receives, normalize known sources into a documented channel group, and retain the unmodified landing URL for investigation. OpenAI says ChatGPT referrals include the parameter `utm_source=chatgpt.com`, which provides one useful signal when it survives the customer’s path.
Referral reporting will be incomplete. A person may read an answer on another device, copy a URL, search the brand later, use a privacy tool, or arrive through an application that does not pass a useful referrer. Report observable AI referrals as a lower bound, not the full influence of answer engines.
Layer four: citation sampling
A citation is not the same as a visit. It can still matter because it shows whether a source is being selected for the questions the company wants to answer. Build a small prompt set from real buyer questions, document the market, language, account state, model or product, date, and exact wording, then record whether the business is mentioned, cited, accurately described, and linked.
This is sampling, not a universal rank tracker. Answer systems vary by model, retrieval index, geography, personalization, and time. Use the exercise to discover content and entity gaps, not to promise a stable position. Save screenshots or exported answers where policy permits so changes can be reviewed with their original context.
- Question and buyer stage.
- Product, model, date, locale, and account conditions.
- Mention, citation, link, and the cited page.
- Accuracy of the description and any unsupported claims.
- Competitors or sources repeatedly selected instead.
Layer five: behavior and qualified outcomes
On-site behavior helps distinguish a curiosity click from useful discovery. Review landing page, engaged session, the next meaningful page, audit or contact starts, booking intent, completed inquiries, and the eventual qualification outcome. Use privacy-conscious event design and avoid collecting more personal data than the decision requires.
Add a plain-language self-reported source question to forms or discovery calls. Options can include Google or another search engine, ChatGPT or another AI assistant, referral, social, podcast or event, and an open text field. Self-reporting has recall bias, but it can recover influence that browser attribution misses.
Use a scorecard that preserves uncertainty
A useful monthly scorecard can show technical exceptions, nonbrand search visibility, qualified organic landing sessions, observable AI referrals, citation-sample coverage, audit or contact conversion, booked conversations, qualified opportunities, and customer-reported discovery. Definitions and known exclusions should sit next to the measures.
Compare trends and cohorts rather than manufacturing precision. If five customers independently say an AI assistant recommended the company, record that as client-reported discovery—not as proof of market-wide citation share. If a technical audit improves, call it readiness—not visibility. Honest labels make the measurement more useful because the team knows what decision each signal can support.
Turn measurement into a content and operating loop
The scorecard should change the backlog. Queries with impressions but weak satisfaction may need clearer pages. Repeated citation competitors may reveal missing evidence or entity corroboration. Strong AI referrals to an irrelevant page may need better internal paths. Qualified conversations can reveal the language buyers actually use before a purchase.
Review the loop monthly: what became more visible, what created a meaningful visit, what converted, what customers reported, and what uncertainty remains. The purpose is not to prove that every channel caused revenue. It is to make increasingly better decisions with evidence proportional to the claim.