Advertisers have spent a year asking whether showing up in ChatGPT, Gemini or Claude answers actually sells anything. Nobody has produced a number. AI visibility vendors such as GetMint, Profound, Scrunch, Meikai and Semrush return scores, not contributions.
73% of marketers have already bought a tool to track their visibility in AI-generated answers, according to a survey by Scrunch and Scribewise reported by Digiday on Aug. 13. Those tools show whether a brand gets cited, how often, in what tone and against which competitors. They do not show whether that visibility generates incremental sales.
Part of the problem is the path to purchase. Someone can ask ChatGPT for a product recommendation, click nothing, then buy on Amazon, at a retailer or in a store days later. The recommendation may well have shaped the decision, but it appears in no attribution tool. Even when a click does happen, assistant traffic stays small and captures only part of the influence.
Measurement has to answer two separate questions. The first: how advertising, PR, creator work, content and site structure move a brand's presence in those answers. The second: whether that presence lifts awareness, consideration or sales.
GEO tools – generative engine optimization, or optimizing visibility inside generative engines – mostly handle the first. MMM could, in theory, handle the second.
Why MMM looks like the right tool
A quick refresher. Marketing mix modeling is an econometric method that lines up the movement of a business metric – sales, leads, subscriptions – against media spend and the other factors that could explain it: promotions, price, distribution, seasonality, weather or competitive activity.
It runs on aggregate data. It does not need to follow each consumer from exposure to transaction. That makes it well suited to clickless journeys and indirect effects, such as a recommendation in ChatGPT followed by a purchase somewhere else.
So a time series for a brand's visibility in AI assistants could be fed into an MMM. The model would test whether swings in that visibility line up with swings in sales, once the other factors are accounted for.
The variable would not necessarily behave like a media channel. Organic presence in answers has no verified impressions and no spend to calculate a ROAS against. It could instead break down the baseline more precisely, the share of sales the model does not tie directly to observed media spend.
But MMM needs three things: a variable that moves enough over time, a long enough history and a stable measurement method. Data from AI assistants barely meets any of them yet.
What MediaROI and GetMint are trying
Measurement firm MediaROI is working with GetMint on a first method for folding AI assistant visibility into its models. The approach is still in beta and does not claim to produce a GEO ROI today. It starts by building an exposure variable the model can test.
The method has two parts.
The first is an index tracking estimated demand inside AI assistants for the advertiser's category. No reliable public data exists on how many queries ChatGPT, Gemini or Claude handle in a given sector.
MediaROI rebuilds that curve from three inputs, co-founder Antoine Szalewski said: "the size of the search market, estimated with Google Keyword Planner, the seasonality of demand, observed with Google Trends, and public data on the general growth of queries and usage in LLMs."
The goal is not the real query volume but a relative index that moves daily or weekly. For an MMM, a variable's dynamics can be usable even when its absolute level is unknown.
The second part is GetMint's visibility score. MediaROI focuses on prompts close to purchase: where to buy a product, which brand to pick, which supplier to go with. The estimated demand index is multiplied by that score to create the exposure variable.
MediaROI started collecting the data over the summer with several advertisers. The company says it needs about six months of history and expects first results by the end of the year. "The method could also work with scores from platforms other than GetMint if the advertiser already has enough history," Szalewski said.
The approach turns a blind spot into a testable hypothesis. More to the point, it starts building the data series the market lacks today. It also rests on assumptions that still have to hold up.
LLM demand, largely reconstructed
The first limit is the jump from search data to AI assistants. A Google query and a ChatGPT conversation do not always serve the same need. Search engines were built around short queries. Assistants take longer, more contextual questions. Their role shifts by category too: exploration in travel, comparison in automotive, technical help in electronics, recommendation in beauty.
Average LLM growth says nothing about how much people actually use them to buy car parts, pick a telecom operator or compare groceries. Szalewski accepts the limit. His indicator is an estimated trend, not an observed volume.
The second unknown is incremental usage versus substitution. The first version of MediaROI's method assumes some assistant queries add to traditional search. Consultancy Ekimetrics says part of that usage may gradually replace searches once run on Google.
The distinction decides everything. If ChatGPT creates new discovery moments, the LLM variable adds information to the model. If it mostly diverts search queries, the two variables move as two faces of the same behavior. The MMM would then split their contributions badly.
Scores that change with the tool and the prompt
The method then depends on how reliable the visibility score is. Results vary by vendor, prompt panel, models queried, language, test frequency and the way answers get aggregated.
A GEO KPI from one partner is not comparable to another's
"A GEO KPI from one partner is not comparable to another's," said France Hassenforder, partner at Ekimetrics. Before measuring any business contribution, the score going into the MMM has to be defined precisely, with a methodology stable enough to hold.
A brand can score high on a set of generic questions and be absent from the prompts people actually use before buying. Changing the panel can also move the indicator up or down without the brand's real visibility changing at all.
LLM answers carry their own variance. The same question can return different recommendations depending on the moment, the account, the location or the model version. ChatGPT or Gemini updates can break the historical series outright.
"Our clients are keen to test, then cool off because the KPIs are not stable. It is still early," Hassenforder said.
Visibility or brand strength
Causality is the hardest problem. A well-known brand is more likely to be recommended by an AI assistant, because it has more content, more press citations, more social conversation and more authority signals. That same brand strength probably explains part of the sales the MMM observes.
A correlation between LLM visibility and revenue does not prove the recommendations drove the sales. Both can come from a third factor: brand equity.
Branding campaigns, PR, creator work and content production can lift the GEO score and sales at the same time. If those effects are not separated properly, the MMM may credit assistant visibility with a contribution that belongs to the marketing spend that created it.
It can run the other way too. LLMs lean heavily on content already on the web. Their answers can reflect a brand's past popularity rather than fresh exposure capable of changing behavior. The GEO score then works as a brand health indicator, not an independent lever.
Why Ekimetrics and m13h are holding off
Ekimetrics calls the subject a priority and keeps it exploratory. The firm is testing several datasets and weighing econometric methods alongside causal approaches built on control groups. "We do not yet have an indicator robust enough to go into MMMs in a standardized way," Hassenforder said.
The first question is which KPI to use: a citation, share of voice, position in the answer, tone, presence in a purchase recommendation, a price mention. The right one depends on the assistant's role in the journey and the marketing outcome the advertiser wants to explain.
Hadrien de Nijs of measurement consultancy m13h is more cautious still. Across his clients, traffic attributed directly to assistants runs between 0.1% and 0.5%, and often below 0.2%. The number understates their influence, since it misses clickless journeys. It also shows how weak the directly observable signal is next to the other variables in an MMM.
History is the real blocker. "An MMM is usually built on two years of data or more," de Nijs said. Most advertisers hold a few months of GEO scores, sometimes only a handful of data points. "So the variable shows up abruptly at the end of the series, when consumers have been using ChatGPT for years," he said. Putting the LLM piece into MMMs today is "cavalier," in his view. The risk is asking a model for a precise contribution from a variable that is recent, unstable and partly reconstructed.
Build the data now for the measurement later
Neither m13h nor Ekimetrics is ignoring the subject. Their main advice is to start collecting immediately. "It is important to start monitoring now, because you cannot rebuild the history after the fact," Hassenforder said.
Step one is defining a prompt panel that matches real usage and keeping it stable over time. De Nijs splits it into three dimensions: "salience, or the probability the brand gets cited in a category search, favour, how positively or negatively it is presented, and power, its ability to show up when its main competitors are cited."
The measurement conditions then have to be documented precisely: assistant queried, model version, language, country, any persona, test frequency and changes to the panel. Raw answers, citations and sources should be kept so the indicators can be recalculated if the methodology changes.
Advertisers can also dig into the small volume of assistant traffic they already get: conversion rate, basket size, new customers and on-site behavior. Panels and surveys can fill the gap by identifying consumers who used an LLM without clicking a link.
Tests on control markets or control populations will show whether a better GEO score actually comes with a change in behavior.
Advertising inside ChatGPT will supply more data on the paid side: impressions, clicks, conversions and spend. It will not solve organic visibility. Advertisers will have to separate the performance of a campaign running in an assistant from the influence of a recommendation the model generated on its own.
MMM remains one of the more credible candidates for measuring that influence without individual tracking. But being able to add an "LLM visibility" column does not make the contribution reliable. The immediate job is stabilizing the indicators and building history. The business impact comes after.


