The IAB finally handed the industry the right scorecard for AI search. Its 2026 framework names the four things that actually matter, Presence, Prominence, Portrayal, and Persuasion, and in nearly the same breath, its review of the roughly twenty tools measuring them found that no two agree on a score. Only about 16% of brands track AI visibility at all. So the useful move is not to explain the 4 P’s one more time. It is to take the tools you are actually being sold and score them against the framework.
Do that and a pattern shows up fast. Most tools nail Presence, fudge Prominence, barely touch Portrayal, and not one of them can honestly prove Persuasion. Being cited is not being chosen. A dashboard full of mentions is a vanity metric until it moves a number a CFO recognizes.
16% is the number that should scare you
Start with the adoption gap, because it frames everything else. The overwhelming majority of brands are not measuring how they appear in AI answers at all, and the minority that are cannot compare notes, because no two tools score it the same way. Feed the same brand into four platforms and you will get four different visibility numbers.
For a while, the excuse was that there was no shared framework. That excuse is gone, the IAB shipped one. The problem now is not a missing scorecard. It is that the tools built to fill it in are early, inconsistent, and mostly selling the easiest metric as if it were the whole picture.
The 4 P’s, fast
Presence, how often your brand shows up or gets cited across AI queries. This is the floor. And there is a real floor beneath the floor: the IAB calls anything under 50 queries “exploratory,” which is a polite word for not decision-grade. Below that volume you are reading tea leaves.
Prominence, placement and treatment. Are you the highlighted, recommended answer, or one entry buried in a list of alternatives? Being mentioned in position seven of a comparison is not the same as being the pick, and Prominence is where that distinction lives.
Portrayal, how you are described: sentiment and accuracy. Split this in two. Hallucinations, where the model invents something false, are the loud problem and, mercifully, a declining one. The quiet, dangerous problem is factual error from stale training data, the model confidently describing your old pricing, discontinued product, or a policy you changed a year ago.
Persuasion, did the citation actually move anything? Post-citation behavior: traffic, consideration, conversion. Not the citation itself, what it caused. This is the P that connects to money, and the one everything downstream depends on.
Presence is the easy P, and the one everyone oversells
Every tool in the category tracks mentions. It is table stakes, and it is exactly where dashboards manufacture the illusion of coverage. A rising Presence line looks like progress, feels like progress, and demos beautifully. It is also the metric most disconnected from outcomes.
The trap is high Presence with zero downstream signal: you are appearing more, and nothing is happening. More citations, same pipeline. If a tool leads with Presence and gets quieter as you ask about the other three P’s, it is telling you where its real capabilities end.
Prominence is where the tools quietly diverge
Ask why two tools give your brand different scores and Prominence is usually the culprit. Measuring whether you are mentioned is relatively easy; measuring whether you are featured, highlighted, recommended, placed first, is a judgment call, and every tool makes it differently.
They sample different engines, run different prompt sets, and define “prominent” in incompatible ways. One counts any mention in the answer body; another only credits a top recommendation; a third weights by position. None of them is necessarily wrong, but they are not measuring the same thing, which is why their numbers cannot be reconciled. When a vendor shows you a single Prominence score, ask what it actually counted.
Portrayal is the one that should keep you up at night
This is the P that matters most in high-consideration categories and the one almost nothing measures well. When someone is shopping like they’re choosing a surgeon, how the model describes you decides whether you make the shortlist, and the model will describe you from whatever it trusts most, which is frequently not your own site.
The failure mode is not usually a wild hallucination. It is a stale fact stated with total confidence: last year’s price, a feature you sunset, a policy you reversed. It is the same dynamic as the machine gaming its own scorecard, a system optimizing for a plausible answer, not a correct one. Most tools will happily report that your Presence is up while the thing being presented about you is wrong, and they will not flag the error, because checking sentiment plus factual accuracy across engines is genuinely hard and most of them do not really do it.
Persuasion is the P no tool can honestly close
Here is the uncomfortable center of the whole category: none of these tools can prove a citation drove revenue. They can show you appeared. They can sometimes show a click followed. They cannot show the appearance caused the outcome, because they do not run the experiment that would prove it.
This is the AI-visibility version of platform ROAS grading its own homework. A dashboard that draws a line from “cited in ChatGPT” to “traffic went up” is asserting attribution, not demonstrating causation. The honest distinction is directional versus decision-grade: appearing-more-and-traffic-rose is directional, worth knowing, not worth reallocating budget on. Decision-grade means you held something out and measured the lift. Know which one you are buying before you move a dollar.
The scorecard: four tools against four P’s
Here is the category as it actually stands. The pattern to watch is the columns emptying out as you move right.
| Tool | Presence | Prominence | Portrayal | Persuasion | Best for |
|---|---|---|---|---|---|
| Semrush (AI Toolkit / Enterprise AIO) | Strong | Partial | Limited | No | Teams already in Semrush; AI Overviews tie-in |
| Ahrefs Brand Radar | Strong | Partial | Thin | No | SEO teams wanting mentions inside the Ahrefs stack |
| Profound | Strong | Good | Moderate | No | Enterprise multi-engine coverage and an action layer |
| Peec AI | Strong | Good | Sentiment + source attribution | No | Prompt-level tracking, sentiment, EU/GDPR teams |
Honorable mention: Otterly, the cheapest entry, monitoring only. It tells you what is happening. It does not help you fix it.
The split is roughly this. The SEO incumbents, Semrush, Ahrefs, bolted AI visibility onto an SEO spine, so they are strong on Presence and weaker as you move right. The AI-native tools, Profound, Peec, went deeper on Prominence and Portrayal, and Peec in particular does real work on sentiment and source attribution. But every column on the far right says the same thing. Everyone sells Presence. Nobody sells Persuasion, because nobody can. (This space ships features monthly, confirm each tool’s current Portrayal and Persuasion capabilities against its live product page before you rely on this.)
What to demand before you buy
Treat the vendor call like a measurement audit, not a demo. Ask for multi-engine coverage, a tool sampling only one assistant is measuring a slice of the market. Ask for a real prompt-set methodology, not twelve hand-picked queries; hold them to the IAB’s own floor of at least 50, and prefer far more. Ask whether they flag sentiment and factual accuracy, and whether they distinguish a hallucination from a stale-data error. And ask the closing question plainly: how do you connect this to conversions? If the answer is “attribution,” ask them to show you the holdout. If there is no holdout, it is not attribution, it is a correlation with a confident font.
The honest close
The 4 P’s are right. The tools are early. That is not a reason to skip measurement, with only 16% of brands tracking any of this, the gap is the opportunity, it is a reason to measure all four and to know which ones your tool is actually delivering. This is the same one-machine logic that runs through SXO, SEO, AEO, and GEO and through why AI search is forcing paid back upstream: appearing is not the same as being chosen, and being chosen is not the same as being paid for. Measure all four P’s, or you will end up optimizing the one that is easiest to game.
So take an honest look at your current dashboard: which of the 4 P’s is it actually measuring, and which is it just calling a metric?