How to make Performance Max focus on net new customers

How to make Performance Max focus on net new customers

There’s a trap door waiting for DTC brands that invest in Google Ads that makes your dashboards look amazing, but absolutely wrecks your P&L.

It’s the danger of recycling traffic from Meta.

Thanks to the overlap between paid search and paid social traffic, running Google as a standalone channel is incredibly difficult if you don’t know how to set it up. Ad platforms refuse to share data with one another, and they love to claim credit for the same conversion — even if those sales would’ve happened without the influence of ads.

The DTC brands I speak to are often proud to show off their new customer numbers: month-over-month growth, a steady upward trend, and a fantastic dashboard. But when we go deeper into the data, we often find that a big chunk of those “new” customers are:

  • Conversions that would’ve happened because of brand or content efforts.
  • Customers who aren’t truly incremental because they consumed ads on multiple platforms.
  • The same people signing up with multiple email addresses.

You could argue that these overlapping sales still count as revenue, and they do. But when you look at the contribution margin from those sales, they cost far more than they should and erode actual profit.

In other words, you lose money when you run ads on both platforms without guardrails.

But that doesn’t mean you need to stop or limit yourself to one channel. Instead, you need a better system for measuring actual customer acquisition.

Why the new exclusions matter

If you’re spending five figures or more on Meta, TikTok, AppLovin, or any other top-of-funnel channel, you’ll want to minimize overlap with other channels to drive actual new customer acquisition.

Here’s what that looks like:

  • Someone sees your ad on Facebook or Instagram.
  • They visit your site, browse, and leave without buying.
  • A while later, they search for your brand on Google or get retargeted on YouTube.
  • Performance Max swoops in, grabs the conversion, and reports strong ROAS.
  • You may have won that order anyway, but now Google and Meta both want credit for it.

Now you’re paying two or more platforms to recycle a conversion that you might have earned with just one.

Ever since Performance Max launched, there wasn’t much you could do about this. It’s been a bit of a black box that automatically goes after the warmest traffic it can find: branded search, site visits, email subscriptions, and existing customers.

It lets you bid more for new customers, but you can’t really stop the campaign from defaulting to easy mode.

A while ago, Google began letting you exclude people searching for your brand on Search and Shopping. Performance Max still targeted warm audiences through YouTube, Gmail, and the Display Network.

The latest round of updates from Google has finally addressed this problem. You can now force Performance Max to focus on net new customer acquisition through a combination of brand exclusions, audience exclusions, and Customer Match data. 

First-party audience exclusions, announced in March, are the final piece that makes this possible (though not foolproof – customer list matching is never perfect).

See exactly how your competitors win.

Uncover the keywords, ads, landing pages, and strategies driving your competitors’ paid search success—and find your next opportunity to outperform them.

Analyze your competitors

A four-step framework for net new customer acquisition

Here’s a four-step framework we’re using at my agency to help clients maximize incrementality.

Step 1: Exclude your brand

This one has been around for a while, but it’s the foundation, so we have to start here.

For smaller brands, brand exclusions usually aren’t necessary. But once you’re spending real money and seeing more than 15% to 20% of your cost or revenue coming from brand searches, it’s time to take action. 

There are two parts to this.

Go into your campaign settings and add a brand exclusion. If your brand isn’t already on the list, click New brand list, create one, and add your brand. Google will do its best to block branded queries from this list.

Because brand exclusions aren’t foolproof, go to the Keywords tab inside the campaign and add your brand name as a phrase match negative keyword. Add a few common variations, too. This catches anything the brand list misses.

If you’re excluding brand terms from Performance Max, you need a dedicated brand Search campaign and a brand Shopping campaign to capture those searches. Otherwise, you’re just leaving money on the table for competitors.

Step 2: Exclude website visitors and email subscribers

Even if you blocked brand searches, Performance Max would still retarget people who visited your website, opened your emails, or interacted with your brand on YouTube, Gmail, Discover, and Display. So even with brand exclusions in place, a big chunk of your spend still went to warm traffic.

Now you can change that. Go to your campaign settings and find the new audience exclusions option. Then build a few remarketing lists:

  • All website visitors: Set this up through the Google Ads pixel or Google Analytics. It captures anyone who has visited your site.
  • Email subscribers: Connect Klaviyo (or whatever ESP you’re using) directly to Google Ads. The benefit of the Klaviyo integration is that the audience updates in real time, so new subscribers are added automatically.

Once you exclude these audiences, Performance Max can only go after people who haven’t interacted with your brand in any meaningful way. What we typically do, and what I recommend, is to come up with an engagement metric that fits each account’s business goal, such as cart adds rather than visitors from the past seven days.

What a change from how this campaign type used to work.

Get the newsletter search marketers rely on.


Step 3: Exclude existing purchasers

Same idea as Step 2, but specifically for people who have already bought from you. You can do this two ways.

  • Through a pixel-based audience that captures anyone who has triggered the purchase event. 
  • By uploading your customer list directly. Shopify now lets you set up Customer Match lists right inside the Google Shopping app, and Klaviyo can do this, too.

Add these audiences to the exclusions section of your campaign, and you’re done.

A small caveat to keep in mind: audience matching is never 100%. If you upload a customer list of 1,000 people, Google might only match 900 of them. So you’ll still see some level of bleed. But going from “the campaign is targeting all my existing customers” to “the campaign is targeting maybe 10% of them” is still a huge win.

Step 4: Use ‘New Customer Bidding’ in campaign settings

The last piece is to tell the campaign explicitly that you want new customers.

In your campaign settings under customer acquisition, you’ll see two options: bid only for new customers, or bid higher for new customers. Both require you to connect a customer list (which you’ve probably already done by Step 3).

The “only new customers” option is the most aggressive setting. The campaign simply won’t bid on existing customers. Combined with the audience exclusions from Steps 2 and 3, this gets you as close to pure new customer acquisition as Performance Max will allow.

The “bid higher for new customers” option is more flexible. You set a dollar value that represents the additional value of a new customer, and the system bids more aggressively when it thinks an auction will result in one.

Here’s where you need to be careful. If you tell Google a new customer is worth an extra $100, and you get a $200 sale from a new customer, Google will report it as $300 in revenue. That extra $100 is a fictional reporting value, not real revenue. It will inflate your ROAS numbers and distort your target ROAS bidding.

Our recommendation is to use a small placeholder value, such as a penny or a dollar, when you want to nudge the system toward new customers without distorting your reporting. Or use a number that genuinely reflects the lifetime value premium of a new customer to your business.

What to expect from this approach

It’s still early, so we can’t draw firm conclusions yet. But based on my experience managing PPC for ecommerce brands, here’s what I expect to happen.

Many advertisers who walked away from Performance Max did so because it was simply recycling Meta traffic. By splitting it out, you force it to go after net new traffic.

This will likely benefit brands that don’t have a ton of video creative for YouTube, which is another platform where brands try to drive net new acquisition at the awareness stage.

One of the big differences between Performance Max and Demand Gen is that the former is much more conversion-focused. Any brand considering excluding branded Search and Shopping from Performance Max should also consider this tactic, as it tends to over-index on hot traffic.

In terms of outcomes, I expect the reported ROAS attributed to Performance Max to be lower than what you may have seen in the past.

But when you look at the breakdown of new versus returning customers, it should align much more closely with new customer acquisition. Without advanced configuration, it might be a 60/40 split, even in the best situations.

Limitations and realistic expectations

Nothing about this is foolproof. Audience exclusions don’t match perfectly. Brand exclusions don’t catch every variation. Customer Match has its gaps. So even with all four steps in place, some percentage of your spend will still hit warm audiences.

But for the first time, you actually have the levers to push Performance Max into upper-funnel territory. You can make it work like a real prospecting channel instead of a retargeting channel that takes credit for demand created elsewhere.

This matters most for brands spending heavily on Meta, TikTok, or other channels and wanting Google to actually grow the customer base rather than recycle the traffic those channels generate. If you’re seeing strong ROAS in Performance Max but flat new customer numbers month over month, this framework is for you.

If you’re a smaller brand still trying to find product-market fit or build initial momentum, this is probably overkill. Let Performance Max do its thing and pick up conversions without too many restrictions.

But once you’re scaling and the question is no longer “Can we be profitable?” but “Can we be profitable while growing the customer base?” these settings become some of the most important levers you have.

Every click they win is a customer you lose.

See where competitors are investing, which keywords drive their results, and how to capture more of the market.

See who’s stealing your traffic

Google’s giving you more control over PMax. Use it.

The conversation around brand versus non-brand is everywhere. You can’t throw a dart at a paid media conference without hitting someone with a strong opinion on it. But for some reason, almost no one seems to be testing this new option.

I just finished auditing an account spending $100,000 a month on Search with no Performance Max or Shopping, so they get purely new customer acquisition. We looked at their numbers and said maybe now’s the time to try this, exclude all these segments, and let it rip.

So here’s when I recommend implementing this test: if your ad spend is high enough (it doesn’t need to be $100,000 a month or anywhere near it), or you’re revisiting Performance Max. Your hypothesis should be that this approach increases the proportion of actual new customer conversions.

I think you’ll find that the needle moves further than you think.

Read more at Read More

How to approach build-versus-buy decisions for SEO

How to approach build-versus-buy decisions for SEO

AI has made SEO teams ambitious about what they can automate. Tasks that previously required engineering support can now be solved with the help of Claude or ChatGPT.

That’s exciting, but it also creates a new problem: thinking you can automate everything. In modern language, that often comes down to one question: Should we build or buy this new tool?

This build-versus-buy dilemma has never been simple, and AI has made it even more complicated. The challenge goes beyond cost. It involves security, maintenance, data access, internal capabilities, workflow fit, and whether a custom solution will remain maintainable, reliable, and useful six months from now.

How AI lowers the barrier to building

AI has lowered the barrier to experimentation. Even without technical knowledge, you can now create a custom GPT, build a workflow, connect data sources, or create an internal AI assistant.

But that doesn’t mean the same person can build and maintain a tool that will remain reliable over the next few years.

In most cases, AI can help SEO teams analyze data, identify patterns, summarize information, and recommend actions. It can save a lot of time, and teams that ignore AI are clearly falling behind.

But, at least for now, AI isn’t doing truly creative work in the same way humans do. It works from existing patterns and predicts likely outputs. That may change in the future.

AI also comes with hidden costs. Internally built tools are often treated as free because the invoice usually doesn’t sit with the SEO team. But that doesn’t mean token usage, API calls, infrastructure, engineering time, security reviews, and maintenance don’t cost money.

We are already seeing this effect. Reuters has described it as “corporate AI sticker shock,” with companies struggling to forecast usage-based AI costs. TechCrunch also reported that Uber introduced AI spending caps after blowing through its annual AI budget in four months.

Today, marketing teams aren’t the heaviest AI users, especially compared with engineering teams. But that can change quickly.

And when usage grows, the bills will grow too. That will naturally make companies ask which AI tools and AI-powered workflows create value and which ones only consume budget.

Be the brand AI recommends.

See where your brand appears in AI search, where competitors are winning, and what it takes to become the answer AI recommends.

See your AI visibility

Start by defining what you need

Before deciding whether to build or buy, SEO teams need to define what they really need.

Different ways to use AI and automation

Many teams group these solutions together, but they vary significantly in cost, complexity, and maintenance requirements.

  • A custom tool: A more complex internal system that usually needs engineering support. It is often more about automation, but it can have an artificial intelligence aspect.
  • A custom workflow: A repeatable process built with different tools, such as a custom GPT, Claude project, spreadsheet, reporting template, and so on. It often includes automation, for example, a scheduled task in an AI tool, and it usually has an artificial intelligence layer.
  • A custom layer on top of SaaS: Using data from existing tools and shaping it into your own reporting, prioritization, or recommendation workflow.
  • A true AI agent: A system that can take more autonomous actions. For example, it can scan your Slack and follow up with people you are still waiting on.

These aren’t the same, but people often label them incorrectly. Calling everything an “AI agent” creates confusion and can lead to wrong estimates about cost and complexity.

Look for repetitive, context-rich tasks

We’re still experimenting. Most of what our team has built focuses on daily tasks that require a lot of manual work.

For example, we’ve created a custom GPT that evaluates whether our content matches our personas and their pain points. The goal is not to replace the human copywriter or reviewer. It is to determine whether a piece remains generic and whether a few additions can make it more relevant.

We are also using AI for translations, monthly reporting, and a weekly summary that combines meeting notes, Slack, and Jira, and helps me see whether I have missed adding a task to Jira or where I still need to follow up.

One of our latest workflows transforms recorded internal meetings into organized landing page briefs.

These types of tasks are good candidates for AI-powered custom workflows because they rely on internal context, repeatable processes, and company-specific knowledge.

Get the newsletter search marketers rely on.


Not everything should be built

One example from our team was a prompt tracking tool that my colleague vibe-coded. It worked well as a starting point. But the data presentation was not perfect, and it was hard to create a trend graph without additional manual steps.

Soon, it became a maintenance burden because every external change in any of the LLM tools required fixes, for which we needed engineering help.

The real issue was reliability. For AI visibility and prompt tracking, we needed consistent data in one place, presented in a way we could analyze over time. That is why we moved to a specialized platform like Peec AI instead of continuing to maintain our own version.

That experiment was still valuable. It helped us understand the problem, the complexity, and the features we actually needed from an external vendor.

And this is one of my pieces of advice: whether you want to build a tool internally or buy one, always test what is already available on the market. Only then will you really understand what you actually need. You may think you need 10 features, only to realize you use only three.

For business-critical tools such as rank and AI visibility tracking, and website crawling, small SEO teams without dedicated technical support should usually be careful about building from scratch. If the data is fundamental to decision-making, reliability should be your main decision factor.

Use AI where your data already lives

Buy the crawler, rank tracker, or AI visibility platform. Then focus your internal efforts on connecting data from these tools to custom information, such as your GA and GSC accounts or even CRM data. Once connected, create reports that combine all these sources and enable you to analyze everything in one place.

MCP connections are also worth considering. The Model Context Protocol is an open standard for connecting AI applications to external systems, data sources, tools, and workflows. With MCP servers, you can analyze data from your primary tools directly using AI, taking your current workflows to the next level.

This doesn’t mean you’re required to learn how to code. But they need to know enough to ask the right questions.

If a tool connects to an internal knowledge base, customer data, or proprietary research, you should be aware that this could pose a security risk. And it might turn out that it is better for the company to dedicate an engineer to support you rather than risk exposing sensitive information.

You should also understand what the final cost will be for your company when you decide to go with a custom tool. Custom tools aren’t free just because the invoice doesn’t sit with SEO. Engineering time, security reviews, AI tokens, and API usage are all part of the cost.

Before asking leadership for a tool, SEO teams should be able to explain the workflow problem, the expected value, the cost of buying compared with the estimated cost of building, and what might happen if nothing is done.

The best requests don’t start with: “We need this tool.”

They start with: “Here is the problem, here is why it matters, here is what we’ve tested, and here is the best way we think we can solve it.”

How to prioritize what to build first

There’s no single prioritization matrix that will work for every situation.

A website crawler, a content evaluation tool, a report builder, or a competitive intelligence system can’t be judged by the same criteria.

If you are in a situation where you think you need more than one tool, start by mapping your current workflow and what your ideal situation looks like.

Once you do that, the patterns will be clear. Often, your strongest priorities will fall into two groups.

The first are tools that can support revenue creation. SEO teams are usually part of the marketing organization, and marketing is expected to bring visibility or leads. If a tool can help identify content opportunities, improve conversion rates, increase AI visibility, or surface gaps versus competitors, it can be seen as a priority.

The second group is workflows and tools that can help you minimize repetitive manual work. This category may not create revenue, but it will give your team time back to focus on more strategic work.

Don’t forget that quick wins also matter. Stakeholders don’t want to wait three months before seeing results. A smaller project that can bring value in three weeks will help you build trust and make it easier to get support for bigger initiatives.

Cross-team value should also be part of your decision.

SEO problems are often not just problems for your team. Competitive intelligence, for example, matters to PPC, ABM, content, product marketing, and sales, too. If several teams share the same pain, the business case becomes stronger.

So don’t be afraid to act as a cross-team synchronization layer when needed. Talk to the same teams you have already worked with, and try to understand their workflows and pain points, and where your needs overlap.

And remember, the best tool is not always the most ambitious one. Starting with something small is often the smartest move.

If AI can’t find you, customers won’t either.

Track your visibility across AI search, uncover missed opportunities, and grow your presence where customers are asking questions.

See your AI visibility

Good decisions start with proper scoping

AI has made it easier to build, but that doesn’t mean you don’t need to think about what really needs to be built.

Before deciding whether to build, buy, or customize, take the time to properly scope the work.

  • Understand the problem, the value you expect, who will use the solution, and who will maintain it after launch.
  • Talk to your team and other teams. Determine whether this is only an SEO problem or a wider business problem.
  • Don’t build because AI makes it possible. Don’t buy because a demo looks impressive.

Without proper scoping, you can end up with an expensive SaaS tool that doesn’t fit your workflow or an internal tool your team can’t maintain.

Always think first. Dedicate enough time to scope properly. Then decide whether to build, buy, or customize.

Read more at Read More

Google Search Console AI performance reports rolling out to more users

Google has confirmed that it has expanded access to the new Google Search Console AI performance reports to more users. Google’s John Mueller wrote on Bluesky, “We’re just rolling these out incrementally to sites, and reviewing the feedback along the way. I know everyone wants the new shiny thing immediately… but first, patience.”

AI performance report. The report shows you how well your content and websites are performing in AI responses, AI Mode, and AI Overviews in Google Search. The reporting includes impressions, pages, countries, devices, and dates, but does not include click data. 

Expanding access. This morning, I spotted a number of SEOs posting that they are now seeing the report, and that it is not restricted to sites in the United Kingdom. Some are seeing the report for sites in the United States, India, Switzerland and so forth.

And as I quoted above, John Mueller from Google confirmed the search company is “rolling these out incrementally to sites.”

What it looks like. Here is a screenshot of this report:

Why we care. Site owners and publishers have been asking for controls over whether and how their content is shown in Google’s AI features since Google launched these a couple of years ago. Well, Google is rolling out this feature to more of its users today. It is unclear how soon everyone will gain access to these controls, but I am surprised that Google expanded access this quickly. Specifically within 20 days from its first release.

Read more at Read More

Why some channels reward breadth and others require commitment

Why some channels reward breadth and others require commitment

Many budget allocation strategies assume that every channel follows the same pattern: the first dollar is the most productive, and each additional dollar yields a slightly lower return.

The charts below show what that pattern looks like.

The log shape means that the first dollar is the most productive, and each subsequent dollar is worth a little less. When every channel looks like that, the game plan is to spread the budget to as many channels as possible and equalize the marginal CPAs to maximize profit.

But not every channel looks like that. Some have a warm-up region where the early spend is the least efficient, not the most. On those channels, the logic above breaks, and so does the “test small, scale the winners” playbook that most of the industry runs on autopilot. 

The difference comes down to one question about the channel: Is the response curve C-shaped or S-shaped?

The answer can change how you approach channel testing and channel measurement, including any MMM analysis. Moreover, Google has been incorporating more S-shaped campaign types, and after its Google Marketing Live announcements, this trend seems set to continue.

The two shapes — and the only part that matters

The response curve plots output (conversions, revenue) against input (spend). This generally results in two types of curves in marketing.

  • C-shaped (concave): Diminishing returns from the very first dollar. A log or power curve. Picture the top-left quarter of a circle: steep at the start, flattening as you go.
  • S-shaped (sigmoid): A slow, inefficient start, then an inflection point where it gets steep, followed by a flattening into saturation. A logistic curve.

The response curve itself isn’t what you allocate against. You allocate against the marginal curve, the derivative, which answers the question: “What did the next dollar buy me?” That’s where the shapes diverge in a way that matters.

  • For a C-curve, marginal return is highest at the first dollar and falls in only one direction. Marginal CPA rises from the first dollar onward. If conversions are a*ln(s), marginal conversions per dollar are a/s, so marginal CPA is s/a, climbing in a straight line as you scale. There’s no warm-up. The cheapest conversion you’ll ever buy is the first one.
  • For an S-curve, marginal return starts low, rises to a peak at the inflection point, then falls. Marginal CPA is U-shaped. It’s expensive at the start, bottoms out around the inflection point, then climbs into saturation.

That region of increasing marginal returns is the whole story. It’s the difference between a channel where small budgets are productive and one where they are wasted.

See exactly how your competitors win.

Uncover the keywords, ads, landing pages, and strategies driving your competitors’ paid search success—and find your next opportunity to outperform them.

Analyze your competitors

How this looks in a marketing campaign

Say your CPA goal is $50. Here is an S-shaped channel, modeled as Conversions = 1000 / (1 + e^(-0.25(s – 20))), with spend in the thousands and the inflection at $20,000/month:

Run the $10,000 test that a sane person runs before committing real budget. Average CPA comes back at $132, marginal around $94. If those two metrics are all you look at, you conclude that this channel can’t hit $50, so let’s kill it.

That verdict is wrong. At $20,000 to $25,000, the channel is running at an average of $32 to $40, and the marginal dollar in the $15,000 to $25,000 band costs $18. That’s not “barely viable.” In that band, it’s the best marginal buy you have. The small test fell within the warm-up and reversed the conclusion.

In a C-shaped channel, the small test would have shown you the best the channel can do. On an S-shaped channel, it shows you the worst.

This is the trap. The standard playbook is “test small, scale what works.” On S-curves, small tests systematically condemn channels that would’ve worked at scale because the test is structurally stuck in the inefficient region.

Get the newsletter search marketers rely on.


The allocation logic, restated

C-shaped channels, go wide

The optimization is convex. There’s one global optimum, the equimarginal rule from the marginal-CPA post applies cleanly, and the solution is usually interior, meaning lots of channels get funded.

Even a small allocation is productive because the first dollar is the best dollar. Run many channels lean, reallocate continuously at the margin, and pull back the instant marginal CPA crosses your goal.

S-shaped channels, go deep or skip

The optimization is non-convex. A small allocation can be strictly worse than zero because below the inflection your marginal return sits under your target, and you’ve sunk money to get nowhere.

The decision isn’t “how much.” It is binary: commit past the threshold, or don’t fund it at all. There’s a real minimum viable budget, and it’s often above normal test budgets. You can’t sprinkle an S-curve and expect efficiency, and you can’t evaluate one on an underfunded test.

Those two rules can look like they fight each other, but that’s only true to a certain point. Past the inflection, an S-curve is concave, so the equimarginal rule governs it exactly as it governs a true C. The S-specific instruction — commit a block instead of sprinkling — is only about the trip from zero to past the inflection.

Shape is therefore mostly a launch-and-evaluation problem. Getting a new prospecting channel into its efficient range requires a committed block and patience with ugly early numbers. Once it clears the inflection, you manage it at the margin like everything else, right up until you consider cutting it hard, where shape matters again because the downside is a cliff, not a ramp.

This is the part that’s genuinely counterintuitive, and it echoes the original marginal-return point: The right move isn’t always the one that looks most efficient at a small scale.

Which channels are which?

The historical default was concave. Simon and Arndt reviewed more than 100 studies and concluded that advertising follows the law of diminishing returns, a concave response. 

The dissent came later: Vakratsas, Feinberg, Bass, and Kalyanaram found that threshold effects do exist and that response is not necessarily globally concave. Their explanation for why thresholds were so hard to find is the useful part. Mature accounts already operate inside the effective range, so the warm-up never shows up in the data, and most studies fit a concave model (the double-log) that can’t reject an S-curve even when one is present.

The platform shift has made the threshold visible again. Here is a fuller map, ordered roughly from C to S. The shape column is an inference from how each system targets and learns, not a measured constant, and the right shape for your account still has to be measured.

Two rows do most of the work.

AI Max is the live example of a channel migrating from C toward S. Swapping explicit keywords for broad and keywordless matching means it needs conversion volume to learn which queries convert, so below a data threshold, it explores badly.

The mixed independent results fit that: Google reports about 14% more conversions on average and up to 27% for exact-match-heavy campaigns, while independent testing reports 84% of advertisers seeing neutral or negative results. Much of that spread is accounts that turned it on without the conversion volume to clear the learning region.

Performance Max is the trap, because its curve is a composite. It blends a harvesting layer (branded, retargeting, Shopping against existing intent) with a prospecting layer (keywordless expansion across surfaces). The harvesting layer is a cheap C that pays off on the first dollar. The prospecting layer is the S underneath.

Blended, the early efficiency looks great, because you are mostly skimming demand you already had, and the average hides the prospecting warm-up entirely. That is also why the platform is glad to optimize it for you: the blend flatters the headline number. You can’t read PMax or run the shape analysis on it until you split the harvesting from the prospecting.

The throughline runs in two layers. Rules-based auctions capture the best inventory first, which yields concavity; machine-learning systems must be fed before they are efficient, which introduces a threshold. Underneath both, harvesting existing demand is concave and mostly non-incremental, while creating new demand is the S-shaped part where the real growth and the real warm-up cost both sit.

Average versus marginal: total over spend, or the slope where you stand.

What you allocate against is marginal incremental return, the slope of the incremental curve at your operating point. A holdout fixes the first axis only. Time-sliced marginal CPA on attributed data fixes the second only. A multi-cell scaling test gets both, at a cost. 

MMM (method 1) estimates the whole curve from aggregate data and sidesteps click attribution entirely, but pays in identifiability and modeling assumptions instead. Most arguments about ‘what is working’ are two people standing on different axes.

There are two major cautions, and I would flag both as genuinely unsettled rather than settled facts. 

  • Separating a true S-curve from “concave with a high half-saturation point” is hard, because a concave model will fit S-shaped data well enough to hide the inflection (this is the Vakratsas point, and it applies to your own dashboards as much as to academic studies). 
  • The learning phase may be a one-time fixed cost to train the model rather than a permanent feature of the steady-state curve. If it is transient, the channel may behave concavely at the margin once it is trained, and the S you measured was a startup artifact. The truth is probably a mix: a one-time training cost, plus an ongoing minimum-volume requirement to stay efficient. Treat every shape call as provisional and re-check it.

One more failure mode, and this one is not unsettled science but a matter of where you are standing on the curve. An S only looks like an S if your data spans the inflection. 

Above the inflection, an S is concave, mathematically identical to a C. Look at only the $20,000-and-up rows of the table above: marginal CPA rises monotonically from $18, a textbook C-curve, and the convex warm-up is invisible because you are no longer operating in it. 

Established accounts usually sit past the inflection, which is exactly why Vakratsas found thresholds so hard to detect, and why you can run an S-shaped channel for years, correctly, while believing it is concave. The tell arrives the day you cut hard and fall off the inflection instead of easing down a slope.

When to go wide and when to go deep

The marginal-return post told you to equalize marginal CPAs across the program. That rule is still correct, but the shape of the curve tells you how you’re allowed to get there. 

  • On C-shaped channels, you can get there by sprinkling, because every dollar is productive and breadth is the natural answer. 
  • On S-shaped channels, you have to commit a block of budget past the inflection before the channel earns its place, and then concentrate rather than spread.

Lay the harvest-versus-create cut on top. Harvesting channels (branded, retargeting, non-brand search) are your C-curves: fund the first dollars, then cap them early, because they saturate fast and most of the tail isn’t incremental, no matter how strong the attributed ROAS looks. 

Prospecting channels (Meta, YouTube, LinkedIn, the expansion half of PMax) are your S-curves and your only real source of incremental growth: commit past the warm-up or don’t start, and judge them on incremental lift rather than attributed CPA, or you’ll kill the thing that was working.

Classic search rewards going wide. PMax, AI Max, and Meta prospecting reward going deep on fewer bets and giving each enough volume to clear the warm-up. Run an S-curve like a C-curve and you’ll starve it, read the underfunded result, and kill a channel that would’ve been one of your best.

Read more at Read More

Amazon launches Alexa+ Agentic Ads

Amazon is bringing transactions directly into advertising with a new format that allows consumers to discover products, ask questions and complete purchases entirely through a conversation with Alexa+, potentially shortening the path from ad impression to conversion.

What’s happening. Amazon today introduced Alexa+ Agentic Ads, a new advertising format designed to let customers move from seeing an ad to completing a purchase without ever leaving the Alexa experience.

The format launches with partners including Papa Johns for food ordering and artists Beck, Jill Scott and Omar Courtz for concert ticket sales. The experience is currently available on Echo Show devices.

Why we care. Alexa+ Agentic Ads remove the traditional handoff between an ad and a checkout page, allowing consumers to complete purchases directly within a conversation. For early adopters, that could lead to higher conversion rates, lower drop-off and a new way to capture high-intent customers at the exact moment they’re ready to act.

How it works. Unlike traditional digital ads that redirect users to a website or app, Alexa+ Agentic Ads keep the entire purchase journey inside a conversation.

Users can engage with an ad, ask questions, compare options, check availability and complete a transaction through natural language interactions with Alexa.

The goal: eliminate friction between interest and purchase.

Concert tickets become conversational commerce. Amazon is initially showcasing the format through live event promotions.

Fans who see an ad for an upcoming concert can ask Alexa about show details, review available seats, compare pricing and purchase tickets directly through the device. Purchased tickets are then delivered to their Ticketmaster account without requiring them to open another app or website.

The experience is designed to transform entertainment advertising from an awareness channel into a direct sales channel.

Food ordering gets the same treatment. The format also extends to restaurant ordering.

A customer looking for dinner ideas could encounter a Papa Johns ad and begin placing an order immediately. Because Alexa+ can draw on previous interactions and preferences, it may suggest favorite toppings or commonly ordered meals before completing the transaction.

The entire process—from ad exposure to order confirmation—takes place within the conversation.

What to watch. Alexa+ Agentic Ads could offer an early look at how AI assistants reshape digital advertising. If consumers become comfortable completing purchases inside conversations, brands may increasingly view AI assistants not just as discovery tools but as full-fledged commerce platforms.

Read more at Read More

An open letter to everyone hiring a search leader

Search unicorn

Anthropic’s latest job posting has the SEO industry abuzz. They may as well have titled it Search Gawd. The truth is, it’s everywhere.

To be transparent, I’ve written this job description a few times and interviewed for it. I’ve yet to see any of these roles get filled, but I’ll come back to that in a minute.

Sometimes the title is Head of SEO. Sometimes it’s Director of AI Search, VP of Search, Director of SEO, AEO and GEO, or — wait for it — Agentic Commerce GEO Consultant.

Lots of titles. The assignment is basically the same: own technical SEO, understand paid search, shape content, partner with engineering and product, build measurement, prepare for AI-mediated discovery, explain it to leadership, and turn it into growth.

The predictable reaction is that this is a lot of jobs rolled into one. An entire agency behind a single employee badge. Fair, but it misses the point.

Companies have been looking for this person for years. Generative search is just forcing the issue.

This is not an Anthropic problem

This morning’s search on the job boards: 

  • Victoria’s Secret: Director, AI & Organic Search (AEO, GEO, SEO), $152K–$216K.
  • Publicis / Starcom: VP, SEO (Performance Content).
  • Accenture: Agentic Commerce GEO Consultant.
  • SailPoint: AEO/GEO Manager.
  • AirOps: Senior SEO Manager spanning SGE, Perplexity, ChatGPT, Gemini.
  • Responsive: Senior Manager, Web Strategy — SEO, GEO, plus Next.js, React, Vercel, DNS.
  • Danaher, Experian Health, Amazon News: some version of SEO + AEO + GEO.
  • Anthropic: SEO Lead, $255K–$320K.

Different industries. Different price points. Same job, unwittingly all looking for the same person.

Even the titles are arguing with the job descriptions

Agency X is hiring a “Director, SEO/SEM” whose responsibilities contain no SEO — just paid search, SEM platforms, vendor management, and a team of seven.

Consulting firm Y is hiring a “Director, SEO/AIO,” where AIO appears to be an in-house acronym no one bothered to define.

An indy agency’s “VP/Director, SEO” lists paid search, paid social, and pharmaceutical marketing among the nice-to-haves.

A token research firm is hiring a “Director, SEO & AEO” whose responsibilities actually describe SEO and AEO work — rare enough to be worth mentioning.

If the company can’t agree on what the role is before posting it, the candidate has no chance of meeting expectations that were never written down.

The taxonomy says one thing. The JD says another. The recruiter screens for a third. The hiring manager interviews for a fourth. The ATS filters out anyone worth a shit.

Looking for the missing link

You need someone who can see across technical search, content, PR, product, engineering, analytics, performance media, and brand — and understand that those functions were never as independent as the org chart suggested.

Search has always exposed the seams. A technical problem can look like a content problem. A content problem can be a product problem. A visibility problem may be an authority problem, not an optimization problem. Paid search often surfaces a messaging problem before brand research does.

Generative discovery makes those dependencies impossible to ignore. When results become answers, SEO stops being a traffic function.

At the risk of going full Yoda to avoid AI-slop speak: found, information is, only if infrastructure allows it. Content makes it understood. Brand makes it trusted. Product turns discovery into use — or it doesn’t.

You’re not asking one person to execute every task. You’re asking one person to understand how the pieces connect. That person exists. Your chances of finding that person through a conventional scoring system are slim by design.

The résumé will not look the way you expect

The value of this candidate isn’t captured by years under an SEO title or a checklist of software. The value is judgment:

  • Knowing which technical issue matters and which is noise.
  • Recognizing when the content team can’t solve the content problem.
  • Knowing when to spend, when to automate, when to wait, and when to tell leadership to stop doing that.

That judgment is hard to capture on a résumé. The candidate may have moved through agencies, publishing, product, consulting, and operating roles. Their career may look less focused than a specialist’s. That’s precisely why they can do the job.

Your ATS will screen them out. Your recruiter will flag them as “non-linear.” Your hiring panel will note they haven’t held the title before. Well, the title didn’t exist before. No one can agree on what to call it.

You can see how this search is already going sideways.

A less charitable possibility

Some of these processes may be less about filling a role than learning from the people willing to interview for it.

Senior candidates diagnose. They explain how they’d structure the function, where the organization is weak, what the first 90 days should look like, which tools they’d buy, and which work they’d kill. Invite enough of them in, and a company can collect competing organizational models and strategic priorities without hiring any of them.

Perhaps that isn’t the intent. But when a role stays open for months, gets repeatedly reposted, changes title and scope, and produces interviews that feel more like advisory sessions, candidates are entitled to ask what the company is actually buying: talent acquisition or knowledge harvesting?

The solution isn’t a shorter job description

The breadth is real, so cutting half the bullets doesn’t make the work disappear. Decide what you want. Is it:

  • A specialist who will execute?
  • A leader who will build a team?
  • An executive who can connect search, content, product, brand, and performance?
  • A consultant who can tell you which one you need?

Those are different jobs. Pretending they’re one role and waiting for a unicorn isn’t a strategy.

A closing note, since you asked

I would, however, be very good at the job. So would a handful of others who’d get screened out for the same reason.

The Anthropic job? Not getting it.

Five years under a title that didn’t exist five years ago — I don’t have them. My résumé reads like the job spec itself, in exactly the shape an ATS is built to reject. It’s an easy system to game. So easy that anyone worth their salt knows how.

The missing link is real. Generative search didn’t create it; it just made it harder to ignore. Before you hire someone to connect these systems, make sure your company can recognize them, hire them, and let them do the job.

The company that figures out how to recognize the candidate—not just write the job description—quietly wins the next decade while everyone else argues on LinkedIn about whether GEO is a word.

Read more at Read More

What Is an AI Citation Audit & What Can It Tell You About Your Content

Key Takeaways

  • An AI citation audit tells you, on a per-topic and per-platform basis, where your visibility gaps come from and what type of action closes each one.
  • The majority of citations driving AI responses typically come from third-party sources, not brand-owned pages. Competitors appear because independent sites reference them, not because their own content is being surfaced.
  • High-volume, low-differentiation content faces the highest displacement risk in an AI environment. Generic how-to guides are exactly the type of content AI can synthesize without sending users anywhere.
  • The goal of content strategy shifts from answering every possible question to being present with genuine authority in the specific contexts that matter to your buyers.

If you’ve been tracking your brand in AI tools and wondering why the data isn’t telling you anything useful, the problem is usually upstream: generic prompts, the wrong measurement model, inputs that don’t reflect how real buyers actually search. In an earlier piece, I introduced a structured framework for fixing it. This post is about what happens once the framework does its job.

Once you have well-constructed prompts, two layers of metrics, and a clear picture of where your brand appears across AI platforms, you get a specific and actionable output: a citation audit. Understanding what is an AI audit and what it tells you is where measurement becomes strategy.

The citation audit sorts your visibility gaps into three categories: gaps that require digital PR, gaps that require owned content, and gaps that point to social and community management. Each category demands a different type of response. And the pattern running across all of them points to the same conclusion: the content playbook built around maximizing coverage and keyword volume is losing ground to one built around genuine authority and relevance.

This post makes that argument concrete, and closes the argument with the strategic implication that follows.

What the Citation Audit Actually Shows

Once the structured topical analysis is complete, the methodology exports citation data for the highest-opportunity topics on each platform. That data breaks down across three dimensions.

Third-party content accounts for the bulk of what AI is drawing on. In most audits, well over 80 percent of highly cited pages come from independent sources: sector publications, accounting and advisory firm blogs, business setup consultancies, and regulatory guides. These are not the brand’s own pages. They are pages where the brand (or a competitor) is mentioned in the context of explaining something broader.

Owned content plays a smaller role than most teams expect, but it’s not irrelevant. Specific owned pages, particularly long-form guides that cover a topic with genuine depth, do earn citations. The issue is that most brands’ owned content skews toward service pages and thin category coverage, which AI systems have little reason to cite when better third-party resources exist.

Social and UGC signals are a smaller but growing dimension. Platforms like Reddit and Quora appear in citation data for certain topic types, particularly those involving peer experience, comparisons, and community knowledge. This is an underserved channel for most brands.

The example below shows how this ecosystem applied to one NP Digital client that we worked with.

The executive summary of results from an AI visibility audit NP Digital conducted for a client.

In one audit, roughly 80 percent of highly cited pages for compliance-related topics came from independent accounting, tax, and audit firms. The brand’s own content was rarely surfaced. Competitors appeared not because of anything they had published directly, but because third-party sites were using them as examples when explaining regulations and requirements. Visibility was earned indirectly, through the content ecosystem, not through the brand’s own pages.

The Coverage Trap

To understand why this matters strategically, it helps to understand the model it’s replacing.

The coverage mindset that drove SEO content strategy for the past decade wasn’t irrational. Traffic was the primary currency. Search engines rewarded breadth. The more questions you could answer, the more pages you could rank, and the more traffic you could capture and convert at the margin. Publishing at volume made sense.

Alt text: Two-column diagram contrasting devalued generic content types on the left with high-value authoritative content types on the right, illustrating the shift from coverage to authority in an AI search environment.]

That model is breaking down in an AI environment, and the citation audit is where you see it most clearly.

AI systems are built to synthesize and summarize. Content that exists to answer broad, generic questions is exactly the type of content AI can handle on its own, without sending users anywhere. A page explaining what SEO is, or listing the top ten CRM tools, or walking through a basic how-to process is precisely the type of content that gets absorbed into an AI response rather than cited as a source.

The more your content resembles what an AI would generate from a basic prompt, the less reason an AI has to cite you. This is the coverage trap: scaling the old model doesn’t just fail to improve AI visibility; it actively increases exposure to displacement.

A graphic showing content strategy is shifting from SEO content to original research and authority.

What AI Systems Actually Cite

The citation audit goes beyond revealing gaps to reveal patterns in what earns citations, and that pattern is consistent across topics and platforms.

Citations go to content that demonstrates genuine expertise in a specific context versus the biggest brand or highest-traffic page. Original research with proprietary data. Long-form guides that go deeper than the obvious. First-hand experience presented with authority. Comparison content that places competitors in context rather than avoiding them.

The pattern from real audit work: educational long-form guides consistently outperform service pages. Content that mentions competitors as examples within broader category coverage drives more citations than content focused exclusively on the brand. Pages that answer a specific, high-intent question with real depth earn citations.

This is a function of what the content actually contains. AI systems are drawing on content that has established a genuine association with a concept, problem, or use case. That association is built through depth, specificity, and demonstrable expertise, not through breadth of coverage.

Table showing that for both compliance and banking topics, long-form educational guides from third-party sources dominate AI citations, with brands mentioned as examples rather than as primary sources.]

The practical implication: AI SEO strategy stops being about answering every question and starts being about answering specific questions better than anyone else. That’s a meaningful shift in how content is briefed, produced, and measured. Good AI keyword research makes that brief concrete, identifying exactly which topics and contexts to prioritize.

Three Actions That Close the Gap

The citation audit produces a specific output: for each topic cluster and each platform, it identifies which type of action is most likely to close the visibility gap. Those actions fall into three categories, each with different resource requirements and timelines.

Digital PR Owned content Social / UGC
Earn third-party mentions Partner with publishers AI draws on. Contribute expert commentary. Be included in sector guides. Build authority content Comprehensive guides, comparison pages, original data. Topics the audit identifies as underserved. Community presence Be credible where buyers research before reaching your site. Longest runway, growing signal weight.
Fastest impact Citations driven by external mentions, not owned pages Medium-term Depends on topic gap size and content quality Longest runway Matters increasingly as AI incorporates social signals

Digital PR and third-party mentions are the highest-leverage activity for most brands, because they address the most common finding: that the majority of AI citations are coming from independent sources, not owned pages. The goal is to be embedded in the content ecosystem for your topic. That means partnering with the publications, advisory firms, and consultancies that are producing the content AI draws on. Contributing expert commentary, providing authoritative reference material that others can link to, and collaborating on guides where your brand appears as a contextual example alongside competitors. 

Owned content investment is the right response when the citation audit shows that your owned pages are genuinely absent from the topic, not just outperformed. The priority isn’t more content; it’s better content in the right areas. The audit identifies exactly which topics are underserved. The content itself needs to be the type that AI systems and third-party sites can cite: comprehensive guides that cover a topic with real depth, comparison pages that place your offer in context, step-by-step process guides built around specific use cases, and, where possible, original data or analysis that doesn’t exist elsewhere. Depth and specificity earn citations. Breadth and volume don’t.

Social and community presence is the response when visibility gaps are driven by UGC signals, typically in topics where buyers seek peer experience and independent comparison rather than brand-produced content. Community management in the right channels, credible participation in conversations on Reddit, Quora, and industry forums, and authentic engagement rather than promotional presence. This is the longest runway of the three, but it’s growing in importance as AI systems increasingly incorporate social signals into what they surface.

The Bigger Picture: Presence Over Position

Traditional search was about position. Rank highly, earn traffic, convert at the margin. Visibility was a number: position one, page one, top ten. You knew where you stood, and you optimized to move up.

AI-driven search works differently. A brand can shape what users learn about a category, influence the answer to a high-intent question, and be present at the moment a decision is forming, all without appearing as a link. Visibility is no longer a rank. It’s a probability: how likely are you to be present when it actually matters?

The brands that understand this earliest are building an advantage that compounds. Not because they’ve found a new SEO trick, but because they’ve shifted their content investment toward genuine authority in specific contexts, and that authority is what AI systems consistently draw on.

That’s the conclusion the citation audit points to, and it’s what makes AI visibility tools genuinely useful when they’re used right. They serve as a diagnostic that tells you where authority is missing and what to build next.

Success in this environment is defined by presence, not position. The content strategy implications follow directly from that.

FAQs

How do you audit AI search optimization response analysis?

Start by running structured prompts across the major AI platforms, covering the topics most relevant to your buyers’ decision-making process. Analyze which pages are being cited in responses to those prompts, and categorize them by source type: third-party, owned, or social. The distribution tells you where the gap is coming from and what type of action closes it. Secondary metrics, including run length, entropy, and Gini coefficient, reveal how stable your visibility is and how competitive each topic is.

How do you use AI for a content audit?

An AI citation audit is a specific type of content audit that goes beyond traditional performance metrics. Rather than measuring traffic or rankings for your owned pages, it measures how often your brand and content appear in AI-generated responses to relevant prompts. The output identifies which topics are underserved, which content types earn citations, and whether the gap requires digital PR, new owned content, or community presence. It connects content decisions directly to AI visibility outcomes.

 How do you audit for AI search visibility?

Build a structured set of prompts using the SPIV framework, grounded in your actual buyer personas and intent stages rather than generic category terms.

Pair that with AI keyword research to identify the topic gaps the audit surfaces, and you have a complete workflow from measurement to action.

Run those prompts across ChatGPT, Google Gemini, Perplexity, and Google AI Overviews on a recurring basis. Track both primary metrics from the platform and secondary metrics calculated on top of the export data. The citation analysis, which identifies what sources AI is drawing on and where your brand appears in that ecosystem, is the layer that tells you what to do next.

Conclusion

This series started with a measurement problem.

Most teams tracking AI visibility are using deterministic tools to measure a probabilistic system, running generic prompts that describe buyers who rarely exist in practice. The data looks clean. The picture it paints isn’t representative.

The response to that problem was a methodology: structured prompt construction grounded in real buyer personas and intent stages, a two-layer metric system that separates surface-level visibility from genuine diagnostic insight, and a modular audit format that makes the output actionable rather than overwhelming.

What the citation audit adds to that is the strategic implication. AI visibility is built primarily through third-party mentions, not owned pages. Coverage-first content is the most exposed to displacement. Genuine authority in specific, high-intent contexts is what earns consistent citations. The content investment that follows from that is about producing the right things, in the right depth, for the contexts where decisions actually happen.

The brands that make that shift now will hold ground as search continues to change. The ones that don’t will keep producing content that looks healthy in their dashboards while becoming invisible in the moments that matter most.

Read more at Read More

How We Rebuilt AI Visibility Measurement From The Ground Up

Key Takeaways

  • The core problem with most AI visibility prompts isn’t that they’re wrong; it’s that they’re missing the context real users bring. Generic inputs produce generic, unactionable data.
  • The SPIV framework (Segment, Persona, Intent, Variable) structures prompts around four variables drawn from real user data, turning stateless AI visibility tracking inputs into high-fidelity user proxies.
  • Once prompts are grounded in real context, the variation you observe in model responses becomes informative rather than noise. Visibility can then be expressed as a probability distribution.
  • Measurement operates on two layers: primary metrics from the tracking platform, and a secondary layer of calculated metrics (run length, Shannon entropy, Gini coefficient, and KL divergence) that reveal the stability and competitive dynamics behind the surface numbers.
  • This approach naturally connects measurement to business priorities. It becomes much harder to justify tracking low-intent queries with no connection to how your product is actually bought.

The first post in this series made the case that most AI visibility tracking is built on the wrong foundation: generic prompts measuring hypothetical users, deterministic tools applied to a probabilistic system. If that diagnosis is right, the obvious next question is: what does a better approach actually look like?

That’s what this post covers. What we built at NP Digital to address both the measurement problem and a second issue that compounded it: early AI visibility audits were trying to do too much at once, producing outputs so dense that clients couldn’t identify a single clear action to take. The rebuild addressed both problems together.

The result is a methodology built around structured prompt construction, two layers of metrics, and outputs that point to specific, defensible actions. Here’s how it works.

Why the Old Audit Approach Wasn’t Working

Before explaining what we built, it helps to explain what we were moving away from, and why.

Early AI visibility audits, including our own initial attempts, were structured like SEO audits. A single document tried to cover everything at once: a content audit, a competitor audit, a structured data review, citation analysis, and strategic recommendations, all bundled into one output. The logic made sense at the time. SEO audits had always worked this way. Why would a GEO audit be different?

The answer, in practice, was that clients couldn’t use them. Data points conflicted. The strategic direction wasn’t clear. The same document had to be re-presented multiple times before anyone could agree on what to do first. We were producing thorough work that left clients more confused than when they started.

Two problems were running in parallel. The first was the measurement problem I covered previously: generic prompts producing data that looked meaningful but wasn’t representative of real buyer behavior. The second was a presentation problem: even if the data had been better, the format buried the signal in too much noise.

A comparison of different approaches to building topic clusters.

The rebuild addressed both. On the measurement side, we moved to structured prompt construction through the SPIV framework. On the output side, we separated the analysis into discrete, digestible pieces: each focused on a specific topic cluster, each pointing to a defined type of action. Clients stopped needing multiple sessions to understand what they were looking at.

Introducing the SPIV Framework

The starting point is familiar data. The same sources that feed traditional keyword research, including People Also Ask results, Google Search Console data, community platforms like Reddit and Quora, and first-party data like customer service transcripts where available, provides the raw material. The difference is what happens next.

Instead of using those inputs as-is, SPIV treats them as raw material and injects four structured variables into each prompt. The practical effect: it turns stateless AI keyword research inputs into pseudo-stateful responses by giving the model the persona context it would otherwise be missing.

The S.P.I.V.framework explained.

Each variable does a specific job:

  1. Segment: The market category or business context. Grounds the prompt in a defined situation: ‘SME owner in the UAE’ rather than ‘business owner.’ This is the broadest layer of context.
  2. Persona: The specific user type, including relevant traits: risk tolerance, level of prior knowledge, geographic or professional context. This is where abstract ‘users’ become real people with real constraints.
  3. Intent: What the user is actually trying to accomplish, not the topic they’re searching but the outcome they need. ‘Understand my compliance obligations’ is different from ‘find the cheapest option.’ Separating these surfaces meaningful differences in how models respond.
  4. Variable: A single modifier that can be shifted to test sensitivity: ‘fastest’ vs. ‘cheapest’ vs. ‘most reliable.’ Isolating one variable at a time makes the data interpretable. Change everything and you can’t explain what moved.

The table below shows what this transformation looks like in practice, using anonymized examples from real audit work:

A prompt optimization metrics for AI visibility audits.

The difference between the raw input and the SPIV-optimized prompt isn’t cosmetic. The raw prompt describes no one in particular. The optimized prompt describes a specific person in a specific situation trying to accomplish a specific outcome. That specificity is what makes the model’s response meaningful as a measurement input.

A well-constructed set of SPIV prompts doesn’t need to be large. Representativeness matters more than volume. A focused set of 15 to 30 prompts mapped to your key buyer personas and intent stages gives more actionable signal than hundreds of generic variations.

The Two Layers of Measurement: Primary and Secondary Metrics

Once prompts are properly constructed, the analysis operates on two distinct layers. Understanding the difference between them is what makes the output useful rather than just interesting.

Primary metrics come from the tracking platforms directly, including Writesonic and Profound. These include visibility percentage, share of voice, and mention frequency. They’re the standard outputs most teams are already familiar with and they provide the baseline picture: how often does your brand appear, and how does that compare to competitors?
 
The four secondary metrics, and what each one tells you:

  1. Run length: The number of consecutive days a brand maintains visibility for a given topic. Short run lengths signal volatile, unreliable presence. Long run lengths indicate that the model has formed a stable association between the brand and that topic, what we’d call persistent authority rather than a transient mention.
A guide to interpret run length in an AI visibility edit.
  1. Shannon entropy: A measure of how evenly visibility is distributed across the brands appearing for a given topic. High entropy means no brand dominates, meaning the model is pulling from a wide, fragmented field. Low entropy means the results are concentrated, and that a small number of brands are taking most of the mentions. Low entropy topics are harder to break into; high entropy topics are more contestable.
  2. Gini coefficient: Where Shannon entropy tells you how distributed results are, the Gini coefficient tells you the degree of concentration. A high Gini score means visibility is dominated by one or two brands. A low score means the field is relatively open. Together with entropy, this gives a picture of whether a topic is winner-takes-most or genuinely shared.
A chart to interpret the Gini coefficient  in an AI visibiity edit.
  1. KL divergence: In a traditional statistical context, this metric measures how a distribution changes over time. We’ve adapted it here to serve a different purpose: measuring how far an individual platform’s results drift from the group average across all tracked platforms. A low score for a given platform means its brand rankings for that topic are broadly in line with the consensus across ChatGPT, Gemini, and Perplexity. A high score means that platform is picking a significantly different set of brands. That’s a meaningful finding. It tells you whether your visibility is genuinely broad or whether it’s concentrated in one model’s view of the world.
A guide on interpreting KL divergence for AI visibility edits.

None of these metrics is useful in isolation. Run length tells you how stable your visibility is; entropy and Gini tell you how competitive the topic is; KL divergence tells you whether that visibility holds across platforms or is fragile in a way your headline numbers don’t reveal. Read together, they give a diagnostic picture that primary metrics alone can’t produce.

What the Data Tells You

With SPIV-structured prompts and both metric layers in place, visibility stops being a single number and becomes a probability distribution. The question changes from ‘where do we rank?’ to ‘how reliably do we appear when the conditions that actually matter are present?’

In practice, this approach surfaces findings across three dimensions that generic tracking misses entirely.

The visibility distribution itself. Some brands are category staples: they appear consistently across multiple runs of the same prompt, across slight variations in phrasing, across different platforms. Others are volatile outliers: they surface occasionally but can’t be relied on. Generic tracking averages this out and produces a headline figure that obscures the difference. The secondary metrics separate the two clearly.

A graphic explaining how visibility should be defined when it comes to AI/LLMs.

The platform dimension. Visibility that holds on Google Gemini but not on ChatGPT is a meaningful finding, not just a data point to average away. Different models draw on different training data, weigh different source types, and respond differently to the same underlying intent. KL divergence makes this visible. A brand that appears strong in aggregate but has a high divergence score on one platform has a concentration risk that matters strategically, especially if that platform is where your buyers actually research.

The topic dimension. This is often the most strategically important finding in the whole audit. Brands regularly show strong visibility in broad, low-intent queries (the general category terms that show up well in standard tracking), but near-zero presence in the specific, high-intent topics their buyers are researching at the point of decision.

In one audit, a brand showed visibility above 65 percent for general licensing topics across platforms. For compliance and banking topics (the two areas most directly connected to their buyers’ decision-making process), visibility was zero across ChatGPT, Google AI Overviews, and Perplexity. The standard tracking looked healthy. The actual picture was that the brand was invisible at the moments that mattered most.

Generic prompts miss this because they aren’t asking the right questions. SPIV-structured prompts surface it because they’re built around the contexts where decisions actually happen.

This is also where the measurement connects directly to AI SEO strategy. Once you know which topics show gaps, which platforms are most divergent, and which competitors are holding the positions you’re not, you have a defensible brief for content and PR investment. The audit doesn’t just tell you where you are. It tells you where to go.

FAQs

How do you track AI visibility?

Tracking AI visibility starts with a defined prompt set run across the major platforms: ChatGPT, Google Gemini, Perplexity, and Google AI Overviews. Tools like Writesonic and Profound automate this process and export visibility data by brand and topic. The critical step most teams skip is structuring those prompts around real buyer personas and intent contexts rather than generic category terms. Generic prompts produce directional data; structured prompts produce data you can act on.

How do you monitor brand visibility in AI?

Brand visibility in AI is monitored by running structured prompts across platforms on a recurring basis and tracking both primary metrics (visibility percentage, share of voice) and secondary metrics (run length, entropy, Gini coefficient, KL divergence). The primary metrics tell you what the numbers are. The secondary metrics tell you whether those numbers are stable, how competitive the topic is, and whether your visibility is genuinely broad or concentrated on a single platform. Monitoring both layers gives you a picture you can act on.

How do I check AI visibility of my brand?

Start by identifying the topics most relevant to your buyers’ decision-making process, not just the broad category terms, but the specific questions they ask when they’re close to a purchase. Build prompts around those topics using the SPIV framework, run them across ChatGPT, Gemini, Perplexity, and Google AI Overviews, and track how consistently your brand appears. The gap between your visibility in general topics and your visibility in high-intent, decision-stage topics is usually the most important finding.

Conclusion

The shift this methodology makes is simple to state but significant in practice: you’re no longer tracking where you rank. You’re tracking how reliably you appear when it actually matters: for the right persona, at the right intent stage, on the platforms your buyers actually use.

SPIV is how you build the inputs that make that measurement possible. The secondary metrics are how you make sense of what the data is telling you. Together, they turn AI visibility from a headline number into a diagnostic that points somewhere useful.

Knowing where you’re visible and where you’re not is only half the equation. In the final post in this series, I’ll cover what this framework reveals about content strategy, and why the old volume-first approach doesn’t hold up in an answer-driven search environment.

Read more at Read More

Web Design and Development San Diego

Help Us Pick the Next Stop in Europe for Search Central Live Deep Dive 2026!

As we mentioned a
few months ago,
we are bringing the Search Central Live Deep Dive format to the EMEA region. This SCL format
requires finding the absolute best home for the event—a place where all of you can truly
connect, learn, and enjoy.

Read more at Read More

AI Brand Visibility: You’re Tracking It Wrong

Key Takeaways

  • Most AI brand visibility tracking today replicates keyword tracking logic, using prompts instead of search terms. The underlying assumption is the same, and that’s the problem.
  • Traditional search engines are deterministic: the same query tends to return similar results. LLMs are probabilistic: the same prompt can produce a wide range of valid answers.
  • Measuring a probabilistic system with deterministic tools produces data that looks clean but doesn’t reflect how the system actually behaves.
  • The prompts most brands are tracking (‘Best CRM in 2026,’ ‘Top accounting software’) describe a user who doesn’t exist, someone with no context, no history, and no specific intent. This is a known gap in current AI SEO measurement approaches.
  • Fixing this requires a different measurement philosophy, not just better prompts.

Have you started tracking your brand in ChatGPT, Perplexity, or Google AI Overviews? Good. You’re thinking about the right problem.

Here’s the harder question: what are you actually measuring?

Most teams doing AI brand visibility tracking today have taken a familiar mental model and applied it to an unfamiliar system. Prompts have become the new keywords. Visibility scores have become the new rankings. Tracking platforms have emerged to show how often your brand appears in AI responses over time. On the surface, it looks like a natural evolution of the work you’ve already been doing.

It isn’t.

The tools built for traditional search were designed for a deterministic system, one where the same query reliably returns the same results. Large language models (LLMs) don’t work that way. They’re probabilistic: the same prompt can produce a range of valid answers, shaped by phrasing, context, model version, and more. Applying rank-tracking logic to a system that doesn’t produce ranks is the core mismatch, and it’s quietly corrupting the data most teams are reporting on.

This post breaks down exactly what’s going wrong and what a better approach looks like. It’s the first in a three-part series on AI visibility measurement. Part two introduces a structured framework for building prompts that actually reflect how your buyers use AI. Part three covers what the resulting data reveals about your content strategy.

The Tool The Industry Reached For (and Why It Doesn’t Fit)

The industry’s current approach to AI visibility measurement wasn’t irrational. It was fast. When a new channel emerges, teams reach for the tools and frameworks they already understand, and in digital marketing, that means rankings, share of voice, and tracked keywords. The logic was simple: prompts are the new search queries, so treat them the same way.

The problem is that search engines and LLMs are fundamentally different types of systems.

Traditional search is deterministic. Submit the same query to Google twice and you’ll get a broadly similar set of results. Position may shift slightly, but the system is stable enough that rank tracking works. That predictability is the entire foundation of AI keyword research and traditional SEO measurement.

LLMs are probabilistic. Run the same prompt multiple times and you’ll get a distribution of responses, not a fixed answer. The model generates each response based on statistical associations, not a retrievable index. There is no ‘rank one’ to hold.

The table below illustrates the mismatch. Applying rank-tracking logic to a probabilistic system doesn’t give you a less accurate version of the right answer. It gives you a fundamentally different kind of measurement entirely.

  Traditional Search LLM Ecosystem
System Type Deterministic Probabilistic
Behavior Predictable / Stable Variable / Generative
Core Metric Rank (Position) Presence (Likelihood)
Same query = same result? Broadly yes Not necessarily

This isn’t a minor calibration issue. It’s structural. If you’re reporting on AI visibility using methods designed for predictable, stable systems, you’re building strategy on a foundation that doesn’t reflect how LLMs actually work.

The User Who Doesn’t Exist

The second flaw in current AI visibility tracking is less obvious but equally important.

Most prompt tracking today relies on generic, decontextualized inputs:

  1. ‘Best CRM in 2026’
  2. ‘Top accounting software’
  3. ‘Best project management tool for small teams’

These prompts are clean, scalable, and easy to standardize. They look exactly like the keywords we’ve always tracked.

They also don’t resemble how real people use AI tools.

Real users carry context. They have prior conversations, professional constraints, specific goals, and levels of knowledge that shape what they’re actually asking. A prompt like ‘Best CRM in 2026’ represents an abstract, anonymous user with no history, no constraints, and no intent beyond the words in the query.

A graphic breaking down the differences between abstract users and how actual users use LLMs.

So when you measure AI visibility using these prompts, you’re measuring how the model responds to a hypothetical person who rarely shows up in real decision-making moments. That’s directionally useful at best.

Real audit work bears this out. In one analysis, a brand showed strong visibility for broad category queries, the kind that show up well in standard tracking. But when prompts were shaped around the specific contexts their buyers actually operate in, visibility dropped to zero in the topics most directly connected to purchase decisions. The tracking looked healthy. The actual picture wasn’t.

Generic prompts measure AI visibility for a user who rarely exists. If you want to know how your brand appears to real buyers, you need inputs that reflect real buyer contexts.

The Scaling Trap

The instinctive response to ‘generic prompts aren’t representative’ is volume. If one prompt isn’t enough, run a thousand variations. Add synonyms, modifiers, intent signals, geographic qualifiers. Cover the space more thoroughly.

This logic leads directly into what we call the scaling trap.

Every topic branches into multiple phrasings, intents, personas, and contextual modifiers. The number of prompts required to meaningfully approximate reality grows exponentially. A topic with five main phrasings, three intent signals, and four persona types generates 60 prompt combinations before you’ve added geographic variation or industry context. Scale that across a full content strategy and you’re looking at tens of thousands of prompts, run repeatedly, across multiple models, on a recurring basis.

A graphic explaining the volume fallacy and how prompts properly reflect reality.

Two problems follow. The first is practical: the cost of running this at scale is significant, and it compounds across every client account and every reporting cycle. The second is more fundamental: even after all of that, there’s no guarantee the resulting dataset is meaningfully more representative of actual user behavior. You’ve scaled the volume without fixing the flaw in the input logic.

More prompts don’t fix a representativeness problem. They just make the flawed measurement more expensive.

What Good Measurement Actually Requires

If the problem is that prompts lack context, and brute-force volume doesn’t solve that, the answer is to improve the quality of the input rather than the quantity.

Good measurement of a probabilistic system requires asking a different question entirely. The old question was: ‘Where do we rank?’ The right question is: ‘How reliably does our brand appear when the conditions that actually matter are present?’

That shift has real implications. A brand that appears 85 percent of the time when the right persona and intent conditions are met has a genuinely strong position, even if its average visibility across generic prompts looks modest. A brand that appears 50 percent of the time on generic queries but near zero percent in high-intent, decision-stage contexts has a problem that average tracking completely obscures.

Visibility, measured correctly, is a probability distribution across specific user contexts, not a single score. Getting to that measurement requires inputs that reflect those contexts: structured prompts built around real user personas, specific intent stages, and the actual questions buyers ask when they’re close to a decision.

That’s the foundation of a better approach to AI visibility measurement. The next post in this series walks through exactly how to build it.

Image related to AI Brand Visibility: You’re Tracking It Wrong

In the next post, I’ll walk through the framework we use at NP Digital to build prompts that reflect how real buyers actually engage with AI and what the data looks like when you do it right.

Why This Matters Now

AI-driven search has moved from a future consideration to a present reality, faster than most marketing teams anticipated.

ChatGPT now has over 700 million users, with exponential growth going on. That’s not a niche research tool. That’s a primary discovery channel for a significant and growing share of your buyers.

Image related to AI Brand Visibility: You’re Tracking It Wrong

Google AI Overviews now appear on roughly 48 percent of tracked queries, up 58 percent year over year according to BrightEdge data. In B2B technology, that figure reaches 82 percent of queries. If your buyers research software, services, or professional categories, AI is already shaping what they find before they ever reach your site.

The competitive dynamics are shifting accordingly. Brands that appear consistently in AI responses for the right queries, at the right intent stages, are building an advantage that compounds over time. Brands that don’t appear, or that appear for the wrong queries, are losing ground in the consideration phase before a sales conversation ever starts.

Every week you’re tracking AI visibility with flawed inputs is a week you’re making content and strategy decisions based on data that doesn’t reflect how your buyers actually use AI. The window to get ahead of this is open now.

FAQs

Why should I track AI brand visibility?

Your buyers are already using AI tools to research options, compare solutions, and form opinions about your category. Tracking AI brand visibility tells you whether your brand is present in those moments or invisible. Unlike traditional search, where a low ranking is visible and actionable, AI invisibility is silent, so you won’t know it’s happening unless you measure it.

What Is AI visibility?

AI visibility refers to how often and how favorably your brand appears in responses generated by AI tools like ChatGPT, Perplexity, Google Gemini, and Google AI Overviews. Strong AI visibility means your brand is being surfaced when users ask questions relevant to your product or service.

What are the top AI visibility solutions?

The most widely used platforms for tracking visibility include Writesonic and Profound, alongside a growing number of specialist tools. Each uses a defined prompt set to measure how often your brand appears across major AI platforms. The quality of your prompt set determines the quality of what you can learn — which is exactly the problem I want to address with this series.

Conclusion

Marketers aren’t doing something foolish by tracking AI visibility. They’re doing something natural: applying the tools and mental models they already know to a new channel. The problem is that those tools were built for a deterministic world, and LLMs don’t operate that way.

The mismatch matters. It means the data most teams are reporting on is structurally limited, not wrong exactly, but not representative of what’s actually happening when your buyers use AI to research your category.

The fix starts with a different question. Stop asking where you rank. Start asking how reliably you appear when it actually matters.

In the next post in this series, I walk through a framework built specifically for that question. This is a structured approach to prompt construction that reflects real buyer contexts and makes probabilistic measurement genuinely useful.

Read more at Read More