Headline formats and Google Discover: What 3.4 million articles reveal

Google Discover headline formats

You’ve probably seen some version of these three claims:

  • Quote-led headlines outperform plain declarative ones by nearly 29%.
  • Question headlines underperform both, sometimes by 24%.
  • Format drives the result: Rewrite a statement as a quote, or add that magic word, and you should expect a real lift.

We tested all three against 1,674,518 English editorial articles and 1,690,295 French articles from the 1492.vision Discover corpus (November 2025 to May 2026): about 3.4 million editorial articles with at least one capture across our fleet.

They share a deeper flaw than any of their numbers.

All three treat headline format as a cause — a lever you pull to gain visibility. But the data shows, layer after layer, that a format’s measured effect is almost entirely a proxy for something else: which publisher used it, for which audience, and on which Discover surface.

The headline is a symptom of those choices, not an independent driver.

The clearest demonstration is Simpson’s paradox. Once you see it, you find it throughout the dataset.

A note on what we measure

Our metric isn’t clicks from Discover; no third party has that data. It’s hits per article: how often an article appears across the 1492.vision fleet we observe, a proxy for visibility.

The corpus is limited to editorial articles. YouTube and X are excluded because their headlines follow different conventions. We’ll return to both at the end—they sharpen the point more than anything else.

A word on why the volume matters: the entire argument depends on being able to slice 3.4 million articles by publisher, Discover surface, topic, and language while still retaining enough data in each segment for meaningful comparisons. That’s the difference between a number and an insight — and between a real format effect and a statistical mirage.

The number is real, at the wrong altitude

Pool all publishers together, and a clean gradient emerges: quote-led headlines at the top, statements at the bottom.

Lang Format Articles Mean hits Median vs statement
EN Quote-led 38,044 13.0 4 +37%
EN Quote inside 75,463 11.5 4 +21%
EN Question 53,081 10.2 4 +7%
EN Statement 1,674,518 9.5 3 baseline
FR Quote-led 179,472 52.8 13 +48%
FR Quote inside 223,052 49.9 12 +40%
FR Question 103,117 41.3 11 +16%
FR Statement 1,690,295 35.7 9 baseline

The commonly cited +29% is conservative for pure editorial articles: quote-led headlines show a +37% lift in English and +48% in French. Questions, far from underperforming, also outperform statements (+7% EN, +16% FR).

At this level of aggregation, claim 1 looks understated and claim 2 looks plainly wrong.

This is the level of aggregation where most headline advice is born. Hold onto that +37% figure — the rest of this piece is about what it’s actually measuring.

Hidden variable 1: which publisher

The aggregate can’t answer a crucial objection on its own: the publishers that use quotes aren’t the same publishers that don’t.

Celebrity media, regional dailies, and buzz-driven sites lean heavily on quotes and earn more Discover hits per article regardless of headline format. Pure-play publishers, wire services, and utility-focused sites favor declarative headlines and tend to sit lower.

The raw comparison, then, isn’t quote versus statement. It’s one publisher population versus another.

This is a textbook Simpson’s paradox: a strong trend in the aggregate that weakens, disappears, or reverses once you segment by group.

To get anywhere near the effect of headline format itself, the grouping variable has to be the publisher.

So make each publisher its own baseline: compare quote versus statement within the same site, holding audience and topic mix constant.

Across 324 English and 439 French publishers with enough of both formats — at least 50 quote and 200 statement articles each:

Lang Publishers Quote wins (median site) Quote wins (mean site) Median within-publisher Δ
EN 324 31.5% 55.9% +3.1%
FR 439 47.6% 57.4% +5.5%

In English, statements outperform quotes at 68% of publishers by the median; quote-led headlines hurt more often than they help. In French, the result is close to a coin flip.

That leaves the underlying format effect at roughly +3% to +5%—about five to nine times smaller than the aggregate figure.

(The mean is higher than the median because a minority of publishers see large gains from quotes. The median is the more reliable measure of the typical publisher.)

Stop here and the lesson sounds like “segment your data.” But the collapse points to something larger.

If three-quarters of a +37% effect was really a publisher effect, the obvious next question is: what else is the headline metric standing in for?

The rest of this article is a tour of those hidden variables. And by this point, the answer to claim 3 is already coming into view: the format itself isn’t the driver.

The same substitution, in reverse: questions

The conventional advice says questions underperform by roughly 24%. The aggregate view of our data says the opposite: questions outperform statements (+7% EN, +16% FR).

Both conclusions are wrong for the same reason. Question headlines are disproportionately used by high-engagement publishers, which inflates their aggregate performance.

Within publishers, the picture settles.

In English, question headlines show a modest real underperformance (-3.7%), winning at only 29.3% of sites. In French, the effect is essentially neutral (-0.5%), with questions outperforming at 46.2% of sites.

The conventional advice gets the direction roughly right in English and neutral in French, but its usual magnitude is about sixfold too large.

The question mark isn’t the cause. The kind of publisher using it is. Same hidden variable, opposite sign.

The effect won’t even hold still

Even that modest within-publisher effect drifts from month to month.

In English, it peaks at +2.5% and turns negative in March 2026, while statements outperform questions at 55% to 60% of sites each month. In French, it ranges from +3% to +12% — strongest in December and February, weakest in March — with no clear trend.

A genuine causal lever shouldn’t wobble like this. A correlation tied to a shifting content mix should.

Hidden variable 2: Which audience

The +3-5% average hides a sharp, consistent split. In English:

  • Gainers: International general news (BBC +85%, Forbes +46%, CBS News +43%, Boston Globe), Yahoo aggregators, mass-market magazines (Parade, Good Housekeeping), Gizmodo.
  • Losers: Specialist sport (RugbyPass, Planet F1, ThisIsAnfield), entertainment (IMDb, TVInsider, People), and factual-leaning dailies (Standard, Washington Post).
Top FR publishers, quote vs statement

French data follows the same pattern in a different market.

  • Gainers: Regional newspapers (La Dépêche, La Montagne, L’Écho Républicain) and general-interest magazines (Grazia).
  • Losers: Specialist sports outlets (Foot National, le10sport, MadeInFoot), technology publishers (Les Numériques), and service-oriented titles (Journal des Femmes, Femme Actuelle).

The pattern is editorial, not algorithmic. Quotes tend to work where the audience comes for commentary, reaction, and framing, and fail where the audience comes for facts.

A publisher built around “what someone said” benefits from a quoted headline. One built around “what just happened” usually doesn’t.

The convergence between English and French is the giveaway. This isn’t a language effect; it’s a reader-intent effect.

What looks like a headline-format effect is, in this case, an audience effect wearing the clothes of a headline.

Hidden variable 3: Which Discover surface

Discover isn’t a single feed. It’s a collection of pipelines, each selecting articles in different ways:

  • Editorial curation (moonstone, mustntmiss).
  • The main topic-personalization engine (aura).
  • Related-reading context (paginationpanoptic, content).
  • Similarity-based recommendation (relatedcontentruby, userpersonascontent).

First, rule out the obvious alternative explanation. Are quote-led articles simply being routed to higher-value Discover surfaces, making the apparent bonus a placement effect rather than a headline effect?

The data says no.

Comparing where quote and statement articles actually appear, the distributions are nearly identical. In English, the largest differences are small: content.f (+2.2 percentage points), aura.f (-1.9), and moonstone.f (+0.6).

Pipeline mix by format

The bonus isn’t about placement: quotes and statements appear on the same surfaces in the same proportions. It’s about intensity — how each format performs once it’s on a surface. There, the overall +3% to 5% breaks into a wide range: from +22% to -14% in EN and from +25% to -12% in FR.

Quote bonus by pipeline, EN, full picture

Grouped into functional families, the pattern is readable:

Pipeline family EN FR
Editorial curation (moonstone, mustntmiss, astria, news…) +3.4% +9.7%
Related reading / context (paginationpanoptic, content…) +2.0% +6.7%
Trends / freshness (deeptrends, freshvideos…) +4.4% +2.3%
Main personalization (aura) +0.6% +1.8%
Similarity-based recommendation (relatedcontentruby, userpersonas…) -1.6% -1.9%

Quote-led headlines win where multiple headlines compete for attention at once — curation carousels, news clusters, and other surfaces where the title carries a social signal: someone said this. They lose on similarity-based recommendations, where the surface sells continuity (“because you read X, you’ll read Y”) and a quote disrupts the topic-clear promise with an out-of-context citation.

The largest pipeline by volume, Aura, ranks on topic affinity and barely reacts to format at all, with gains of just +0.6% to +1.8%.

Why is the net effect so small?

A single quote-led FR article doesn’t get one number; it gets a blend:

  • +10 to +25% on its curation share (moonstone, mustntmiss, astria)
  • ~0% on its aura share, the largest slice of volume
  • -3% on its relatedcontentruby share (≈ 10% of captures)
  • -2 to -6% on shopping/viewer-related surfaces

Integrate those and you land at +4% to +7% net. The curatorial gains are real but partly offset by recommendation losses, which is why the aggregate is nowhere near +29%. The same format is both an asset and a liability, depending entirely on the surface serving it.

And +4–7% overstates how much the format itself matters because each pipeline’s ranking is a compound of signals unrelated to the title: engagement, scroll depth, topic affinity, E-E-A-T, entities, reading history, location, timing, and prior interactions.

A quote in the headline is, at best, one weak signal competing with all of those. Long before an article reaches a feed, it’s largely swamped by everything else.

Questions by pipeline, same story sharper

Question vs statement bonus, by pipeline

These are within-publisher medians (each publisher against itself), so they aren’t a crude artifact of FR using more questions. The format follows the same pipeline logic as quotes, but in a more polarized form:

  • FR curation leans positive on questions; EN curation leans negative. astria.f, the same pipeline in both languages, runs +9% in FR and -1% in EN; FR mustntmiss.f is +14%, EN moonstone.f is -13%.
  • Similarity-based recommendation penalizes questions everywhere, harder than quotes: relatedcontentruby.f FR -11.5% (306 publishers), EN -6.1% (119); itemitemcollaborativefiltering.f FR -14.5%.
  • aura stays neutral in both (+3.5% FR, -0.6% EN).

Two caveats point in the same direction:

  • A fleet-capture metric can’t distinguish an algorithmic penalty from an audience-eviction effect: readers see a question mark, decide “not now,” and scroll past. The fact that relatedcontentruby — which serves already-engaged readers — penalizes questions this heavily points to a behavioral signal, not just ranking.
  • Within-publisher pairing controls for each publisher against itself, but the median is still computed across a different set of publishers in FR and EN, on partly different surfaces. So “FR rewards questions, EN doesn’t” describes the publishers and topics occupying each cell, not an inherent property of the language or the question mark. It’s another hidden variable mistaken for a format effect.

Hidden variable 4: Which editor, and which judgment

Even the honest +3% to 5% comes with a caveat that outweighs its size. When a publisher writes a headline as a quote, they choose the best available quote for that story. So the within-publisher figure compares the best quote an editor selected with the average of all that publisher’s statements, not the same article written two ways.

It’s the subject-line A/B testing problem: a good alternative beats a bad one, but the average alternative doesn’t. Convert every headline to quote-led and you’d be writing average quotes, so most of the gain would disappear. The +3–5% is an upper bound on a selective practice, not the return from a blanket rule.

That’s the final reason “do it everywhere” fails:

  • Not every article has a quote. A sports result, a press release, a market analysis, a product test: forcing one means fabricating it.
  • The editor-selection bias above: The measured bonus is the best quote chosen, not a property of the format.
  • Recommendation pipelines are long-tail levers. relatedcontentruby and friends are how an article redeploys after its initial peak, the main mechanism for extending Discover lifetime. Optimizing the headline for the curation peak while breaking the promise on these surfaces can net negative.
  • The largest pipeline barely reacts. aura is 11% to 15% of FR captures and 7% to 9% of EN, with a +0.6% to 1.8% quote effect. A universal quote rule optimizes secondary surfaces while ignoring that the biggest one runs on topic affinity.

The clincher: the same format, opposite meaning

YouTube and x.com, quote bonus

We excluded YouTube and X from the main corpus, but their results are the clearest proof of the thesis. The same quote-led format produces opposite effects depending entirely on what the title is trying to do.

Domain Lang Quote articles Statement Mean hits quote Mean hits stmt Δ
YouTube EN 43,476 734,986 11.6 10.2 +14%
YouTube FR 16,509 93,912 59.0 29.1 +103%
x.com EN 34,156 268,175 5.2 4.9 +6%
x.com FR 32,201 114,914 21.4 24.6 -13%

On YouTube, the title is effectively a text thumbnail that has seconds to create curiosity. A quote serves as a content promise — “here’s the line worth hearing” — which helps explain the +103% result in French. On X, the title is the post itself, and a detected quote usually indicates that someone is repeating or responding to another person’s words, diluting the original message. That correlates with a -13% result.

Same characters. Same regex. Opposite outcome. The format didn’t change; the job it was doing did.

(Methodological footnote: a naive audit that folded YouTube into the editorial corpus would inflate the overall quote bonus by 20–30 points, while one that folded in X would dilute it. Any serious headline study has to isolate editorial articles before measuring headline effects.)

The headline was never the variable

Put the layers together. Three-quarters of the +37% raw bonus was explained by publisher differences. What remained split again by audience, then by Discover surface, then by which quote the editor selected, and finally reversed entirely when the title served a different function on another platform. At every step, removing context shrank or flipped the apparent format effect.

There’s no clean residue at the bottom where the headline acts independently. The effect is inseparable from the context that creates it.

That’s not a measurement failure; it’s the finding. We just saw the mechanism. Headline format is one weak signal among many stronger ones, all moving through pipelines that often pull in opposite directions.

The consequence is the point. An article’s visibility is the running score of that entire contest, not the verdict of any headline rule. A number measured across publishers is downstream of everything that travels with the format: who published it, what topic it covers, what the audience expects, the newsroom’s style and habits, and the conventions of the language itself.

So when an aggregate reports “+29% for quotes,” it isn’t isolating the quotation marks. It’s measuring a correlation with that whole bundle of factors and quietly relabeling it as causation.

None of this means aggregate data is the enemy. Everything above comes from aggregate data, just analyzed at the right level.

The trap is narrower: treating a single cosmetic variable, averaged across publishers that don’t belong in the same category, as a causal lever.

The same index that exposes that mistake also reveals the signals that genuinely drive Discover: which topics a publisher wins on, which entities are accelerating, who dominates a given surface, and what’s trending before it peaks. Those signals aren’t cosmetic, and they aren’t drowned out by stronger forces. They’re the underlying demand that headline format only weakly approximates.

The lesson isn’t “ignore the data.” It’s “stop averaging the wrong variable across the wrong population.”

This is why no cross-publisher average, corrected or not, converts into a rule for your site:

  • Visibility isn’t traffic. Two sites can earn identical Discover visibility on the same article and see very different CTRs because their audiences click for different reasons.
  • No two audiences are the same. A quote that reads as insider commentary to a magazine reader may read as vague or irrelevant to someone scanning sports scores.
  • A cross-publisher average of one cosmetic feature is the average of audiences you don’t have. Segment by your audience, your topics, and your surfaces, and it becomes information again.

The only test that answers your question is the one you run on your own site, with your own audience. Know who you’re writing for, then measure them. Slice the data by your audience, your topics, and your surfaces — not by a single number averaged across everyone.

So what about the three claims?

Each is real as a correlation and useless as a cause:

  • “Quotes beat statements by ~29%”: True in aggregate — larger than +29%, in fact — but mostly explained by publisher differences. At the publisher level, the residue is +3% to 5%, and even that compares the best quote an editor selected against the average of all statements, not the format itself.
  • “Questions underperform”: Directionally true in EN, neutral in FR, but the magnitude is about 6x too large. The actual effect is roughly -4% in EN and ~0% in FR.
  • “The format itself is the driver”: The claim the dataset refutes. The same article from the same publisher, mechanically rewritten as a quote, would not gain the aggregate effect.

The honest version, if you want one sentence to keep:

A quote-led headline can earn roughly +3% to 7% additional Discover visibility for audiences that value commentary and framing (general news, magazines, regional press), especially on curation surfaces, and lose for factual audiences (sports, tech, utility) and on similarity-based recommendation surfaces. There is no universal gain from quotation marks; the popular ~+29% figure overstates the format effect by roughly an order of magnitude. The useful question isn’t “Should I use a quote?” but “Who am I writing for, and which Discover surface drives my traffic?” The only place to answer that is with your own site, not anyone else’s average.

Methodology

  • Data and period: 1,674,518 EN and 1,690,295 FR editorial articles with Discover visibility from 1492.vision proprietary data, collected between 2025-11-01 and 2026-05-19. Editorial articles only; excludes ads, videos, AI Overviews, and showcases. Domain exclusions: x.com, twitter.com, m.twitter.com, youtube.com, www.youtube.com, and m.youtube.com (reported separately above).
  • Headline format detection (regex): Quote-led: title starts with a multi-word quoted phrase (“…”, «…», ‘…’, or ‘X…’:). Quote inside: a quoted phrase appears but not at the start. Question: ends with ?. Statement: everything else. Titles under 20 or over 300 characters are excluded. Detection deliberately errs toward false negatives in the quote bucket, biasing against finding a quote effect, so the +3–5% is conservative.
  • Three layers of analysis: (1) Raw aggregate: all publishers pooled, producing +37% / +48%. (2) Within-publisher: quote vs. statement inside each publisher with ≥50 quote and ≥200 statement articles; we report the share of publishers favoring quotes and the median per-publisher Δ. This neutralizes publisher-mix bias. (3) Monthly evolution: the same pairing, recomputed monthly with relaxed thresholds (≥10 quote, ≥40 statement).
  • Pipeline layer: Captures come from 1492.vision proprietary data, with each row representing one capture on a specific pipeline. For each (pipeline, format, publisher), captures per article = pipeline captures ÷ distinct articles. Within-publisher pairing includes publishers with ≥20 quote (or question) and ≥60 statement articles on that pipeline. A pipeline is shown only if ≥5 publishers qualify. Pipeline families are an empirical grouping (editorial curation, related reading, trends, similarity-based recommendation, and main personalization) that reflects how each surface behaves.
  • Metric: A “hit” is one capture of an article on Discover by the 1492.vision device fleet. It is a visibility proxy, not a visit.
  • Known limitations: (1) No traffic data: the metric is Discover visibility, not clicks, so a format could affect CTR independently without appearing here. (2) Regex detection misses edge cases and is biased toward under-counting quotes. (3) Within-publisher effects compare the best quote an editor selected against the average statement, not the counterfactual of making every headline quote-led. (4) Some negative pipelines have small publisher samples (<10); the consistent direction matters more than any individual magnitude.

Read more at Read More

Why TikTok Is Expanding Its Premium Ads Push and What That Means for You

Key Takeaways

  1. TikTok launched four new or expanded premium ad formats at its 2026 Newfronts: Logo Takeover, Prime Time, TopReach, and expanded Pulse offerings.
  2. More than 200 million Americans are on TikTok, and the platform reaches 1.99 billion monthly active users globally.
  3. Early results on Logo Takeover showed double-digit lifts in brand awareness and purchase intent.
  4. TikTok’s engagement rate of 3.7 percent is nearly eight times higher than Instagram and twenty-five times higher than Facebook.
  5. The platform is positioning itself as a full-funnel engine, with commerce and lower-funnel capabilities maturing alongside its reach.
  6. TikTok-native creative authenticity remains essential, even within premium placements.

TikTok-native creative authenticity remains essential, even within premium placements TikTok’s 2026 IAB NewFronts presentation made one thing clear: the platform is no longer asking brands to treat it as a social experiment. It is asking for a seat at the table alongside TV and streaming budgets, and the new ad products it unveiled give it a credible case to make.

If you are still running TikTok as an afterthought in your media mix, it is time to reassess.

The New Formats, Explained

TikTok’s NewFronts announcement introduced a set of formats specifically designed to capture premium brand investment.

Logo Takeover places your brand at the moment users open the app, before anything else on the screen competes for attention. It is co-branded with TikTok itself, which carries an implicit credibility signal alongside the raw reach. Early tests showed meaningful lifts in both awareness and purchase intent, giving advertisers an actual benchmark to work from rather than just a pitch.

The logo takeover format.

Source

Prime Time is a sequential format that delivers up to three ads from the same brand to the same user within a 15-minute window, timed to high-engagement periods or major cultural moments. The ability to tell a continuous story across multiple exposures in a short window has historically been a TV strength. TikTok is bringing that capability to a mobile-first, creator-driven environment.

TopReach combines two existing high-visibility placements into a single buy: the first ad users see when opening the app, and the first in-feed ad in the For You feed. For brands running a major launch or trying to dominate a cultural moment, maximizing unique daily reach through a single purchase is a genuine efficiency gain.

The Top Reach format.

Source

The expanded Pulse offerings include Pulse Mentions, which places brands adjacent to conversations already happening about their category, and Pulse Tastemakers, which lets brands align their ads with specific creator communities. Both formats lean into what TikTok does better than any other platform: making ads feel like they belong inside the content experience rather than interrupting it.

Pulse Mentions.

Source

TikTok Has Grown Past Its Early Reputation

There is still a version of TikTok in many marketing budgets that looks like a niche social channel with unpredictable ROI. That picture is outdated.

The numbers tell a different story. TikTok generated $33.1 billion in global advertising revenue in 2025, a 43 percent increase from the year before. Its engagement rate of 3.7 percent sits well above every major social competitor. More than half of TikTok users have purchased from brands after seeing their products featured on the platform. TikTok Shop generated $15.82 billion in U.S. sales in 2025, growing at 108 percent year over year.

Only 26 percent of marketers currently run TikTok campaigns. For brands not yet on the platform in a serious way, that gap is the opportunity.

Commerce capabilities have matured to the point where lower-funnel performance is genuinely measurable. Creator-led storytelling has proven to drive purchase behavior in ways that traditional video placements often cannot. And now, with premium formats designed to deliver the kind of reach and sequential storytelling that TV has historically owned, TikTok is a legitimate alternative for budgets flowing toward linear and streaming video.

The brands that shifted budget toward digital video early, before it was obvious, built advantages that took competitors years to close. The same opportunity exists here.

Why Cost Efficiency Matters

Beyond reach and engagement, the cost structure of TikTok advertising makes it worth serious consideration. TikTok ads average a CPM of around $9, compared to Meta’s average Facebook CPM of roughly $15. That cost advantage combined with the platform’s higher engagement rate means dollars spent on TikTok tend to produce more interaction per dollar than on competing platforms.

TikTok vs Meta vs Google comparison.

Source

That advantage will not last forever. As more advertisers move budget onto the platform, auction competition will increase and CPMs will rise. The brands that establish their TikTok presence and learn what works now will be building that knowledge at a lower cost than those who wait.

How to Approach This

The most common TikTok mistake is importing creative from other channels. A CTV spot or a YouTube pre-roll that performs well will not automatically translate. TikTok rewards content that feels like it was made for the platform and the moment. Even within premium placements, the native feel of the content matters.

Research backs this up. Spark Ads deliver 34 percent higher conversions than standard in-feed ads. The best-performing brand content on TikTok does not look like advertising. It looks like something a person would make and share. Getting that balance right, particularly within premium, high-production formats, is the creative challenge.

That does not mean sacrificing production quality. The new format are built for exactly the intersection of high production value and platform-native storytelling. Getting both right is the challenge, and it requires thinking about creative from a TikTok-first perspective rather than adapting assets designed for other channels.

A few practical steps worth taking now:

  • Test Logo Takeover and TopReach early, while competition for the placements is lower and cost benchmarks are more favorable.
  • Revisit your media mix model. If TikTok is still sitting in a social budget silo, it may be underweighted relative to what it can deliver against video and streaming objectives.
  • Align your paid social and commerce teams. TikTok’s lower-funnel capabilities only deliver their full value when both sides of the house are working toward the same goals with the same data.
  • Pay attention to creator selection. Pulse Tastemakers gives you the ability to align placements with specific creators. Treat that as a targeting decision, not a creative one. The right creator community for your brand will outperform a broad placement every time.

FAQs

How is TikTok’s ad audience different from other platforms?

TikTok reaches 1.99 billion monthly active users globally, with the 25 to 34 age group now its largest single cohort at 40 percent of users. The audience is maturing, meaning the perception that TikTok skews very young is increasingly outdated. The platform also sees daily active users return an average of five to fifteen times per day, making frequency of exposure higher than most other social channels.

What makes TikTok advertising different from Meta or YouTube?

The key difference is how ads fit into the platform experience. TikTok’s ad formats, at their best, look and feel like the content people are already watching. This native quality drives higher engagement and, in many cases, better conversion performance. The platform’s algorithm also rewards content quality over account size, which means strong creative can reach audiences far beyond your existing follower base.

Is TikTok Shop worth investing in alongside paid ads?

Yes. With $15.82 billion in U.S. sales in 2025 and 108 percent year-over-year growth, TikTok Shop has crossed the threshold from experiment to serious commerce channel. Research shows that 25 percent of users who bought from TikTok Shop found the item through a TikTok ad. Paid media and shop strategy work best when they are planned together.

What budget should I start with on the new premium formats?

There is no universal answer, but the general principle applies: treat initial spend on new formats as learning investment rather than expecting immediate ROAS. Get in early while competition is lower, build benchmarks, and scale from a position of knowledge rather than guesswork.

Conclusion

TikTok is not pitching itself as a social media platform with ad inventory, but a full-funnel engine where entertainment, commerce, and performance meet. The numbers back that up: global ad revenue growing at 43 percent year over year, engagement rates eight times higher than Instagram, and a commerce operation that grew by more than 100 percent in a single year.

The brands that take that seriously now and build creative and budget strategies to match will be harder to catch as the platform continues to mature. The window for establishing a cost-efficient early presence is still open. It will not stay that way indefinitely.

Read more at Read More

New: Yoast releases performance optimizations for larger websites

With our latest 27.8 release, we introduced performance optimizations that should reduce loading times throughout the plugin’s functionalities, especially noticeable in large sites with lots of posts and users.

Note: This post contains technical content and implementation details.

Offering well-tuned software with minimal overhead in servers and fast loading times is always at the forefront of everything Yoast developers do. However, Yoast SEO is installed in millions of websites so the variance of setups that we must be well-tuned for is big. This means we should be continuously going back to search for windows in optimizing the performance of the plugin. We’ve been known to do that consistently in the past, like when we improved our database system.

The 27.8 release is the outcome of one of those targeted reviews. We deliberately picked features whose behavior at scale offered the most headroom and reworked them to be leaner and faster. From modifying queries to make pages faster for sites with many users and shaving heavy operations in the admin for sites with many posts, to reducing rounds trips to the database for multiple features and generally applying performance best practices, this is a release meant to improve the user and developer experience in the Yoast SEO plugin.

We would also like to offer a technical summary of the improvements in this release here, focusing on their nitty-gritty details because it’s always nice to raise awareness about performance best practice (not to mention that it’s always fun to talk about code).

Significantly reduce loading times of the root sitemap on sites with many users

For context, for Yoast SEO to calculate the Last Modified value of the author sitemap, when it outputs the root sitemap, it uses the usermeta of the all the users that are eligible to be included in the author sitemap.

Calculating the eligible users was traditionally done by checking user capabilities. This was done by adding the ‘capability’ => [ ‘edit_posts’ ] argument in the get_users() call that was used. As a result, a very heavy query with multiple joins and no use of the indexes of the database was triggered.

Specifically, the resulting query added a clause like this:

AND ((((mt1.meta_key = 'wp_capabilities'
        AND mt1.meta_value LIKE '%"edit\_posts"%')
       OR (mt1.meta_key = 'wp_capabilities'
           AND mt1.meta_value LIKE '%"administrator"%')
       OR (mt1.meta_key = 'wp_capabilities'
           AND mt1.meta_value LIKE '%"editor"%')
       OR (mt1.meta_key = 'wp_capabilities'
           AND mt1.meta_value LIKE '%"author"%')
       OR (mt1.meta_key = 'wp_capabilities'
           AND mt1.meta_value LIKE '%"contributor"%')
       OR (mt1.meta_key = 'wp_capabilities'
           AND mt1.meta_value LIKE '%"wpseo\_manager"%')
       OR (mt1.meta_key = 'wp_capabilities'
           AND mt1.meta_value LIKE '%"wpseo\_editor"%'))))

Since LIKE ‘%…%’ cannot use any B-tree index, MySQL must read each matching wp_capabilities row in full and do seven substring scans of the serialized PHP meta_value per row.

By modifying that calculation from using the capability check to looking for users with published posts (via using the ‘ has_published_posts ‘ => true argument), we instantly turned the resulting query to be one that uses indexes and that performs way better in sites with many users.

In fact, on one of our tests, on a site with around 2 million users, the time it took to complete each query (so approximately the time that took the root sitemap to render), went from over 300 seconds down to just 25 milliseconds! This means that the change has the potential for drastic improvements in loading times of root sitemaps in similar sites.

Finally, considering that the ‘ has_published_posts ‘ => true argument was already used in a later stage of the sitemap generation, the change itself should have little to no negative impact on the actual functionality of the feature.

Reduce loading times of the author sitemap on sites with many users

For Yoast SEO to render author sitemaps, it needs to calculate the eligible users. On sites with many users, this can be a very heavy operation. Aside from the above optimization, we noticed that while Yoast SEO was calculating eligible users, it also added a meta query to check whether the user_level of each user was over 0.

It turned out that this was a remnant from old times, because the user_level framework had been deprecated by WP core since version 3.0. While this didn’t break things in our sitemap feature, it unnecessarily added an INNER JOIN in the resulting query without much purpose and in sites with very big user and usermeta tables that was degrading performance. So we went and removed the unnecessary JOIN:

INNER JOIN wp_usermeta AS mt1 ON wp_users.ID = mt1.user_id 

... 

AND ( mt1.meta_key = 'wp_user_level' AND mt1.meta_value != '0' )

Since the user_level framework was deprecated a long time ago, we made the deliberate call to drop support for it, especially since doing so would make our feature smoother. In fact, we are comfortable shipping this optimization and expect minimal disruption as a result, exactly because of how old that deprecation is.

Prevent unnecessary expensive database queries in admin pages

In order to timely notify admins that they need to perform the necessary actions for their site data to be indexed optimally in our internal storage, Yoast SEO used to run a database query daily while admins navigated throughout the backend. For big sites, that database query had the potential to run for several seconds, slowing the rendering of admin pages periodically.

Specifically, the following function:

Limited_Indexing_Action_Interface::get_limited_unindexed_count()

This can run complex queries like the ones below, which were running periodically on admin pages, slowing rendering on larger sites.  

SELECT Count(P.id) 
FROM   wp_posts AS P 
WHERE  P.post_type IN ( 'post', 'page' ) 
       AND P.post_status NOT IN ( 'auto-draft' ) 
       AND P.id NOT IN (SELECT I.object_id 
                        FROM   wp_yoast_indexable AS I 
                        WHERE  I.object_type = 'post' 
                               AND I.version = 2)

We managed to re-arrange the logic of the code responsible for the notification that told admins about pending actions in such a way that those heavy queries now run only once, at the moment it’s first detected that such a notification should be created.

That way, we effectively cache the results of the

Limited_Indexing_Action_Interface::get_limited_unindexed_count()

and rely on cache invalidation that existed before our changes, but weren’t properly utilized. As a result, a potentially very heavy database query went from being triggered daily (and, on very busy sites with lots of concurrent users, once per 15 minutes) to being triggered only once in most sites. 

Optimize expensive database queries in admin pages

Related to the above query-preventing change, not only did we manage to avoid running that aforementioned heavy database query more than once per site, but we also managed to optimize the query itself. An added benefit from that is that we made the SEO optimization tool much faster in sites with lots of posts.

Specifically, we went from:

AND P.ID NOT IN ( 
    SELECT I.object_id FROM wp_yoast_indexable AS I 
    WHERE I.object_type = 'post' 
)

To:

AND NOT EXISTS ( 
    SELECT 1 FROM wp_yoast_indexable AS I 
    WHERE I.object_id = P.ID 
      AND I.object_type = 'post' 
)

Since NOT IN (subquery) builds the entire list of object_ids, while the second query short-circuits the moment one row matches, the query runs considerable faster in sites with multiple thousands of posts.

Reduce roundtrips to the database

As a rule of thumb, roundtrips to the database are considered to be expensive operations that should be reduced to a minimum whenever possible. Our reviews discovered instances where we were retrieving data for multiple posts in sequential SELECT queries where we could have done a single batched SELECT query to gather data for all posts at once. 

For example, a piece of code that looked like this:

$indexables = []; 

foreach ( $post_ids as $post_id ) { 
	$indexables[] = $this->repository->find_by_id_and_type( (int) $post_id, 'post' ); 
}

was refactored into something that looked like this:

$ indexables = $this->repository->find_by_multiple_ids_and_type( 
	array_map( 'intval', $post_ids ), 
	'post', 
);

That meant that for a chunk of 1000 posts, instead of performing 1000 SELECT queries that yielded a maximum of one row, we now perform a single SELECT query that yields a maximum of 1000 rows. Naturally, we made sure that the posts that will be requested each time do not exceed a certain threshold, to avoid reaching MySQL usage limits.

As a result, sites with e.g. 1000 posts would save 960 roundtrips to the database for certain operations like part of their SEO optimization or part of the output of the schema aggregation feature.

Improve post editor performance by preventing unnecessary re-renders

The WordPress editor re-renders Yoast’s sidebar panels whenever the data they pull from the store appears to have changed. Unfortunately, “appears to have changed” is decided by reference equality (JavaScript’s ===) not by comparing values. A selector that returns { items: [‘foo’] } looks identical to a human, but if it’s a fresh object literal each time, React treats it as new and re-renders the panel. And if we multiply that by a busy editor that dispatches state updates on every keystroke, the result is panels that re-render constantly for no reason.

With the 27.8 release, we identified multiple instances where data that weren’t actually changed triggered unnecessary re-renders in the post editor and patched them, making our editor integration much more robust and performant.

The post New: Yoast releases performance optimizations for larger websites appeared first on Yoast.

Read more at Read More

AI-generated content & SEO 🤖 Everything you need to know in 2026

AI-generated content – Loved by some, feared by others… Ever since tools like ChatGPT, Gemini, Claude, and many other AI writers became part of everyday marketing, SEOs have been…

The post AI-generated content & SEO 🤖 Everything you need to know in 2026 appeared first on Mangools.

Read more at Read More

Query Fan-Out: What It Is and How It Affects AI Visibility

Your content can rank on the first page of Google and still never be cited or mentioned by LLMs.

This makes sense once you understand query fan-out, a background process AI systems use to build answers.

When someone asks ChatGPT or Perplexity a question, it doesn’t default to the best-ranking page.

Instead, it runs related searches behind the scenes, pulling from the most relevant and reliable sources, regardless of position.

User query

If your brand doesn’t show up in those searches (whether through your own content or third parties), you’re unlikely to make it into the answer.

High rankings don’t hurt, of course.

But in AI search, coverage and retrievability are king.

In this guide, I’ll teach you how to optimize your content strategy for query fan-out to help increase your AI visibility.

You’ll learn:

  • Why LLMs use query fan-out
  • How it behaves differently across major AI platforms
  • Why it changes how you create and structure content
  • A 6-step workflow for earning more citations in AI search

Free template: Our Query Fan-Out Audit Template includes ready-to-use spreadsheets for logging money prompts, sub-queries, and content gaps — plus a checklist to keep you on track. Download it now to follow along.


First, I’ll dive deeper into how query fan-out works.

What Is Query Fan-Out?

Query fan-out is a process AI search systems use to break a single user query into multiple sub-queries to create the most helpful response.

In other words, the AI “fans” the query out into a series of related sub-questions to build a more complete picture of the topic.

How query fan-out works

It then pulls information from multiple sources — editorial sites, Reddit threads, comparison and product pages — and synthesizes it into a single comprehensive answer.

The query fan-out process

AI systems use query fan-out for a few reasons:

  • Confirm information: A single source might be wrong or biased. Running parallel sub-queries allows the system to cross-reference multiple sources and find consensus before committing to an answer.
  • Handle complex, specific queries: When a question has multiple layers, like comparing two products across price, reliability, and long-term value, fan-out breaks it into manageable pieces that the system can research independently.
  • Answer the real question: Someone searching “best toothbrush” probably also wants to know about price, battery life, and durability, even if they didn’t say so. Fan-out anticipates those needs and gathers evidence upfront.

For example, a search for “best toothbrush” might trigger sub-queries like “best electric toothbrushes [year]” and “best toothbrushes for sensitive gums.”

This helps the AI build a more complete and useful answer:

Sub-Query What It Contributes to the AI Response
Best electric toothbrushes Top-rated picks and editorial consensus
Best toothbrushes for sensitive gums Use-case recommendations
Oral-B vs. Philips Sonicare Head-to-head comparison data
Best eco-friendly toothbrushes Value picks and pricing information

The AI then synthesizes those findings into a single answer that covers everything the user might want to know: top picks, price ranges, use-case breakdowns, and comparisons.

In this way, it anticipates the user’s needs, even though the original prompt (best toothbrush) was just two words.

ChatGPT – Best toothbrush

What Query Fan-Out Is NOT

Now that we’ve covered what query fan-out is, let’s clear up a few common misconceptions.

Query fan-out is not:

  • Keyword research: This is the process of finding terms your audience searches for. Query fan-out is something AI systems do automatically, behind the scenes, every time someone asks a question.
  • People Also Ask: PAA is a visible SERP feature that shows users what else they might want to search. Fan-out happens in the background whether you can see it or not.
  • A fixed set of queries: Only 27% of fan-out sub-queries remain consistent across repeated searches, according to a SurferSEO study. Sub-queries vary by phrasing, user context, and platform.

Why Query Fan-Out Matters for AI Visibility

Understanding what query fan-out is only gets you so far. The real question is: What does it mean for your content strategy?

Here are four shifts that should make you rethink how you approach content.

You Don’t Need Top Rankings to Get AI Citations

Top rankings don’t automatically translate to AI citations.

When AI breaks a query into sub-queries, it pulls the most relevant and complete source for each one, regardless of where it ranks.

ChatGPT cites pages in position 21+ almost 90% of the time, according to a Semrush study.

Perplexity and Google show the same pattern.

Ranking Positions of LLM-Cited Search Results

AI Retrieves Passages, Not Pages

Rather than directing users to a page, AI systems scan your content and synthesize the exact passage that resolves a query.

This means that the earlier you answer a question, the better your chances of being extracted.

The data backs this up.

44.2% of citations in ChatGPT responses come from the first 30% of a page, while 31.1% come from the middle, and 24.7% from the final third, according to growth advisor Kevin Indig’s analysis of 1.2 million ChatGPT responses.

ChatGPT – Citations from intros

You’re Competing Across a Whole Topic, Not Individual Keywords

SEO often revolves around individual keywords. Query fan-out revolves around comprehensive coverage.

That’s why broad, well-connected coverage across a topic (think pillar pages and topic clusters) can help you earn more AI visibility.

Topic clusters

Pro tip: Pages that rank for fan-out queries (not just the main query) are 161% more likely to get cited, according to a SurferSEO AI Overviews study.


Query Fan-Out Collapses the Buying Journey

We were taught that buyers move linearly — awareness, consideration, decision — and have long optimized content for each stage.

The Marketing Funnel

With AI, those stages collapse into one.

A single high-intent question triggers the system to fan out.

It pulls awareness-level context, consideration-level comparisons, and decision-level specifics into one answer.

The entire buying journey can now happen in a single interaction. So your content needs to work across the full funnel, not just the stage you’re targeting.

Pro tip: Want to work through these steps as you read? Our free Query Fan-Out Audit Template has spreadsheets for tracking your money prompts, sub-queries, intent buckets, and content gaps — plus a checklist to keep the full workflow on track.


The Query Fan-Out Workflow: 6 Steps to Earn More AI Citations

This six-step workflow shows you how to earn more AI citations by identifying and targeting high-impact sub-queries.

It’s repeatable, so you can follow these steps for every topic that matters to your business.

Note: Each AI platform handles fan-out differently, from the number of sub-queries it runs to how it cites sources. We cover the platform differences in depth after the workflow.


Step 1: Find Your Money Prompts

Money prompts are the conversational phrases or questions your ideal customer would ask an AI tool when trying to solve the problem your product or service addresses.

Money prompts are:

  • Typically long-tail and highly specific
  • Tied to a real use case or constraint
  • Close to a decision, not just browsing

Think of money prompts as the AI SEO equivalent of money keywords: high-commercial-intent keywords designed to drive sales.

For example, “noise-canceling headphones ” is a keyword.

“What noise-canceling headphones are best for working from home with kids around, and cost under $300?” is a money prompt.

Noise canceling headphones

Look for money prompts where your audience asks questions:

  • Customer support tickets
  • Community forums
  • Sales call transcripts
  • Internal chat logs
  • Google Search Console queries

For example, when I searched for noise-canceling headphones on Reddit, I found multiple money prompts in real users’ posts.

Like this one that asks for the best noise-canceling headphones for telehealth:

Reddit – Telehealth noise cancelling headphones

And this one asking for durable headphones that will last longer than 2 years:

Reddit – Durable noise cancelling headphones

Forums and transcripts are a good starting point. But you’ll need a dedicated tool to find money prompts using real AI search data.

Semrush’s AI Visibility Toolkit tells you exactly what users type into AI tools, along with the AI’s response.

To show you how it works, I’ll use Bose, a well-known headphone brand, as an example.

Note: I’ll be using Semrush to show you how to complete the query fan-out workflow. If you don’t have a subscription, sign up for a free trial of Semrush One, which includes the AI Visibility Toolkit and Semrush Pro.


First, I searched Bose’s domain in the Visibility Overview tool.

The “Topics & Sources” report revealed over 123.7K prompts where the brand already appears in AI answers.

Visibility Overview – Bose – Prompts

Filtering by “noise canceling” let me dig deeper into topic-specific money prompts like “noise-canceling headphones for sensory issues.”

Visibility Overview – Bose – Prompts – Noise canceling

Clicking the prompt provides a full breakdown: the AI’s response, every brand mentioned alongside yours, and the exact sources it cited.

Visibility Overview – Bose – Prompt details

Follow the same process for your own domain.

These prompts are your highest-priority money prompts — your audience is already searching them, and AI is already answering them.

Don’t have AI visibility yet? Use the Prompt Research tool.

Enter a broad topic to see the prompts that generate the most AI results in your industry.

Prompt Research – Noise canceling headphones

As you find relevant prompts, add them to your spreadsheet.

Even a few money prompts give you enough to work with for the next step.

Fan-Out Audit Template – Money Prompts

Step 2: Generate Your Fan-Out Set

There are two ways to generate fan-out sets: manually or with a dedicated fan-out tool.

The manual approach is free and helps you understand how fan-out behaves, while tools are faster and better suited to working at scale.

I’ll start with the manual method.

Paste this prompt template into any AI platform to get a fan-out set:

Expand this question into the sub-queries an AI system might search to answer it: [your money prompt].


When I ran my Reddit money prompt through ChatGPT, it returned sub-queries grouped into categories:

  • “Core Product Category”
  • “Durability & Longevity”
  • “Battery & Hardware Lifespan”
  • “Reliability & Failure Rates”

ChatGPT – Money prompt

Each category is a potential content gap you’ll address in Step 4.

Run your money prompt through multiple AI tools to get a more complete picture, since each platform tends to expand prompts differently.

Pro tip: Manual research is a solid starting point, but outputs can contain inaccuracies or hallucinations. A dedicated fan-out tool simulates how different AI platforms expand your query and returns an organized list of sub-queries you can act on immediately.


For a faster option, Backlinko’s free ChatGPT Query Fan-Out Tool is worth trying.

Install the Chrome extension, open ChatGPT, and ask your money prompt. The extension captures the response in real time and breaks down every sub-query ChatGPT ran behind the scenes.

When I ran a prompt through it, the panel showed:

  • Each sub-query the model generated
  • The metadata behind the response, including model version
  • Every URL cited, categorized by type: sources, products, images, and news

As you gather sub-queries, assign a query type to each — this tells you what kind of content you’ll need to create in the next step.

Use these definitions to categorize them.

Query Type What It Means
Reformulation A reworded version of the original prompt
Comparative Weighs two or more options against each other
Implicit Addresses a need the user didn’t explicitly state
Personalized Tailored to a specific situation, constraint, or preference
Entity expansion Drills into a specific brand, product, or person mentioned
Related A connected topic the AI anticipates the user might want next

Step 3: Bucket Sub-Queries by Intent Type

Bucketing by intent tells you what types of content to create and the ideal format for each.

To categorize a sub-query, answer this question: What does the person actually want to do after getting an answer?

Consider an example from the noise-canceling headphones query fan-out set: “Sony vs Bose Noise Canceling Headphones.”

Someone asking this is weighing two specific products against each other, so it’s a “comparison” query.

Fan-Out Audit Template – Intent Buckets

The right format for this query is a head-to-head comparison page or table, not a general buying guide or listicle.

The intent isn’t always this obvious, and some sub-queries may fit more than one bucket.

When that happens, place it where the strongest intent lies.

Here’s a general guide to the main intent buckets and what each one calls for:

Bucket Description Example Sub-Query Content Format
Definitions / Basics What is X? How does X work? “how do noise canceling headphones work” Explainer article, glossary section
Comparisons / Alternatives X vs Y, alternatives to X “apple airpods max vs sony wh 1000xm4” Comparison page, head-to-head section
Best for X / Recommendations Best option for a specific use case “best noise canceling headphones for working from home” Listicle, buying guide
Problems / Troubleshooting How to fix X, why does X happen “how to get rid of background noise in audio” How-to guide, FAQ section
Pricing / Value How much does X cost, is X worth it “are there any good wireless headphones with noise cancellation under $150?” Pricing page, value comparison section
Social Proof / Discussions Reviews, Reddit opinions, user experience “best earbuds for calls in noisy environment reddit” Review roundup, user feedback section

Step 4: Audit Your Existing Content for Gaps

Once you’ve bucketed your sub-queries by intent and format, check which ones your site already covers and which ones it doesn’t (aka content gaps).

Start by searching your own site.

Type “site:yourdomain.com [sub-query topic]” into Google.

For example, running “site:bose.com noise canceling headphones” surfaces all their pages on that topic.

Google SERP – Bose – Noise canceling headphones

From here, evaluate each page against the sub-query it should cover:

  • Coverage: Does it directly answer the sub-query, or just mention the topic in passing?
  • Format: Is it the right content format for the intent?
  • Self-contained answers: Can the answer stand on its own, without the reader needing to look anywhere else?

Categorize each page by its coverage level:

Coverage Level What It Looks Like What to Do
Not covered No page on your site addresses this sub-query at all Create new content targeting this sub-query directly
Partially covered A page mentions the topic in passing but doesn’t resolve the sub-query directly Add a dedicated section to the existing page that fully answers the sub-query
Fully covered A dedicated section or page answers the sub-query completely and can be extracted and cited by AI without needing surrounding context Monitor for AI citations and update regularly to stay current

For each sub-query, you’ll also want to know which competitors are showing up for your money prompts.

Run your money prompts through AI platforms to gather this information manually. Or refer back to your research from the AI Visibility Toolkit in Step 1.

Click any prompt to see which brands were mentioned and the exact sources the AI cited.

Bose – Prompt details – Brands & Sources

Already showing up alongside competitors? That’s a prompt worth protecting — focus on strengthening your coverage so you stay in the answer.

If competitors are showing up and you’re not, that’s a gap worth closing before they own it.

Fan-Out Audit Template – Content Audit

Step 5: Structure Your Content So AI Can Extract It

Creating the right content is only half the job. The other half is making it easy for AI to find, parse, and use.

Start by filling the gaps you identified in Step 4.

For sub-queries with no coverage, create dedicated pages or sections that target them directly.

For partial coverage, add self-contained answers to existing pages that resolve the sub-query without needing surrounding context.

Then, structure everything so AI can extract it cleanly:

  • Address specific questions directly — lead with the answer, not background context
  • Use content chunking: Break content into focused sections with clear headings, short paragraphs, and bullet points
  • Front-load key information early in the page or section
  • Use clear, precise language, including specific product names, figures, and use-case-specific wording
  • Add FAQ sections

Here’s what this looks like in action.

Bose has over 63.9K mentions across AI platforms in the U.S. alone:

Visibility Overview – Bose

It helps that they’re a household name. But their content is also built to be extracted.

Their product pages front-load specific claims as scannable elements — “24 hours of battery life” and “legendary noise cancelation” — rather than burying them in copy.

Bose – Product features

Key specs are organized into structured comparison tables:

Bose – Product specs

And they build dedicated landing pages for use cases like flying, using descriptive, scenario-specific language.

This matters because AI fans out into use-case-specific sub-queries.

Bose – Noise cancelling headphones for flights

When I searched “best noise-canceling headphones for flight anxiety,” AI Mode recommended Bose, using nearly identical language from Bose’s flight landing page.

Google AI Mode – Noise canceling headphones

When a user’s prompt matches the scenario your page was built for, AI systems may be more likely to pull from it.

This is a clear example of that in action.

You don’t need a complete site overhaul to make this work.

Even restructuring a few high-priority pages to address your fan-out gaps can improve your chances of being extracted and cited.

Step 6: Measure Your Performance in AI Search

Once your content is structured and live, track your performance in LLMs.

Start with the money prompts you identified in Step 1.

For each one, you want to know:

  • Are you showing up? Is your brand mentioned or recommended in the response?
  • Is what it says accurate? Are the claims the AI makes about your brand correct, or is it pulling outdated or wrong information?
  • How do you compare? Which competitors appear in the same response, and how are they positioned relative to you?

If you’re tracking manually, run them through multiple LLMs (in a private or incognito window) and record what you find.

ChatGPT – Bose headphones

But once you’re tracking dozens of sub-queries across platforms, manually tracking gets messy (and time-consuming).

I use Semrush’s Prompt Tracker to automate the process.

It alerts you to changes in mentions for your money prompts, so you don’t have to keep re-running them yourself.

Position Tracking – Keywords

Another helpful tool is the Visibility Overview.

It provides an AI visibility score that tracks how often you’re showing up in AI answers compared to competitors.

Visibility Overview – Bose

The Perception tool tracks sentiment so you know how LLMs describe your brand — and if they mention competitors more favorably.

Perception – Bose – Sentiment

It also breaks down the factors driving that sentiment.

For Bose, “industry-leading noise cancellation” shows up as a strength, while “over-the-ear models not sweatproof” flags a use-case they could address with targeted content.

Perception – Bose – Key sentiment drivers

Tracking should be an ongoing process.

Revisit your money prompts regularly and update your content as new sub-queries emerge or competitors gain ground.

How Query Fan-Out Works Across Different Platforms

How content surfaces in an AI answer depends on several factors:

  • Whether the system searches the live web or draws from its training knowledge
  • How many sub-queries it runs
  • Which sources it favors, and how it cites them

Understanding those patterns helps you make smarter decisions about content structure, format, and where to focus your optimization effort.

Plus, if a competitor outperforms you in a specific LLM, understanding how that platform handles fan-out can help you figure out why.

Platform How Fan-Out Works
ChatGPT Reasons internally, then runs live web searches when a question requires fresh data, comparisons, or current information
Perplexity Combines conversation context with real-time web search
Claude Clarifies intent first; relies mostly on training data
Google AI Overviews Synthesizes Google’s index into condensed, featured-snippet-style summaries
Google AI Mode Breaks complex prompts into multiple searches across Google’s index

Note: Some of the behavior described below is based on how each system describes its own reasoning when prompted. LLMs aren’t always reliable narrators of their own processes, so treat these observations as directional rather than definitive.


ChatGPT

For simple, informational queries, ChatGPT usually responds from its training data without running a live search.

ChatGPT – Compound interest

But that changes when the question requires fresh information, comparisons, or real-world data.

When I asked which car I should buy (Toyota vs. Honda) in Thinking mode, ChatGPT spent about 22 seconds reasoning through the question.

Then, it produced an answer drawn from 41 cited sources

ChatGPT – Toyota vs Honda

That’s query fan-out in action: one prompt, varied sources, and multiple sub-queries running behind the scenes.

By default, you can’t see the sub-queries ChatGPT runs. But I’ll show you how to find them (don’t worry — it’s easier than it looks).

Note: This DevTools method only works in the web version of ChatGPT. You can’t access sub-query data on mobile or in the desktop app.


First, search a money prompt in ChatGPT.

Then, look at your browser’s address bar and copy the slug that appears after chatgpt.com/c/ — that’s the unique ID for your conversation

ChatGPT – URL

Next, right-click anywhere on the page and select “Inspect.”

ChatGPT – Inspect

A developer panel will open on the side of your screen:

  • Click “Network” at the top of that panel
  • Paste the slug you copied into the filter bar
  • Refresh the page

Click on the fetch version of the slug (here, it’s the second option under the Name column).

Chrome DevTools – Network

Then, open the Response tab.

Chrome DevTools – Network – Response

Once it loads, press Ctrl+F (or Cmd+F on Mac) and search for the word “queries.”

Chrome DevTools – Network – Response – Queries

What appears is the exact set of internal searches ChatGPT ran before producing its answer.

For the Toyota vs Honda prompt, ChatGPT generated queries around:

  • Vehicle specifications
  • Fuel economy
  • Reliability
  • Safety ratings
  • Long-term ownership costs

Once you have the sub-queries, cross-reference them against your content.

Are you targeting each one? Do your pages use the same language ChatGPT is searching for — “long-term ownership costs” rather than just “value”?

ChatGPT often pulls from third-party sources like Reddit threads, review sites, and comparison pages.

So topical authority matters here — not just what’s on your site, but whether your brand shows up across the sources ChatGPT is likely to retrieve.

Perplexity

Perplexity runs two types of fan-out simultaneously:

  1. Internal fan-out — scans your prior conversation history for relevant context
  2. External fan-out — searches the external web for relevant information

The final answer draws on both layers, which means your content needs to work for a range of user situations, not just one.

For the Toyota vs. Honda question, Perplexity’s first batch of sub-queries had nothing to do with the cars.

Perplexity – Toyota vs Honda

Instead, it checked whether I’d previously mentioned anything that could shape its recommendation.

Perplexity – Toyota vs Honda – Subqueries

Like budget constraints, driving habits, or past questions about either brand.

Perplexity – Toyota vs Honda – Subqueries – Details

Only after that internal scan did it launch external searches about reliability, ownership cost, and safety ratings.

What this means for your content: Perplexity may pair your page with context you can’t predict: a user’s past questions, constraints, or preferences.

Your content needs to be specific and self-contained enough to remain accurate and useful no matter the surrounding context.

Claude

Claude takes a different approach.

Rather than immediately running sub-queries, it asks clarifying questions first. Then, it generates a response tailored to your answers.

When I asked the Toyota vs. Honda question, Claude presented a preference widget before producing an answer.

Claude – Toyota vs Honda

Once I responded, it generated a recommendation tailored to my priorities.

Claude – Toyota vs Honda – Answer

Because it clarifies intent before searching, Claude tends to generate fewer, more targeted fan-out sub-queries than other platforms.

The implication for your content: Answer specific, well-defined use cases directly rather than trying to cover every angle on a single page.

Google AI Overviews and AI Mode

AI Overviews appear as concise, AI-generated summaries with sources listed in a clickable sidebar.

Google SERP – Toyota vs Honda – AI Overview

They work by synthesizing Google’s existing web index into a tighter, more contained summary.

AI Mode, by contrast, is a dedicated conversational search tab designed for complex, multi‑part questions.

Google AI Mode – Toyota vs Honda

Like AI Overviews, it draws on Google’s index to generate answers, but it offers more interaction and depth.

Neither platform exposes the sub-queries it runs.

But SEOs have found a way to extract Google’s fan-outs using Screaming Frog configured with a Gemini API. Watch Dan Hinckley’s tutorial for a full walkthrough.

For both, the optimization focus is the same: Front-load your answers, use descriptive subheadings, and structure content so individual passages stand on their own.

AI Search Runs on Query Fan-Out — Your Content Strategy Should Too

High rankings alone won’t earn AI mentions.

The brands showing up are the ones covering the questions their audience is actually asking and making that content easy for AI to extract and cite.

You’ve got the query fan-out framework. Now it’s about execution.

Start with one money prompt, map the sub-queries, and audit where your content stands.

Then work through the gaps, one topic at a time.

Next, dive deeper into how to get your brand seen and trusted across AI platforms with our AI search strategy guide.

The post Query Fan-Out: What It Is and How It Affects AI Visibility appeared first on Backlinko.

Read more at Read More

Using AI to Support and Defend Your Brand

Key Takeaways

  • AI-generated answers have compressed brand discovery into a single moment. One summary can now serve as a customer’s entire first impression.
  • AI systems pull from a wide range of sources, including forums, review sites, and outdated content, not just your owned properties.
  • The most repeated claim tends to surface in AI outputs, not necessarily the most accurate one.
  • Inconsistent messaging gets amplified by AI, not smoothed over.
  • Content governance, proactive publishing, and continuous monitoring are the new foundations of brand reputation management.

    Brand management has a new problem. Everything you have built, your positioning, your messaging, your reputation, can now be summarized by an AI system before a customer ever visits your site, reads your content, or talks to your team. That summary may be accurate. It may not be. The person reading it likely has no way to tell the difference.

    This is not a hypothetical risk. It is happening continuously, across every major AI platform, for brands of every size. The question is not whether AI is shaping how people perceive your brand. It is whether you are doing anything to influence what AI says.

    The First Impression Problem

    People used to form impressions of brands gradually. They encountered coverage, read reviews, visited a website, spoke with someone. Perception built up over multiple interactions, giving brands time to shape it.

    That process is being compressed. An AI-generated answer can now stand in for all of those touchpoints. A prospective customer asks ChatGPT or Perplexity about your company, gets a two-paragraph summary, and walks away with a complete impression, accurate or not, before ever interacting with anything you control.

    A graphic showcasing brand hijackings in AI search ads on ChatGPT.

    What makes this genuinely difficult is how AI builds those summaries. It does not prioritize your owned content. It pulls from whatever it can find: your website, press coverage, review platforms, social media, forum discussions, complaint boards. It weighs those sources by factors that are not always intuitive. A high volume of low-quality negative content can outweigh a smaller volume of accurate positive content. Old information that has not been addressed or replaced sits alongside current content, with no timestamp visible to the user.

    Your brand’s AI reputation is shaped by your entire content footprint, not just the parts you have invested in carefully.

    The Risk Goes Beyond False Information

    Most brands are not facing outright fabrication. The more common risk is partial truths: accurate statements pulled out of context, outdated information that was once correct, nuanced positions simplified into something that no longer reflects where you actually stand.

    Partial truths are more insidious than false information because they are harder to dispute and easier to spread. Once an AI system has assembled a narrative from the sources it has found, that narrative gets reinforced every time someone asks a related question. It becomes what people know about you, and correcting it requires more than just publishing accurate content. It requires replacing the sources the AI is drawing from.

    A ChatGPT query about the best plumbing companies in the Chicago area.

    There is also a compounding effect to be aware of. AI-generated summaries get shared across platforms. Screenshots get posted. Those shares become new inputs that reinforce the same narrative in future AI outputs. A problematic summary does not stay contained.

    The practical consequence is straightforward: the most accurate claim does not automatically rise to the top in AI outputs. The most repeated claim does.

    Content Governance Is Brand Protection Now

    The practical response to this challenge starts with content governance, and governance needs a different frame than it typically gets in marketing organizations.

    Most brands treat governance as an internal process concern: who approves content, how brand guidelines get followed, what templates teams use. Those things matter. In an AI-mediated environment, though, governance is the mechanism that determines whether AI systems can accurately summarize who you are. It is infrastructure, not administration.

    As one brand governance expert put it: this “ensures that the core signals of your brand are clear enough to survive the compression that happens through an AI component.” When brand signals are inconsistent or vague, AI amplifies that inconsistency rather than resolving it.

    Messaging consistency across every touchpoint. If different teams, regions, or channels are publishing different descriptions of your product, your mission, or your positioning, AI will find all of them and combine them into something that may not accurately represent any of them. A unified source of truth that every piece of external content draws from is the foundation.

    Content that explains rather than claims. AI systems have no way to evaluate vague marketing language. Terms like “industry-leading” or “innovative” mean nothing to an AI summarizing your brand. What does register is specific, plain-language explanation of what you do, how you work, and why it matters. Replace generic claims with clear explanations throughout your owned content.

    Your website treated as AI infrastructure, not just a marketing asset. Most organizations still build their websites primarily as human-facing experiences. For AI systems, your website is often the first place used to understand your organization. Review your key pages with one question in mind: could an AI produce an accurate summary of your brand from what we have published here? If the answer is no, you have content work to do.

    Taking an Active Role in What AI Says About You

    Governance handles internal consistency. The external picture requires a more active approach.

    Start by auditing what AI systems are currently saying about your brand. Prompt ChatGPT, Google AI Overview, and Perplexity with the questions a prospective customer, investor, or journalist would ask. Capture those outputs. Then trace the narrative back to its sources. Are those sources accurate? Current? Are there negative or outdated sources being weighted heavily because you have not published sufficient structured content to counter them?

    Using our Chicago plumber example from before, we see Angi is heavily weighted as a source in that ChatGPT answer.

    An Angi landing page dedicated to Chicago plumbers.

    That audit gives you a content agenda. Gaps in AI representation can often be addressed by publishing clear, well-structured content that gives AI systems better information to pull from. If outdated claims are being surfaced, identify the sources driving them and address those sources directly. Claims spreading on Reddit or social platforms can be addressed on those platforms. 

    A Reddit post axsking about Chicago plumbers with responses.

    Structured explanations published through FAQs and policies give AI systems better, more current information to draw from.

    Third-party credibility carries significant weight. Earned media, analyst coverage, and credible reviews are treated as high-trust signals by AI systems that evaluate external validation. Proactive brand publishing and digital PR work are not just marketing tactics in this environment; they are inputs that shape what AI says about you before a narrative hardens.

    Spokespeople and executives also need to think about this. In a traditional media environment, journalists contextualize statements. In an AI-mediated environment, those statements get pulled directly into summaries. Specificity and context matter more than polished soundbites. Complete explanations travel better than compressed talking points.

    Monitoring Cannot Be Periodic

    One of the most common mistakes brands make with AI reputation management is treating it as a project with a completion date. You audit, fix the gaps, and move on. That approach misses how dynamic the AI reputation environment actually is.

    New coverage, a viral social post, a competitor’s messaging shift, or a change in how your content is indexed can all alter what an AI says about your brand. The only way to stay ahead of narrative shifts before they harden is to monitor consistently, not quarterly.

    Brand-based prompts in Writesonic.

    Build a standing practice of prompting major AI tools with brand-relevant queries on a regular cadence. Track what changes. Create workflows for responding to misinformation on the platforms where it originates, before it has time to proliferate. Think of AI reputation management the same way you think about SEO: something that requires continuous attention, not a one-time fix.

    FAQs

    How often should I audit what AI says about my brand?

    Monthly at minimum, with closer attention during periods of significant company news, product launches, or any event that generates substantial external coverage. AI systems update as the web updates, so the outputs you capture today may not reflect what users see in six weeks.

    What content is most effective at influencing AI summaries?

    Clear, specific, well-structured content that directly addresses the questions people ask about your brand. FAQs, plain-language product explainers, executive Q&As, and detailed company descriptions all register more effectively than vague marketing copy. Third-party coverage from credible sources also carries high signal weight.

    What should I do if AI is saying something inaccurate about my brand?

    Identify the sources driving the inaccurate narrative. Address misinformation directly on the platforms where it originated (forums, review sites, social media). Publish structured, authoritative content that provides AI systems with better information to draw from. Building third-party credibility through earned media helps establish accurate narratives as the dominant signal over time.

    Conclusion

    The question brand managers need to be asking has shifted. It is no longer just “what message do we want to put out?” It is “what will AI tell someone about us, and is that accurate?” Answering that question requires consistent messaging, clear content, active monitoring, and a willingness to treat AI reputation as a standing business function rather than a marketing add-on.

    The brands that build that infrastructure now will have a meaningful advantage as AI-mediated discovery continues to grow. The brands that do not will find their reputation increasingly shaped by whatever AI happens to find first.

    Read more at Read More

    Google Is Testing Sponsored Shops in SERPs: What This Means for Advertisers

    Key Takeaways

    1. Google is testing “Sponsored Shops,” a format that groups multiple products from a single retailer into one branded unit inside Shopping results.
    2. This moves competition from the product level to the retailer level, changing what it takes to win visibility.
    3. Feed quality, seller ratings, and assortment depth become more critical than ever.
    4. The format introduces multiple click paths within one ad unit, which could complicate attribution and traffic flow.
    5. Performance Max is a likely vehicle through which Sponsored Shops placements will be accessible when the format formally launches, but nobody knows for sure.
    6. Brands that build strong store-level signals now will be better positioned if and when this rolls out broadly.

    Google is running a Shopping test that could change how brands compete for visibility in product search. If it scales, the rules shift, and advertisers who see it coming will have a head start.

    Here’s what’s happening and what you should be doing about it right now.

    What Is Google Actually Testing?

    Google’s Sponsored Shops test groups several products from one retailer into a single ad unit inside Shopping results, alongside the store name, ratings, and brand signals. Think of it as a mini storefront sitting directly inside the search results page, rather than a row of individual competing products.

    Sponsored shops results for backpack.

    Source

    It is still a test. Google has not confirmed a broad rollout. The direction it points toward matters, though, and Shopping advertisers should be paying close attention.

    The test does not exist in isolation. It is part of a broader shift Google has been building toward for a while: more brand-centric, discovery-oriented, and AI-mediated shopping experiences. In 2025, Google introduced the Merchant Brand Profile feature, which lets retailers build brand-presence pages in search with lifestyle images, videos, and business descriptions. 

    An example business in Google Sponsored shops.

    Source

    Sponsored Shops looks like the logical next step in that direction, bringing brand identity directly into the Shopping ad unit itself.

    Why the Format Change Is a Bigger Deal Than It Looks

    Right now, Shopping competition is largely a product-level game. Your listing competes against a competitor’s listing. Better feed, stronger bid, you take the placement.

    Sponsored Shops changes the terms of that competition. Instead of a single product earning a spot, your entire store is on display at once: assortment, brand presence, and ratings together. A competitor with a stronger catalog and better seller signals will have a structural advantage that no amount of bid optimization can fully offset.

    That’s a meaningful shift. Brands that have been winning through finely tuned individual product listings will need to think harder about how their store presents as a whole. Brands that have invested in feed quality, customer experience, and assortment depth will find that investment paying off in ways it didn’t before.

    There’s also a measurement angle worth flagging. A single ad unit with multiple clickable elements (store name, individual products, ratings) creates multiple potential click paths. How traffic splits across those paths, and how that maps to your current attribution model, is an open question every Shopping advertiser should be thinking through before this format scales.

    What This Signals About Where Google Is Headed

    Google has been explicit about where it wants Shopping to go. In its own communications about 2026 priorities, the company described its goal as making search “a more powerful tool for discovery, where ads can inspire and answer all at once.” AI Mode already surfaces organic shopping recommendations based on query relevance, and Google has confirmed it is testing a new ad format inside AI Mode that showcases retailers offering relevant products, clearly marked as sponsored.

    A ChatGPT result for men's running shoes black.

    Source

    Sponsored Shops fits squarely into that roadmap. It moves Shopping slightly up the funnel, making it as much about brand discovery as product comparison. Rather than a format designed purely to capture demand-ready buyers, it is designed to let brands show up with range and identity in front of people who are still forming their consideration set.

    For users, the format is intuitive. Browsing several products from the same retailer without leaving the results page is a better experience than clicking in and out of individual listings. Google tends to expand formats that improve user experience. That’s worth taking seriously.

    The PMAX Connection

    As of right now, we don’t know what vehicle is going to power sponsored shops. Performance Max is a likely bet based on volume and Google’s push for PMax adoption, but nothing is confirmed. PMax already accounts for roughly 62 percent of Google Shopping spend among major advertisers, and it is already designed to surface both store-level and product-level assets dynamically across Google’s ecosystem.

    With this said, though, AI Max for shopping is still in beta, so that might impact what plays a role. We also know that Google does tend to favor some of their newer products which likely helps adoption rate (e.g. AI Max, PMax, & Broad being eligible for AIO ad placements).

    What to Do Before This Rolls Out

    You do not need to wait for a full launch to get ahead of it.

    Start with your product feed. Feed quality has always mattered in Shopping, but a storefront format makes weak data much more visible. Every title, description, image, and availability signal is part of how your store presents in that unit. Get it right now. Research consistently shows that product titles, images, and product identifiers are the three highest-impact feed optimizations, and all three will matter even more in a store-level display format.

    Google results for gymshark tshirts.

    Source

    Take stock of your seller ratings. In a storefront format, ratings are far more prominent than they are in individual listings. If you have not been actively managing reviews and customer experience signals, that needs to change. A store-level placement that leads with a weak rating is a self-defeating ad.

    Look at assortment depth. A Sponsored Shops unit showing three products when a competitor shows ten is a losing presentation. Review whether your full catalog is properly represented in your feed and close any gaps.

    Audit your PMax asset groups. Given that PMax is the likely vehicle for Sponsored Shops placements, your asset groups should be fully built out with all image formats, high-quality lifestyle images alongside product images, accurate brand descriptions, and audience signals that represent your full customer base rather than just buyers of individual products.

    Revisit your attribution setup. Multiple click paths inside a single unit means your current reporting may not capture traffic flow accurately. Think about how you will measure this before the format exists in your account at scale.

    FAQs

    What exactly is a Sponsored Shops unit?

    A Sponsored Shops unit groups multiple products from a single retailer into one ad block inside Google Shopping results, displayed alongside the store name, ratings, and brand signals. Rather than individual product listings competing side by side, the format presents a mini storefront for a single brand.

    Is Sponsored Shops live now?

    As of now, Sponsored Shops is still in testing. Google has not confirmed a broad rollout timeline. The format is worth preparing for regardless, since the steps that improve your eligibility for it also strengthen your existing Shopping performance.

    Which campaign type will Sponsored Shops use?

    Performance Max is the most likely vehicle, given that it already accounts for the majority of Shopping spend and dynamically surfaces store-level and product-level assets across Google’s ecosystem. Making sure your PMax asset groups are fully built out is the right preparation move.

    Will smaller retailers be disadvantaged?

    Formats that reward assortment breadth, seller ratings, and feed quality tend to favor established retailers with larger catalogs and more customer reviews. That said, a well-optimized feed and a strong seller rating matter more than raw catalog size. Smaller retailers with tight assortments and excellent customer experience signals are not automatically excluded.

    What should I do right now?

    Focus on feed quality, seller ratings, and PMax asset completeness. These are the fundamentals that will determine Sponsored Shops eligibility and performance when the format expands, and they are also the fundamentals that determine your current Shopping performance.

    Conclusion

    Sponsored Shops is still in testing. Google Shopping is clearly moving toward a model where brands compete as storefronts, not just as individual products. The shift fits a broader pattern: more AI-mediated discovery, more brand-level visibility signals, more emphasis on the full store experience rather than the individual listing.

    The time to build those store-level signals is before the competition catches up, not after. The good news is that everything you do to prepare for Sponsored Shops makes your existing Shopping campaigns stronger right now. There’s no downside to starting.

    Read more at Read More

    Google zero-click searches hit 68% in early 2026: Study

    Google zero

    Google searches ended without a click 68.01% of the time in the U.S. during the first four months of 2026, according to new SparkToro research based on Similarweb clickstream data. That’s up from 60.45% in 2024, a 7.56-point increase in two years.

    Fewer searches result in clicks. The share of searches generating at least one click fell 9.51 percentage points between 2024 and 2026 (a 22.9% decline), according to SparkToro. This includes clicks to organic results, paid ads, and Google-owned properties such as Maps and YouTube, but excludes follow-up searches within Google.

    • Over the same period, the share of searches that led to another Google search rose 7.2 percentage points.
    • This trend reflects Google’s growing ability to answer questions directly in search results while encouraging users to refine or continue their searches within Google, according to SparkToro.

    AI Overviews and zero click. SparkToro believes AI Overviews are likely contributing to the increase in zero-click searches, though the study doesn’t isolate the extent to which the overall rise between 2024 and 2026 can be attributed specifically to AI Overviews.

    • AI Overviews now appear on more than 20% of Google searches, according to the research. When they do, click-through rates drop by nearly 60%.

    AI Mode and zero click. It appears to have played only a limited role during the January to April study period. SparkToro found that just 0.34% of searches transitioned into AI Mode during that time.

    • However, Google said at I/O 2026 that AI Mode had surpassed 1 billion monthly users and that query volume was more than doubling each quarter, suggesting its impact on search behavior could grow significantly.

    Zero click history. SparkToro has tracked zero-click search behavior for years, though its underlying data sources have changed over time. Because the studies rely on different providers, panels, and methodologies, long-term comparisons are not directly equivalent. Still, the available data consistently points to a rise in zero-click behavior over time, according to SparkToro.

    Why we care. The findings suggest Google is increasingly satisfying user needs without sending users to external websites. However, you should interpret direct comparisons across years cautiously because SparkToro’s historical analyses rely on different clickstream data providers and panels.

    SEO still matters, but… SEO alone may be insufficient for many publishers seeking to regain historical levels of Google-referred traffic. SparkToro co-founder Rand Fishkin recommended investing in brand awareness and influence on the platforms where your audience already spends time, regardless of whether those efforts drive direct website visits.

    • Some categories continue to benefit significantly from SEO, including branded searches, local business queries, and high-intent transactional searches, Fishkin said.

    About the data. The study used Similarweb desktop and mobile web panel data covering U.S. Google searches from January through April 2026. SparkToro assumed that two-thirds of searches occurred on mobile devices and one-third on desktops. The analysis excludes searches conducted in Google’s mobile search app, where SparkToro said zero-click behavior may be even higher.

    The study. In 2026, Less than One Third of Google Searches Still Send a Click

    Read more at Read More

    Web Design and Development San Diego

    How AI forms opinions about your brand

    How AI forms opinions about your brand

    AI forms opinions about your brand from what it can see online. That’s your digital footprint.

    The problem is that AI often sees only fragments of your business. It sees your website, content, reviews, and mentions, but much of the expertise, customer insight, and operational knowledge that makes your business valuable never makes it into the digital footprint.

    The solution is to surface that knowledge, organize it into a single source of truth, and turn it into machine-readable signals. Here’s how to collect it, organize it into a single source of truth, and distribute it across the channels AI uses to understand, evaluate, and recommend brands.

    What you feed the machines is understandability, credibility, and deliverability (UCD)

    Everything you put into your footprint is fodder for three things AI has to decide about you. Together, they provide the fodder for the whole funnel.

    Understandability

    Does AI know who you are, what you do, and who you serve? You already know where your understandability comes from: 

    • Your about page.
    • Your product pages.
    • Your structured data. 

    What often gets missed is the operational detail that explains what you actually do once a client is inside.

    Credibility

    Does AI believe you’re good at it? This is N-E-E-A-T-T credibility — notability, experience, expertise, authoritativeness, trustworthiness, and transparency, an extension of Google’s E-E-A-T.

    You know what credibility signals you currently feed: your case studies, your credentials, and your testimonials. What many businesses don’t realize is how much N-E-E-A-T-T credibility is already embedded in their day-to-day operations.

    Deliverability

    Does the AI engine have the content to hand you to the subset of its users who are your audience? 

    You know where your deliverability comes from: the topical content, the marketing, and the authority pieces you commission. Deliverability is often hiding in plain sight, in the content generated by your business operations and offline activities.

    Your customers search everywhere. Make sure your brand shows up.

    The SEO toolkit you know, plus the AI visibility data you need.

    Start Free Trial
    Get started with

    Semrush One Logo

    5 streams of business data feeding every commercial surface

    All three elements of the UCD trio are fed by the five inputs below, and how much each contributes varies by business.

    The point isn’t to file each input under one letter. Organized and codified, the five together give AI the fodder it needs from top to bottom of the funnel.

    5 streams of business data feeding every commercial surface

    1. Products and services: What you sell, and you already do it

    Your products and services data: what you sell, at what price, under what conditions, and with consistent names and identifiers. This is mostly about understandability, with credibility riding alongside it.

    Most businesses already do this, so the work is in the depth, not the effort. Don’t just list what you sell. Describe who each offering is for, what problem it solves, what it costs, what it doesn’t do, and how it differs from the next option.

    A thin product page tells AI a product exists. An exhaustive one tells it when to recommend that product and to whom.

    Keep it accurate, complete, and consistent with everything else in your footprint. A price or product name that differs across pages reads as doubt.

    2. Authority content: Your expertise, and almost everybody does it

    This is the marketing you already create to show you know your field: your articles, videos, guides, data studies, and the thought leadership you publish to tick the box marked “content created.”

    People put effort into it to build authority, rank, do SEO, and position themselves as experts. That’s fine. It leans toward deliverability because it’s what tells AI which territory to surface you in.

    But everybody does it, which is exactly why it’s the least differentiating of the five on its own. It earns its weight only when it’s tied to the rest: the same expertise proven by your operations and corroborated by third parties, not just asserted in a blog post.

    It’s necessary, but it’s not where your advantage hides.

    3. Brand narrative and voice: Who you are, who you serve, and why you’re the best

    All marketers create brand narratives, so the work here is about consistency and clarity rather than invention. Everybody communicates who they are, what they do, and who they serve, and keeping that clear and consistent matters enormously. 

    But three things are often left out, and AI needs all of them.

    • Intent: It isn’t enough to name your ideal customer profile (ICP). You have to pair your ICP with what they’re after: the cohort-to-intent combinations from the funnel query pathway. AI has to know not just whose problem you solve, but which problem, and at which moment, before it can hand you to them.
    • Credibility: The thing that feeds your N-E-E-A-T-T. Many people leave it out because they feel awkward saying it. You have to set it out because AI won’t work out your true value on its own. Be clear and bold about why you’re credible, then make sure you can back it up with evidence.
    • Making the relationship with your clients explicit: Validation from the people you serve that you deliver on what your narrative and cohort-to-intent mapping promise. Say who you are, what you do, and who you serve. Then explain why a customer should choose you and prove it.

    Voice is the part corporations get wrong most often. Narrative is what you say. Voice is how you say it. One team may write the narrative once, but voice escapes through every rep, every support reply, every social post, and every deck. 

    When it drifts, and in most large companies it drifts constantly, AI reads the same brand as five different brands and loses confidence in all five.

    So standardize your voice and keep it consistent everywhere. Consistency is a credibility signal in itself. Inconsistency is a tax you pay without seeing the bill.

    In short, make sure your brand narrative clearly sets out your ICP, who you are, and why you’re the best fit for them, in a voice that stays consistent wherever AI finds it.

    4. OPID business operations: The stream almost nobody harvests

    This is everything your business generates by running: onboarding, performance, integration, devotion, and all the day-to-day activity around them. 

    It’s the most powerful of the five because the material comes from your clients and from the work your team does to serve them, which is exactly the material that rarely makes it online. It sits behind closed doors, buried in a CRM, parked on a platform nobody values, and almost nobody harvests it.

    It feeds all three elements of understandability, credibility, and deliverability more effectively than anything else you own. 

    • Understandability comes from the granular detail of what you actually do and the exact circumstances in which you help. Most of that is only ever discussed inside the business. A review where a client describes precisely what they got from you puts something on the record you’d never say about yourself, and the machine reads it as fact.
    • Credibility is your N-E-E-A-T-T, and this is the most convincing kind because it comes from clients themselves, not from your marketing.
    • Deliverability comes from the match. The content here aligns exactly with your cohort-to-intent combinations because it was created around the clients you attracted and served well. Whether it comes from you or from them, it fits the audience and intent you need to communicate to the engines.

    Once you start looking, you’ll find the richest material you own:

    • Customer voice is the highest signal because it’s real questions in real language: reviews across every platform, written and video testimonials, FAQs, unpublished support questions that should become FAQs, support and sales call transcripts, onboarding and churn-exit interviews, and free-text survey responses.
    • Evidence and outcomes provide the proof you need: case studies with real before-and-after numbers, patent filings, academic deposits that are public but underused, and independent third-party studies that corroborate your claims.
    • Methodology covers the rest. SOPs, playbooks, training materials, glossaries you currently keep private, and long-form spoken content such as webinars, keynotes, and podcast appearances, transcribed.

    Look for material that answers a question an assistive engine or agent actually gets asked, in the questioner’s own words, with a verifiable fact attached. 

    A support ticket, churn interview, or sales call transcript will often outperform polished marketing copy in that test because it’s already phrased the way real people ask questions.

    That’s the whole point of harvesting OPID business operations: taking information from a place AI can’t see and moving it to a place where it can, while making it visible to your human audience, too. It’s convincing to both because it’s true and because it matches the cohort-to-intent combination exactly.

    5. Bringing the offline online: The stream almost nobody runs

    This section is all about the marketing and audience engagement you do offline: the talks you give, the festivals or hackathons you sponsor to support your community, the interviews, the panels, and the rooms full of clients. It’s obvious to you, but largely invisible to AI.

    Bring the offline online and feed it to the machines by publishing self-reporting content and linking to the social posts and summary articles others write. That’s a huge win most brands miss.

    But it works the other way, too. Your codified source of truth can feed your offline communication, so the story a client hears from you at a conference, in a newspaper, on the radio, or face to face is consistent with the story you’re telling AI on the web.

    That matters more than it seems. If the two differ, you lose the person because the gap reads as doubt to a human and as low confidence to a machine.

    Clarity and consistency over time, online and offline, is the name of the game.

    Get the newsletter search marketers rely on.


    Organize and codify the five into one source of truth

    Once you’ve harvested all five streams, organize and codify them into a single source of truth: a database you build to output whatever format each surface needs, including HTML, schema, MCP, RDF, prose, audio, video, and images.

    Organize the data once, centralize it, set up a system that codifies it on the way out, and from there you can distribute it in a few clicks while your digital footprint stays clear and consistent as it grows.

    Then distribute it across your digital ecosystem in the format your human audience expects and packaged so machines can ingest it cleanly.

    Where you publish affects how much the machine believes you, and the rule is simple: the less of you there is in it, the more it trusts it. You’re working across three tiers.

    First-party: You claim 

    You publish on your own properties, in your own voice. You state who you are and set the frame. It’s the baseline, and on its own it proves nothing because you wrote it and you published it.

    Second-party: You corroborate

    Here, you’re still publishing, but across a broader footprint and with other voices in the mix. Two things widen here.

    • The platform: In addition to your own entity home website, you publish on platforms where you own the account, such as YouTube, LinkedIn, Medium, and press releases. You’re stating your case the same way you would on your website, just on another property you control.
    • The voice: You can publish your own words, or you can publish what a client or user said, such as a review, quote, or case study, on your own site and across those other accounts.

    It’s a step up from first-party because the substance is no longer solely your own assertion, even though you’re still the one choosing it and publishing it.

    Third-party: They prove you

    A third party publishes in its own voice, on its own site or social accounts, or on a neutral platform such as Trustpilot, with no involvement from you. 

    Think clients and partners sharing their experiences, journalists, analysts, academics, and the long tail of user-generated content that assistive engines lean on.

    It’s the strongest evidence because you had no hand in creating it.

    You can’t write that third tier, but you can feed it. Your clients publish because you’ve served them well enough that they want to, so earn it.

    Independent publishers can’t see inside your business, so give them something to work with: a client story they can build on, a view into your operation, or data about your business and industry they can cite.

    Giving outside parties a true, detailed version of your business to publish is what PR, marketing, and content teams have always done. The only thing that’s changed is that now you do it so machines read the result as proof, not just so humans read it as coverage.

    Point all three tiers at the same picture — you, your audience, and the independents — and they align into one answer the machine can’t miss.

    Author x Publication

    Read the grid by how much of you is in the publication.

    • First-party is all you. Your words on your own site. It’s pure claim, and the machine treats it as the baseline because you wrote it and you published it.
    • Third-party is none of you. Someone else’s words on a platform you don’t control. That’s why it’s the strongest proof.
    • Everything in between is second-party corroboration. Your own words carried onto an account you run elsewhere, or someone else’s words that you chose to publish on your own page.

    The same review is second-party when you surface it on your site and third-party when the client publishes it on their own account. The words are identical. The weight is different. The difference is determined entirely by who publishes it.

    Step back, and you have a powerful loop: You harvest your operations, codify them into a single source of truth, and distribute them across the tiers machines read. Then the machines recommend you, your ICP arrives, and serving them generates the next round of operations to harvest.

    Each turn feeds the next, so your digital footprint compounds instead of resetting.

    A simplified version of the flywheel

    The mirror principle is why this is the whole game

    When an AI engine recommends a brand, think of it as an impartial broker. Much as a travel agent carries every airline or a mortgage broker has the whole market on screen, an AI engine carries every brand in your category and recommends whichever it judges to be the best solution for the person asking.

    That impartiality is why buyers trust it. It’s also why the engine recommends your competitor without hesitation. It was never on your side. It’s on the buyer’s.

    That’s good news once you see it the right way. An AI engine can only recommend what it clearly understands and trusts. You don’t need to trick a rigged system. You need to provide the clearest, most complete picture of who you are, what you do, who you serve, and why you’re the right fit.

    Build a clearer, better-corroborated case than your competitors, and, on merit, you become the name the engine reaches for throughout the funnel. Many brands aren’t losing because they’re being outspent. They’re losing because the picture AI has of them is incomplete.

    And that picture comes from your digital footprint. AI forms its view of you from the world’s view of you: the reviews, coverage, and corroboration scattered across the market. What it shows about you is its opinion of the world’s opinion of you. That’s the mirror principle.

    You can try to flatter the system, trick it, or lean on it, and that might work for a while. But the approach that lasts is changing what the world can see. When you do that, you’re not manipulating anything. You’re providing proof: something that was always true, but underrepresented or invisible.

    That’s exactly what this article has laid out. Harvest the five streams, organize and codify them into a single source of truth, and distribute them across the channels AI reads. Do that, and you’ve provided the fullest, truest, and best-corroborated picture of your business at the moment that matters most: when someone is looking for what you sell, and AI is deciding what to recommend.

    Do it consistently, across everything AI can see, and you shape how it understands your business over time.


    This is the 17th piece in my AI authority series.

    Read more at Read More

    Web Design and Development San Diego

    What server logs reveal that SEO tools miss

    What server logs reveal that SEO tools miss

    For large websites, server logs often reveal technical SEO problems long before rankings decline. They show how search engines crawl your site, where crawl budget gets wasted, how quickly servers respond, and whether important pages remain accessible.

    Unlike Google Search Console, analytics platforms, and third-party crawlers, server logs capture every request search engines make to your infrastructure. 

    Yet many organizations never analyze them — missing one of the most valuable sources of technical SEO data available.

    Why server logs reveal what other SEO tools miss

    Many SEO teams rely on Google Search Console, Bing Webmaster Tools, third-party crawlers, and analytics platforms. Those tools help, but they all rely on data samples, delayed reporting, or simulated crawls. 

    Server logs capture direct interactions between crawlers and infrastructure. That distinction matters on websites with hundreds of thousands or millions of URLs.

    A log file records every request processed by a server. For SEO purposes, the most useful entries come from crawlers such as Googlebot, Bingbot, GPTBot, Applebot, and other verified search engine bots. 

    Each request generates operational data, including the requested URL, response code, timestamp, user agent, and response timing. Over time, those records form a detailed crawl history.

    Your customers search everywhere. Make sure your brand shows up.

    The SEO toolkit you know, plus the AI visibility data you need.

    Start Free Trial
    Get started with

    Semrush One Logo

    Hidden SEO issues in crawl data

    Most technical SEO issues begin as crawl inefficiencies that gradually compound over time. A search engine crawler may:

    • Request a page and receive an unexpected response.
    • Encounter a category section that slows under heavy load.
    • Follow redirect chains that expanded after a deployment. 

    In other cases, product pages disappear from inventory while still returning a 200 status code. These problems rarely occur as isolated incidents. 

    Search engines encounter them repeatedly across thousands or millions of crawl requests, creating patterns that can quietly erode crawl efficiency, indexing, and visibility.

    Server logs expose those patterns clearly. 

    • On large ecommerce platforms, logs often show crawlers spending excessive time on filtered navigation URLs while strategic product pages receive limited recrawling. 
    • On publisher websites, crawlers sometimes revisit outdated archive paths more aggressively than newly updated content. 
    • SaaS platforms frequently expose staging environments or parameter-driven duplicate URLs through internal systems without realizing how heavily those URLs consume crawl activity. 

    Without logs, those problems remain hidden behind aggregate reporting.

    Server logs also provide historical visibility. Unlike Google Search Console data, which expires over time, retained logs reveal crawl trends tied to migrations, infrastructure changes, indexing shifts, and platform redesigns.

    Where crawl resources go

    Search engines don’t crawl every page equally. Large websites compete internally for crawl attention. 

    Search engines allocate resources based on perceived importance, internal linking, infrastructure quality, content freshness, and historical performance. Logs reveal those crawl decisions directly.

    A retailer with five million URLs may assume high-value category pages receive regular crawling because they appear in XML sitemaps and navigation systems. Log file analysis may show Googlebot spending a disproportionate share of crawl resources on parameterized URLs created through faceted filtering instead.

    Another site may discover crawlers revisiting redirected legacy URLs years after a migration. These situations are common because search engines work from observed behavior rather than internal assumptions.

    Server logs also help identify sources of crawl waste that quietly consume large portions of crawl activity. Common examples include:

    • Infinite URL combinations.
    • Session parameters.
    • Crawlable internal search pages.
    • Open faceted navigation systems.
    • Duplicate mobile URLs.
    • Exposed staging environments.
    • Broken canonical structures. 

    As web platforms expand over time, crawl efficiency increasingly becomes an infrastructure challenge as much as a traditional SEO problem.

    When infrastructure limits crawling

    Response timing data is among the most valuable information in server logs. Search engines monitor how efficiently servers respond during crawling. Slow or unstable infrastructure affects how aggressively crawlers move through a site.

    A difference between 300 milliseconds and 3 seconds may appear minor on a single request, but across hundreds of thousands of crawler requests, the impact becomes substantial. Response timing analysis helps isolate infrastructure bottlenecks under real crawl conditions and exposes performance issues that traditional SEO tools often miss.

    In production environments, these patterns appear frequently. Product pages may bypass cache layers and generate database-heavy responses, image optimization services can slow down media crawlers, and API-driven templates often create inconsistent latency during crawl spikes. JavaScript rendering systems may delay crawler access to content, while regional CDN routing can introduce performance issues in specific markets.

    Synthetic monitoring tools often miss these patterns because simulated testing doesn’t fully replicate crawler behavior. Logs capture what crawlers experience at the request level. Timing analysis also helps separate isolated incidents from persistent operational issues.

    A temporary deployment issue differs from a structural bottleneck. Logs reveal the difference through historical request patterns.

    Search engines, particularly Google, tend to reward reliable infrastructure with more consistent crawling. Fast, stable responses support efficient crawl allocation and improve recrawl frequency on important pages.

    On enterprise systems, response timing analysis frequently influences infrastructure planning beyond SEO. Operations teams use log data to prioritize cache improvements, CDN adjustments, scaling decisions, and deployment scheduling.

    Get the newsletter search marketers rely on.


    Soft 404s at scale

    Soft 404s remain one of the most overlooked yet highly consequential SEO issues for large online brands. Unlike a standard 404 page, which correctly returns an HTTP 404 status code, a soft 404 returns a 200 OK response while serving thin, empty, or functionally useless content.

    To search engines, these pages appear crawlable and indexable despite offering little or no value, which can quietly waste crawl budget and dilute overall site quality signals.

    Common soft 404 examples include:

    • Out-of-stock product pages that remain live without meaningful replacement content.
    • Empty category templates created through faceted navigation.
    • Broken internal search result pages.
    • Placeholder inventory URLs with little usable information.
    • Expired listings that still return a 200 OK status code. 

    Failed rendering can create similar issues when JavaScript content doesn’t fully load for crawlers. On large web platforms, these low-value pages often accumulate quickly and consume significant crawl activity without contributing meaningful search visibility.

    Search engines eventually classify many of these pages as low quality. The issue becomes operational when crawlers continue revisiting those URLs repeatedly. Document size analysis within logs provides one way to identify potential soft 404 patterns at scale.

    Landing pages with nearly identical response sizes can sometimes indicate templated low-value responses. A group of 60,000 product URLs all returning responses smaller than 100 bytes after inventory expiration usually points toward placeholder templates rather than meaningful content.

    Internal search systems create another common example. Empty search result pages often generate highly consistent response sizes because the template loads correctly while no actual content appears.

    Response codes alone rarely expose the full pattern of crawl behavior. A clearer operational picture emerges when HTTP status codes are analyzed alongside response sizes, crawl frequency, and URL patterns. Together, these signals reveal how search engines interact with different sections of a web platform and where crawl inefficiencies begin to accumulate.

    Large publishers, such as news websites, also encounter soft 404 issues through broken pagination systems or empty archive states. 

    SaaS platforms sometimes expose onboarding placeholders through crawlable public URLs. 

    Marketplace websites frequently generate thin pages for inactive listings while still returning successful responses. Document size analysis helps identify these patterns quickly across large datasets.

    The case for log retention

    Short log retention periods limit the quality of server log analysis. Many crawl patterns develop gradually, with search engines adjusting crawl allocation over weeks or months rather than days. 

    Historical log data reveals long-term shifts in crawl behavior, including:

    • Changes in crawl frequency.
    • Legacy URL activity.
    • Migration effects.
    • Infrastructure instability.
    • Seasonal crawl patterns.
    • Redirect persistence.
    • Broader crawl budget fluctuations.

    For large websites, six to 36 months of logs often provide meaningful operational history.

    Historical data is especially valuable during migrations. Teams compare crawler behavior before and after structural changes to determine whether important sections gained or lost crawl visibility. Without retained logs, those comparisons disappear permanently.

    Many organizations still overwrite logs quickly or don’t retain them at all. Once lost, historical crawl data can’t be reconstructed later.

    Separating search crawlers from bot noise

    Raw server logs contain large volumes of automated traffic unrelated to SEO. Many bots impersonate Googlebot or Bingbot, making accurate filtering essential before meaningful analysis can begin. Effective validation typically combines user agent analysis, reverse DNS checks, and trusted IP verification to separate legitimate crawlers from scrapers, monitoring systems, and malicious automation.

    Once filtered correctly, server logs reveal clear behavioral differences between crawler types, including Googlebot Smartphone, Googlebot Image, Bingbot, Applebot, AdsBot, and newer AI-oriented crawlers. Each interacts with web platforms differently, creating distinct crawl patterns, resource demands, and indexing behavior.

    Image crawlers place heavier demands on media infrastructure. Mobile crawlers focus more heavily on rendering consistency. AI-focused crawlers often revisit large archive sections repeatedly.

    Crawler segmentation helps technical teams prioritize infrastructure improvements based on actual crawl demand rather than assumptions.

    Monitoring migrations with log data

    Migrations are one of the highest-risk periods in technical SEO, as even well-tested launches can introduce crawl instability. 

    Server logs provide direct visibility into how search engines respond after deployment, including which redirects crawlers continue to follow, whether redirect chains form, which legacy URLs remain active, and where 404 spikes occur. 

    Logs also reveal how crawl allocation shifts across the platform, whether response times begin to deteriorate, and which sections search engines continue to prioritize after the migration goes live.

    A migration may appear successful during browser testing while crawlers encounter entirely different behavior through caching systems, CDN routing, or redirect logic.

    Large ecommerce migrations often reveal persistent crawl activity on old URL structures weeks or months after launch. International platforms sometimes discover regional redirect inconsistencies affecting only certain crawlers. Logs expose those failures early enough to correct them.

    Collecting the right log data

    Useful log analysis depends on complete records. At a minimum, logs should include:

    • Remote IP address, including originating IP and optional (X-)Forwarded-For information.
    • User agent string.
    • Request protocol, such as HTTP, HTTPS, or WSS.
    • Request hostname.
    • Request path.
    • Request parameters.
    • Request time, including date, time, and time zone.
    • Request method.
    • Response HTTP status code.
    • Response timings.

    These fields create the operational baseline required for meaningful crawl analysis.

    Hostname and protocol fields often receive less attention than they deserve. Missing these values creates blind spots on multilingual websites, subdomain-heavy platforms, and CDN-driven architectures.

    Many organizations simplify analysis by storing the full request URL as a normalized field containing protocol, hostname, path, and parameters.

    Additional fields can further improve analysis quality:

    • Response byte size.
    • Cache status.
    • Referrer.
    • CDN edge location.
    • Upstream timing.
    • Compression type.

    Response size data becomes especially valuable during soft 404 investigations and duplicate content analysis.

    Why logs remain underused

    Server logs often fall between departments. Infrastructure teams view them as operational data. Security teams use them for threat monitoring. SEO teams focus on crawling and indexing. Analytics teams prioritize user behavior reporting.

    As a result, one of the most valuable technical SEO datasets within an organization often remains completely unused. Yet server logs answer operational questions that few other systems can.

    They reveal which pages absorb the largest share of crawl resources, which sections return unstable responses, and which deprecated URLs continue receiving heavy crawler activity years later. 

    Logs also expose latency issues affecting specific crawler groups and low-value pages that dilute crawl efficiency. These insights directly influence rankings, crawl allocation, and search visibility.

    Technical SEO and GEO increasingly overlap with infrastructure engineering because search engines continuously evaluate operational quality. Server logs expose those operational realities in detail. 

    For large websites, log analysis stops being optional once crawl scale reaches enterprise complexity. The data already exists. The advantage comes from retaining it, structuring it properly, and using it consistently.

    See the complete picture of your search visibility.

    Track, optimize, and win in Google and AI search from one platform.

    Start Free Trial
    Get started with

    Semrush One Logo

    The business value of server logs

    Ultimately, server log retention delivers value far beyond SEO alone. In particular, preserved log data can strengthen buyer confidence by providing verifiable operational evidence of site performance, infrastructure stability, and historical activity. 

    That additional transparency can materially support due diligence and even contribute positively to company valuation, making a compelling case that the cost of recording and retaining server logs is often outweighed by their long-term strategic value.

    Read more at Read More