Buying AI Ads Is Easy. Winning Them Isn't.
OpenAI Ads Market Report
The money is already moving. Most of it is moving blind.
Getting an AI ad live takes an afternoon. Knowing where to compete, what to say, how to localise it and how to win consistently is a different problem entirely.
Based on Whitebox's own monitoring of the OpenAI Ads surface: 153,514 sponsored placements observed across 7,171 questions, 3,352 advertisers and more than 200 brand categories in the United States, Canada, Australia, Japan, South Korea, Brazil and Mexico, between April and August 2026.
Key Takeaways
Are brands actually advertising inside ChatGPT yet?
Mostly not. Only 16.4% of the brands ChatGPT already names in its answers have ever run a single ad on the surface beneath. In 48 categories, not one of the brands the AI names is buying the slot underneath.
How much competition is there on a single question?
Less than most people assume. The median question draws 7 advertisers, and 17.3% of monitored questions have had exactly one advertiser ever appear.
Why is performance so hard to read on this surface?
Because position rotates. Ask the same question twice and 84.2% of the time a different advertiser comes back, and 99.9% of ad-bearing answers carry exactly one ad. A short window of delivery data cannot separate a weak result from ordinary rotation.
What is the most avoidable mistake in the data?
Absence from your own name. 61.1% of questions naming a brand did not surface that brand's own ad, even though the brand advertises on the platform.
Is the creative any good?
Rarely. Wrong language, one headline forever, generic landing pages and absent brand defence recur across thousands of advertisers. 97.1% of ads served against non-Latin-script questions came back in Latin script.
Do advertisers stick around?
Most do not. 67.3% ran fewer than ten ads in total, and the median advertiser with real volume was active on five days.
So where is the opportunity?
Not in low-competition intent as such. It is conversational context with meaningful demand, lower-than-expected competition, and a defensible relevance advantage.
Executive Summary
ChatGPT advertising has been live in production for months. Thousands of brands are already running AI ads. Most are doing it without visibility into the competition around them.
A budget, a headline and a landing page puts a brand live within the hour. What an advertiser cannot see from the buying and delivery view is the competitive environment around it: whether a question drew one rival or fifteen, which questions are under-contested relative to demand, and where the brand is missing from questions that name it directly.
The result is a market with two features that rarely coexist: a great deal of unclaimed inventory, and a great deal of poorly executed advertising.
What the Data Shows
| Most brands have not shown up. Only 16.4% of the brands ChatGPT already names in its answers have ever run a single ad on the surface beneath. | 16.4% |
| Competition per question is thin. The median question draws 7 advertisers, and 17.3% of monitored questions have had exactly one advertiser ever appear. | 17.3% |
| Position rotates. Ask the same question twice and 84.2% of the time a different advertiser comes back. | 84.2% |
| Brands are absent from their own name. 61.1% of questions naming a brand did not surface that brand's own ad, even though the brand advertises on the platform. | 61.1% |
| Execution is weak across the board. Wrong language, one headline forever, generic landing pages and absent brand defence recur across thousands of advertisers - quantified in section 5. | see §5 |
| Most advertisers do not stay. 67.3% ran fewer than ten ads in total; the median advertiser with real volume was active on five days. | 67.3% |
1. A Live Market Almost Nobody Can See Into
1.1 How Advertising Works Inside ChatGPT
This is not a search results page. The user asks a question in natural language; ChatGPT composes an answer that names brands, weighs trade-offs and frequently reaches a recommendation. Only then does a single sponsored placement appear beneath it, labelled and separated from the organic response.
That sequencing is the whole game. By the time the ad renders, the organic answer has already told the user what to think. In our sample, 99.9% of ad-bearing answers carried exactly one ad - there is no second position to settle for, which is why position rotates rather than accumulates.
1.2 Who Sees the Ads
Ads are shown to users on the Free and Go tiers. Plus, Pro, Business and Enterprise remain ad-free. Ads are not shown to minors, and are not eligible to appear near sensitive or regulated topics including health, mental health and politics. The rollout began with adult users in the United States.
A note on scale. Third-party estimates put ChatGPT at roughly 900 million weekly users, with the free tier around 81% of the active base. OpenAI publishes no tier-level breakdown, and we could not substantiate a "90%+ of users" figure, so this report does not use one. What is defensible: the ad-eligible tiers are the free tier plus a low-cost paid tier.
Tier eligibility, ad format and topic exclusions are sourced to press coverage of OpenAI's January-February 2026 announcements (Axios, MacRumors, Forbes, Analytics Insight). User-base figures are third-party estimates, not OpenAI disclosures. Everything else in this report is measured directly from Whitebox monitoring.
1.3 Early, but Filling
We observed 3,352 distinct advertisers, with roughly 40% of those active in any given week appearing for the first time. Weekly ad volume grew from about 1,400 to 22,000 between mid-May and late July. The market is early and structurally underpenetrated - and the competitive window is narrowing.
2. The Opportunity Most Brands Are Missing
2.1 Only One Brand in Six Has Shown Up
Across more than 200 categories, Whitebox tracks 330 named competitor brands - the companies that appear in the organic answers. Only 54 of them, 16.4%, have ever been observed advertising. In 48 of those categories, not one of the brands the AI names is buying the slot underneath.
2.2 Competition per Question Is Thin
The median question draws 7 advertisers, and even where one leads, the median leader holds only 33.3% of that question's impressions. Which of these contexts are actually worth competing in is the subject of section 3.2.
2.3 Participation Is Unstable
Two thirds of advertisers never got past nine ads. Among those that reached real volume, the median advertiser was active on just five days, and 13.0% appeared on a single day and never returned. Most brands are not losing this market - they are leaving it.
3. Why Brands Cannot Find the Opportunity
This is the heart of the problem. The inventory is open and the mistakes are visible - but not from inside an ad account, which reports what you spent and what you received, never the competitive environment around it.
3.1 Position Rotates on Every Ask
Of 21,465 questions we asked two or more times, 84.2% returned a different advertiser on the repeat. On Google a brand buys a keyword and holds it; here there is nothing to hold. The consequence is that a short window of delivery data cannot separate a weak result from ordinary rotation.
3.2 Choosing Where to Compete
Delivery here is not decided by bid alone. An ad appears when it is both commercially competitive and relevant to the conversational context the user has created. That is two levers - and the second is the lever that does not depend on outspending the field.
So the question is not "what should I bid?" but "which conversational contexts should I compete in?"
The Relevance Lever
That is where relevance becomes a practical lever. Three positions may offer a structural relevance advantage:
| Contexts naming your brand - 61.1% of brand-name questions do not surface that brand's own ad. | 61.1% |
| Contexts in the user's language - 97.1% of ads against non-Latin-script questions came back in Latin script. | 97.1% |
| Contexts matching a stated qualifier - a generic competitor has a weaker relevance claim to that specific constraint. | - |
Competition is not evenly spread across them:
| Intent type | Questions | Median advertisers | With ≤3 advertisers | Share of ad volume |
|---|---|---|---|---|
| Sentiment / reviews | 1,729 | 8 | 29.6% | 22.2% |
| General | 611 | 7 | 34.5% | 8.5% |
| Comparison | 707 | 6 | 36.2% | 8.0% |
| Unbranded / category | 5,281 | 6 | 32.9% | 59.7% |
Questions are counted per tracked context, so a question monitored in more than one category or market appears more than once; the column therefore sums above the 7,171 distinct questions in the sample. The four intent types shown exclude three minor categories accounting for 1.5% of observed ad volume.
Comparison stands out. It is the context closest to a decision, yet it draws fewer advertisers than reputation and sentiment contexts - a median of 6 against 8, with 36.2% of comparison questions drawing three advertisers or fewer. Strong decision proximity alongside below-average competitive density is an unusual combination. Comparison contexts account for just 8.0% of observed ad volume.
Competition also falls as the user supplies more context: questions of one to six words draw a median of 10 advertisers, those of eleven to fifteen draw 6. Two real questions from the sample: "Best project management tools for small teams?" drew 13 advertisers. "Which project management platform scales from a small team to a larger company?" - the same underlying need, more context supplied - drew 4. The second is not the better target because it is longer; it is the better target for the advertiser that genuinely fits it. Length is only a proxy for supplied context, not a targeting strategy.
Persona framing does not work the same way. "Which smart home platforms integrate easily with existing systems for small businesses?" drew 15 advertisers, and "Smart home solutions designed for enterprise-level security and employee productivity" drew 15 - the persona changed completely and the field did not move, against an observed range of 1 to 39 advertisers per question. Across the sample, persona-framed contexts drew a median of 7 advertisers, the same as generic ones. Persona is not a route to a thinner field; it is worth targeting when the advertiser is genuinely better suited to that persona than the incumbents are.
Specificity and persona are not shortcuts to lower competition. They matter when they sharpen a relevance advantage.
The constraint that governs all of it: across 7,171 questions, advertiser count and ad volume correlate at 0.87. Uncontested usually means small - of the 2,732 questions with three advertisers or fewer, only 79 carried meaningful volume. What matters is not low competition, but competition that is low relative to demand.
The Test
Score a conversational context on three things:
- Meaningful demand - is this context large enough to matter?
- Competition relative to demand - is advertiser density lower than the volume would predict?
- Defensible relevance advantage - is there a concrete reason you should be especially relevant here: a brand mention, the user's language, a stated qualifier, a use case you fit?
Method note: competitive density is measured as distinct advertisers observed per question; specificity is proxied by question length. OpenAI has not published its ad-ranking mechanics, so the relevance advantages described here are inferred from observed delivery patterns rather than from a disclosed ranking formula.
3.3 Brands Are Missing From Questions About Themselves
Across 355 brand-and-question pairs where the brand is a known advertiser on the platform, the brand's own ad was missing 61.1% of the time. Competitors appear in those answers unopposed - the most avoidable failure in the dataset, and invisible unless somebody is watching the questions rather than the campaign.
3.4 What an Ad Platform Cannot Tell You
- Whether you were competing against one advertiser or fifteen.
- Whether a weak result reflects the competitive field or ordinary rotation.
- Which questions in your category are uncontested, and which carry decision intent.
- Where you are absent from questions that name your own brand.
4. Worked Example: Ford
What that visibility looks like in practice, from a Whitebox workspace tracking Ford in the United States. On the headline numbers the brand is in good shape: 26.2% share of voice, the number one advertiser in its category, first position in five of its eleven tracked ad groups, and a total ad rate of 69.7% - meaning roughly seven in ten questions in this category now return an ad at all.
The category view is where it becomes useful. Sorted by Ford's own share, one pattern runs the length of the table:
| Ad group | Ad rate | Ford share | Opportunity | Leading advertiser |
|---|---|---|---|---|
| Ford vs. General Motors | 82.5% | 48.1% | Medium | Ford 48.1% |
| Ford vs. Tesla | 67.5% | 47.6% | Medium | Ford 47.6% |
| Performance, Off-Road & Adventure | 73.0% | 40.3% | Medium | Ford 40.3% |
| Ford vs. Stellantis | 75.0% | 39.3% | Medium | Ford 39.3% |
| Sentiment: Ford | 72.0% | 35.5% | Medium | Ford 35.5% |
| Trucks & Towing Capability | 64.0% | 23.6% | Medium | Capital One 30.9% |
| Electric Vehicles | 69.0% | 21.0% | Medium | Volkswagen 38.7% |
| Sentiment: General Motors | 76.0% | 13.2% | Low | Volkswagen 23.7% |
| Sentiment: Tesla | 72.0% | 9.4% | Medium | CarMax 25.0% |
| SUVs & Family Vehicles | 62.2% | 8.7% | Medium | Capital One 28.3% |
| Sentiment: Stellantis | 66.0% | 7.4% | Low | Volkswagen 25.9% |
Ford share of voice, by ad group
Ford wins where its own name is in the question, and loses where the user describes a need. It holds 48.1% against General Motors, 47.6% against Tesla and 35.5% on questions about its own reputation. But on Trucks & Towing Capability - Ford's core product territory - it holds 23.6% and the leading advertiser is Capital One, a financial services company. On SUVs & Family Vehicles Ford holds 8.7% against Capital One's 28.3%. Volkswagen leads Electric Vehicles at 38.7% and leads the conversation about two of Ford's three named rivals.
Read against section 3.2, this is the framework in one page. Ford has a defensible relevance advantage wherever it is named, and translates that advantage into higher observed share of voice. Where the context is a need rather than a brand - a family shopping for a three-row SUV, a contractor sizing towing capacity - that advantage does not carry, and advertisers outside the category take the position instead.
The commercially significant gaps are not the ad groups Ford is losing narrowly. They are the ones carrying real demand where its share sits in single digits, and where the advertiser ahead of it is not a carmaker.
Figures in this section are a product snapshot from the Whitebox Ads Research view for Ford, United States, over a two-week window, and use the platform's brand-level attribution. They therefore differ in window and grouping from the sample-wide figures used elsewhere in this report.
5. Where Execution Breaks Down
The market stays winnable partly because the brands already in it treat it like a display network. Five failure modes repeat across thousands of advertisers.
| Failure mode | Frequency | Basis |
|---|---|---|
| The ad is in the wrong language | 97.1% | of 3,740 ads served against non-Latin-script questions came back in Latin script |
| The brand is absent from its own name | 61.1% | of 355 brand-and-question pairs |
| The click lands somewhere generic | 52.0% | 21.2% on a bare homepage, 30.8% on a category page (n = 129,558 ads with a URL) |
| One headline for every impression | 17.3% | of advertisers with 20+ ads; 27.5% use one headline on 90%+ of their ads |
| Show up, then vanish | 5 days | median active life among advertisers with real volume; 67.3% ran fewer than ten ads |
Each frequency rests on its own basis, shown alongside it; the five are not measured against a common denominator and are not directly comparable.
Among advertisers with 20 or more ads the median has run 4 distinct headlines; across all 3,352 advertisers the median is 1. More than half of ads with a landing URL point at a homepage or category page - the user asks about one product and arrives at a search box.
5.1 Three Real Placements From the Sample
"What are the main pros and cons that users report about Fiverr?"
Six advertisers appeared - ZoomInfo, Payoneer and Jobber among them. Fiverr was not one of them.
"How does Hyundai stack up against Toyota?"
The closest thing to a purchase decision in the category. Neither carmaker showed. The ad came from a local dealership group.
A Japanese-language question about forex trading for beginners
It was served an English-language ad from a code editor: "The Leading AI IDE." Seventeen times.
6. Where the Opportunity Differs
6.1 B2B Has Arrived. B2C Has Not - and Neither Is Executing Well
Adoption has been led by software - and the segment with the most competition has the weakest execution.
| Measure | B2B | B2C | Mixed |
|---|---|---|---|
| Categories | 38 | 47 | 8 |
| Ads observed | 74,562 | 55,780 | 13,586 |
| Distinct advertisers | 1,516 | 1,498 | 624 |
| Median advertisers per question | 10 | 3 | 12 |
| Questions with a single advertiser | 6.5% | 29.9% | 3.1% |
| Median category leader share | 22.3% | 40.7% | 25.2% |
| Categories with a majority leader | 2 of 32 | 11 of 30 | 1 of 8 |
| Median headlines per advertiser (20+ ads) | 3 | 6 | 4 |
| Lands on a bare homepage | 28.1% | 12.8% | 21.2% |
| Missing from own-brand questions | 63.3% | 52.2% | too few |
Median advertisers per question, by segment
Selling to consumers: far less contested - a median of three advertisers per question, and 29.9% of questions with only one advertiser ever. But 11 of 30 measured B2C categories already have a majority holder, so the window is narrowing.
Selling to businesses: presence alone achieves little. Ten advertisers on the median question means the advantage lies in what you say and where you point it - and B2B runs half the creative variation of B2C with more than twice the homepage landings.
Segment is an analyst classification of the monitored categories, not a database field. Eight genuinely ambiguous categories - marketplaces, SMB fintech, prosumer tools, home services - are held in a separate Mixed bucket rather than forced either way.
6.2 The Answer Adapts. The Ad Often Doesn't.
One enterprise security category, one 31-question set, translated and run in three markets:
| Market | Question language | Ads | Advertisers | Leading advertiser | Leader share |
|---|---|---|---|---|---|
| United States | English | 858 | 76 | SentinelOne | 21.0% |
| Japan | Japanese | 747 | 34 | Shinobi Security | 28.5% |
| South Korea | Korean | 497 | 29 | SuperGRC | 47.5% |
Fewer advertisers, harder concentration
Distinct advertisers
Leading advertiser share
Japan and Korea drew less than half the advertisers the United States did for an identical question set, and concentrated more than twice as hard. No advertiser appeared in both the US and either Asian market.
We had expected ads to stay global while organic answers localised. In this sample the roster does the opposite - advertiser sets diverge across borders more sharply than the organic answers do:
| Comparison | Organic brands in both | Advertisers in both |
|---|---|---|
| United States vs Japan - enterprise security | 21.0% | 0% |
| United States vs Australia - devtools, both English | 38.5% | 5.4% |
The roster changes by geography. The creative frequently does not. In the Japanese categories we monitor, only 3.4% of ads were written in Japanese and only 3.1% pointed at a Japanese-localised landing page; in Korea both were 0.0%. The question is localised, the answer is localised, the ad is not.
Want the full report as a PDF?
22 pages, every figure and the methodology behind it.
7. What Whitebox Does About It
Everything above is invisible from an ad account and visible from the surface itself. Whitebox monitors the questions, the organic answers and every ad served against them - so a brand sees the competitive environment it operates in, not just the invoice.
7.1 See Where Your Competitors Are Advertising
- Share of voice by intent. Every intent in your category as a share-of-voice bar: who holds it, how much, and where you sit. In the automotive example above, the Electric Vehicles intent carried 62 observed ads split across nine advertisers.
- Their actual creative. The headline, description and image a rival ran, with their share of appearances - you read a competitor's positioning directly.
- Ad rate over time. How often any ad appears against your category's questions, so you can watch inventory fill before it is full.
7.2 Find Where the Opportunity Is Strongest
- Ad group opportunities. Every category and competitor intent scored on ad rate, your share and an explicit opportunity level, with the top three advertisers beside it - as in section 4.
- The opportunity map. The gaps that matter - meaningful demand with lower-than-expected competitive density - are visible at a glance.
- Under-contested question detection. Questions and contexts where advertiser density is low relative to observed demand are surfaced as targets, rather than simply prioritizing places where nobody is competing.
7.3 Know What to Say, Not Just Where to Say It
- Creative briefs written from your organic position. Whitebox reads what the AI already says about you before writing the ad, quoting the organic answer about you and about each rival side by side. The generated brief for Ford's comparison ad group reads: "I chose to highlight specific engineering advantages like Pro Power Onboard and Super Duty towing capacity to address user intent for reliable work vehicles. I also leveraged the Explorer's IIHS safety rating to appeal to family buyers. These points directly contrast with reported powertrain concerns and reliability ratings for competitor trucks."
- Context hints from real questions. Each ad group carries the actual phrasings users bring - "tradespeople comparing work truck reliability", "contractors seeking jobsite power solutions", "parents researching top safety rated SUVs" - so creative answers real demand, not a keyword guess.
- Localisation built in. Creative generated in the language of the question and pointed at the matching local page, addressing the 97.1% failure directly.
7.4 Test It, and Track Whether It Worked
- A/B testing that survives rotation. Where 84.2% of repeat asks return a different advertiser, single-campaign readings are noise. Controlled A/B tests let a performance difference be attributed to the creative rather than to rotation.
- Campaigns built from the finding. Opportunities move to execution in one step - campaigns, ad groups and ads assembled as drafts, published by the customer.
- Mention rate against ad share of voice. Organic presence in the answer and paid presence beneath it on one timeline - is paid covering an organic weakness, or duplicating a strength?
- Brand defence monitoring. Continuous checking of the questions that name your brand, so the 61.1% failure becomes an alert rather than a discovery.
8. Methodology and Limitations
The Sample
153,514 sponsored placements captured across 7,171 distinct questions, 3,352 advertisers and more than 200 brand categories belonging to 69 tracked customers in the United States, Canada, Australia, Japan, South Korea, Brazil and Mexico, between April and August 2026.
- A convenience sample, not a census. Categories are the brand contexts Whitebox customers pay to monitor; verticals no customer operates in are absent entirely. All figures describe the monitored sample.
- Non-US observations carry a caveat. OpenAI's rollout began with adult users in the United States. Our Japan, Korea, Canada, Australia and Mexico observations reflect monitoring configured to those locales; we cannot confirm how OpenAI's serving geography maps to them. Mexico returned 4 ads and is excluded.
- The dataset grows daily. Percentages are stable; absolute counts move. All figures are as of 6 August 2026.
What This Report Does Not Measure
- No spend, bid or auction data. Nothing here measures cost. "Share" always means share of observed impressions, never share of spend. Categories are described as less contested or open - never as cheap. Lower observed competition does not prove lower cost; that would require bid-level data we do not hold.
- No click or conversion outcomes. We observe what was served, not what worked. Claims about creative quality are structural - wrong language, one headline, generic destination - never performance-proven.
- No rendered creative. We capture headline, description and destination, but no rendered ad unit or image dimensions - so image-format problems such as mis-sizing or cropping cannot be quantified here, and are not claimed.
- Brand matching is exact. The 16.4% penetration figure undercounts brands advertising under a subsidiary or agency name, and is a floor rather than a point estimate.
- Geography rests on two categories. The zero US-Japan advertiser overlap held in a second, all-English category, but a broader multi-market panel is needed before generalising.
9. What to Do Next
This surface will not stay this open. Between 270 and 480 advertisers appear for the first time every week, and competition is increasing. The categories being claimed now are being claimed for a long time - and on the evidence above, not especially well.
Three questions worth answering about your own category this quarter:
- Where are you absent? Which questions in your category name your brand and return somebody else's ad?
- Where is competition low relative to demand? Which high-intent questions carry meaningful demand but remain under-contested?
- What does the answer above your ad actually say? And does your creative agree with it, in the language the user asked in?
Whitebox answers all three continuously - competitor share of voice by intent, opportunity scoring per ad group, creative written against the organic answer, A/B testing that survives rotation, and brand-defence monitoring.
See your own category.
Competitor share of voice by intent, opportunity scoring per ad group, creative written against the organic answer, and brand-defence monitoring - for the questions your buyers are actually asking.
Frequently Asked Questions
How was this data collected?
Whitebox monitors the OpenAI Ads surface directly: the questions, the organic answers and every ad served against them. This analysis covers 153,514 sponsored placements across 7,171 distinct questions, 3,352 advertisers and more than 200 brand categories belonging to 69 tracked customers, between April and August 2026. All figures are as of 6 August 2026.
Does the report say how much ChatGPT ads cost?
No. Nothing here measures cost. There is no spend, bid or auction data in the sample, so "share" always means share of observed impressions, never share of spend. Categories are described as less contested or open - never as cheap. Lower observed competition does not prove lower cost.
Does it measure clicks or conversions?
No. We observe what was served, not what worked. Claims about creative quality are structural - wrong language, one headline, generic destination - never performance-proven.
Who sees ads inside ChatGPT?
Ads are shown to users on the Free and Go tiers. Plus, Pro, Business and Enterprise remain ad-free. Ads are not shown to minors, and are not eligible to appear near sensitive or regulated topics including health, mental health and politics. The rollout began with adult users in the United States.
Why does position change every time I ask the same question?
Because there is only one slot and it rotates. 99.9% of ad-bearing answers carried exactly one ad, and of 21,465 questions asked two or more times, 84.2% returned a different advertiser on the repeat. The practical consequence is that a short window of delivery data cannot separate a weak result from ordinary rotation.
Is this a representative picture of the whole market?
It is a convenience sample, not a census. The categories are the brand contexts Whitebox customers pay to monitor, so verticals no customer operates in are absent entirely. Every figure describes the monitored sample. Non-US observations carry an additional caveat, and the geography findings rest on two categories.
See the questions, the answers and the ads in your own category - and what to do about them.


