Skip to content

13 min read

The New Bottleneck: You Can Now Produce Thousands of Creatives. You Can't Test Thousands of Creatives.

Why you need the inversion: treat production as infinite, tests as scarce, and learning as the asset on the balance sheet.

Bernard Bontemps · Founder, HyperScale

For a decade, the craft of Meta advertising was targeting. You won by knowing your buyer better than the auction did: stacked interest audiences, lookalike ladders, exclusion lists tuned like race engines. The creative mattered, but it was the payload, not the guidance system.

Three things ended that era, in sequence.

App Tracking Transparency (2021) cut the signal that made granular targeting work.

Meta pulled targeting into the model. Broad audiences, Advantage+, “give us room and we’ll find them.” The levers you used to hold moved inside the black box.

Andromeda (December 2024) made the new architecture explicit: a retrieval engine rebuilt on NVIDIA and MTIA silicon to evaluate tens of millions of ad candidates in milliseconds, designed, in Meta’s own words, to handle “the exponential growth of creatives.” Personalization moved into the machine’s selection among your ads. The system stopped asking you who to reach. It started asking you for more raw material: more distinct creative concepts, so it could match each one to the pocket of people it suits.

“Creative is the new targeting” stopped being a conference line and became a literal description of the ranking stack. The audience you reach is now a function of the ad you made. The hook, the face, and the claimed benefit are the targeting parameters.

When creative became the targeting, volume became the bottleneck

If each distinct concept is a targeting instrument, then a team running four concepts a month is targeting four audiences a month. The fix looked obvious: make more, make them different, feed the machine.

So the volume race started, and for most of 2023 to 2025 production was the constraint everyone actually felt: briefs queued behind editors, UGC creators behind casting calls, iteration cycles measured in weeks.

Generative AI led the race out of that bottleneck, and the numbers got absurd. AppsFlyer’s 2025 creative report, built on 1.1 million video variations across 1,300 apps and $2.4B in ad spend, found high-spending non-gaming apps now averaging 2,365 creative variations per quarter, up 18% year over year and growing faster than gaming’s output. Apps spending $7M+ per quarter produce nearly three times more creatives than the tier just below them, and AppsFlyer attributes the non-gaming surge “largely” to AI adoption. Volume became table stakes.

By Meta’s own count, more than 8 million advertisers now use at least one of its generative AI creative tools, roughly double a year earlier, and adoption of its video-generation tools grew 20% in a single quarter. The production side isn’t just assisted anymore; it’s winning. At BoomBit, one of the more transparent mid-size studios, about half of the winning creatives are already AI-generated, at least on hooks. In categories where GenAI formats took hold, comparison-style AI creatives were pulling 40% of spend.

A competent two-person growth team can now produce five hundred variants a month for the price of API credits, and the pitch deck of every creative tool shows the same slide: output, up and to the right.

Here is the uncomfortable part: win rates didn’t move.

In that same AppsFlyer dataset, the top 2% of creatives still absorb 43% of total spend for non-gaming apps (53% in gaming). The distribution of winners is as brutal as it was before the volume explosion. There is simply a taller pile of creatives underneath it that never earned delivery. Producing more didn’t produce more winners, because production was never where winners were found. They’re found in testing. And testing didn’t get cheaper.

You can now produce thousands of creatives. You can’t test thousands of creatives.

Testing has physics that production doesn’t

Producing a creative is a compute problem, and compute got cheap. Testing a creative is a statistics problem, and statistics stayed expensive, because a fair test is paid for in conversions. Conversions cost money.

Start with the constraint Meta states outright: an ad set needs roughly 50 optimization events within a 7-day window to exit the learning phase. Below that floor, delivery is unstable and costs run high. The number you’re reading is closer to noise than verdict.

For a game optimizing to installs, 50 events is pocket change. For a subscription app optimizing to trial starts or purchases, it isn’t. RevenueCat’s State of Subscription Apps puts the median install-to-trial rate between 1.4% and 2.8% depending on price point, which means your optimization event sits 35 to 70 installs deep. The event you test against is structurally scarce.

Now price a fair read. If a creative gets its own cell, one week at the learning-phase floor costs about 50 times your trial CPA:

Cost per trial startOne 7-day readFair reads per month on $50k of test budget
$20~$1,000~50
$40~$2,000~25
$80~$4,000~12

Those are generous assumptions: one clean week per verdict, no re-tests, no waiting on downstream truth. Against a production stack shipping 500 variants a month, a $50k testing budget gives a fair read to somewhere between 2% and 10% of what you make. That gap is the new bottleneck. Production scaled a hundredfold. Test capacity scaled not at all.

And signal scarcity is only the first tax. Three more stack on top.

The delayed-truth tax. A trial start is a proxy. The event you actually care about, trial converting to paid, resolves days or weeks later, and the two regularly disagree. AppsFlyer’s report found the GenAI comparison-style formats that captured 40% of spend also had the lowest Day-7 retention in their categories. A verdict rendered on day 2 optimizes for trial collectors, not subscribers.

The privacy tax. On iOS, a large share of your signal arrives modeled, delayed, or thresholded through ATT-era measurement. Small test cells are exactly the ones that fall under privacy thresholds and report zeros. The smaller you slice your tests, the blinder you get.

The attention tax. A dead cell dies cheap: the auction starves losers on its own, and automated rules cap what any test can burn. The tax lands on the decisions no rule automates. Every concluded cell demands a read (what did it teach, which belief does it grade, what should the next round probe because of it), and a human team still tops out at a few dozen concurrent cells because conclusions pile up unread: winners scale with their why unextracted, and the next round waits on the last one’s paperwork. The latency doesn’t eat budget; it eats test capacity, which is the scarcer asset.

Buy signal at the altitude you can afford

Every number above prices a read in conversions, and that pricing feels like a law of nature. It isn’t. The event you optimize a test cell to is a choice, and it is the single biggest lever on what a read costs.

Every subscription marketer knows the first rung: optimize test cells to trial start, not purchase or renewal, because it is the highest-frequency event that still correlates with revenue. That is table stakes, and it is already priced into the table above. The less obvious move is that the ladder keeps going up. You can buy readable signal above the install, at attention altitude, where events arrive at impression scale and the learning-phase arithmetic barely applies.

Attention metrics alone (CTR, hook rate, hold rate) are cheap, fast, and dangerous: they select for clickbait, and the scroll-stopper that pulls the wrong crowd wins every one of them. The unlock is pairing them with a user-quality event you create yourself. Ask one question at onboarding, who are you, with your personas as the options, and fire the answer as a custom event. Route your web-to-app flow through your own domain so the event attributes per creative, and every test cell now reads as attention metrics qualified by who the creative pulled in.

When I ran growth at Mojo (YC W18), our onboarding question offered Creator, Business owner, Marketing professional, and Other. Business owners carried the highest LTV, so every test tracked one number next to the attention metrics: the share of business owners each creative recruited. I ran the same system again at Beside. We call that event a user-quality signal, UQS, and it is a first-class metric in every account our agents run.

Now do the physics again at this altitude. The quality event fires at onboarding, so it sits one onboarding deep instead of 35 to 70 installs; a fair read costs a multiple of your install CPA rather than your trial CPA, and the same $50k of test budget reads several times the cells. The verdict is a persona share you have already calibrated against LTV, not a day-2 trial count, so the delayed-truth tax shrinks instead of compounding. And because the event lands on your own domain’s pixel rather than inside thresholded SKAN reporting, small cells stop reporting zeros. None of the taxes disappear. Every one of them gets cheaper at the altitude where you have done the calibration work.

The calibration is the price of admission: UQS only reads as signal after you have mapped persona to LTV on your own cohort data. Do that once, and testing stops being rationed by your trial CPA.

Why not just upload everything and let Meta sort it out?

The tempting cheat: skip the discipline, flood the account, let the delivery system be your test bench. Meta half-invites this. Andromeda exists to pick from enormous slates, and GEM, the LLM-scale foundation model Meta shipped across its ads stack in 2025, lifted conversions about 5% on Instagram largely by getting better at exactly this kind of selection. The machine really is getting better at picking from whatever you give it.

But delivery is not testing, for three reasons.

Delivery concentrates instantly. Upload 100 ads and Meta will give meaningful impressions to a handful. The rest die at a few hundred impressions, which is statistically nothing. You didn’t run 100 tests. You ran five tests, plus 95 coin flips decided by a cold-start prior.

Near-duplicates collapse. The ranking system reads creative content. Thirty permutations of one concept occupy one semantic slot: they cannibalize each other in the auction and fatigue as a block. This is why Meta’s guidance keeps landing on one word, diversification. Genuinely different concepts for different people, not more versions of the same idea.

The platform explores for its objective, not yours. Meta’s job is the probability of a conversion in this auction. Your job is figuring out which belief about your market is true, so next month’s creatives are better than this month’s. The delivery system will happily tell you that ad #34 won. It will never tell you why, and a winner without a why doesn’t compound. “The insomnia angle beats the productivity angle for women 35+” generates your next ten creatives. “Ad #34 won” generates nothing.

Flooding gets you the worst of both: you pay cold-start spend on noise, and the signal you get back answers a question you didn’t ask.

Ration tests like the scarce resource they are

Accept the inversion (production is abundant, tests are scarce, learning is the asset) and the operating model follows. The teams doing this well run some version of five disciplines.

1. Test beliefs, not assets

The output of a creative test isn’t a winning ad; it’s a validated statement about your market. Structure the space explicitly, angle by audience by format, and make every creative a deliberate probe of one cell. Ten creatives probing ten distinct beliefs teach more than a hundred permutations of one belief, because beliefs compound: each verdict prunes the space every future round draws from. Assets don’t compound. They fatigue.

2. Make the market pay for your first filter

Before a concept deserves a test slot, it should survive two screens that cost you nothing.

First, the revealed answer key. In Meta’s Ad Library, every ad that has survived 30+ days of continuous spend in your category is a test someone else already paid for. We publish x-rays of these winning sets for top subscription apps, and reading your market’s survivors before briefing is the cheapest signal that exists.

Second, your own account history. Most accounts carry years of concluded experiments nobody has read as a corpus. Angles that have died three times will die a fourth; don’t rediscover them at $2,000 a discovery.

The same instinct drives operators like BoomBit to run $5 pre-tests in cheap geos before committing real budget: buy the cheapest available signal first. Proxies don’t validate. But they rank, and ranking before spending is the point.

3. Protect the read

Pick the highest altitude your calibration has earned (attention plus UQS where you have it, trial start where you don’t), then defend the cell. Cap costs so a bad cell can’t overspend its read. One variable per cell, or the verdict is unreadable. Never touch a cell mid-read: budget moves past roughly 20%, creative swaps, and audience edits each reset learning and torch the spend already burned. Verdicts come at the fixed maturation window, not from day-2 panic.

4. Bank the learning, not just the winner

The standard post-test ritual (scale the winner, archive the losers, move on) throws away most of what you bought. A clean negative that permanently kills a bad angle is worth its full test spend. The metric that compounds isn’t win rate; the concentration data above says win rate stays ugly for everyone. It’s cost per validated belief, and whether validated beliefs actually redirect next round’s production. Write verdicts where the next planning session must read them, or you’ll pay to learn the same thing twice.

5. Cadence beats bursts

Testing capacity is a pipeline: concurrent cells times verdict rate. Run rounds on a fixed cadence, each one a portfolio. Mostly probes of adjacent, promising territory; always a couple of genuine explorations; occasionally a wildcard. Twenty tests this week and none for three weeks buys you learning-phase churn and a decision pile-up. The same money on a steady weekly rhythm turns into compounding verdicts.

What this does to the org chart

Run those five disciplines by hand and count the roles: someone reads results daily, someone maintains the ledger of what’s been learned, someone designs each round as a portfolio, someone briefs and produces against the gaps, someone launches, monitors, kills, promotes. That’s a planner, an analyst, a producer, and a media buyer running a daily loop. It’s an operations problem now, not a campaign problem. Human attention caps it at a few dozen concurrent cells while the production stack can feed thousands. That gap is the frustration most growth teams are living inside right now.

Full disclosure of where we stand: closing that gap is the product we build. HyperScale is an agent team that runs the loop autonomously for subscription apps. It reads the account and the market’s winning sets, plans each round as a portfolio of test hypotheses, produces the ads, and writes every verdict back into the account’s belief ledger that the next round plans from. Humans keep exactly two moments: swiping verdicts on the ads themselves, and consenting to launches. (If you want the technical case for why this loop is about to matter even more, we’ve written up what generative retrieval does to creative economics.)

But you don’t need our product to act on this essay. You need the inversion: treat production as infinite, tests as scarce, and learning as the asset on the balance sheet.

The week-one checklist

If you run growth for a subscription app, this week:

  1. Count last quarter honestly. Creatives produced; creatives that got 50+ optimization events in a week (a fair read); verdicts anyone wrote down. Those three numbers are the diagnosis.
  2. Write down the ten beliefs your account currently operates on. Mark which were actually tested and which were inherited from a 2023 gut call.
  3. Kill near-duplicates before launch. If two concepts probe the same belief, one of them is spending your scarcest resource to teach you nothing.
  4. Ship a user-quality event. One onboarding question, who are you, your personas as the options, fired as a custom event and attributed per creative through your own web-to-app domain. It buys you attention-altitude reads qualified by user quality, days before trial data matures.
  5. Re-price your tests. Compute cost-per-fair-read at your real trial CPA, and budget tests as concurrent cells, not a monthly lump.
  6. Put “cost per validated belief” next to CPA on the dashboard. It’s the only number in this essay that compounds.

Production is solved. Testing is the moat now. The teams that internalize this first will spend the next two years compounding beliefs, while everyone else compounds render minutes.