A Creative Testing Framework for Health Ads: Budgets, Kill Criteria, and Reading Results Without Fooling Yourself
Ask a health brand’s media buyer how creative testing works in the account and you’ll usually get one of two answers: “we put new ads in and see what happens,” or a testing framework so elaborate it hasn’t produced a decision in six weeks. Both fail the same way - money goes out, and nobody can say what the account learned.
Testing isn’t launching ads. Testing is a structure that converts spend into decisions. Here’s the framework AdBoost Health runs across health, telehealth, and supplement partners, including the numbers.
How should you structure a creative test cell?
The cardinal rule: one variable per test. If a new hook, a new angle, and a new format launch in the same ad, a win teaches you nothing - you can’t reproduce what worked because you don’t know what worked.
Our standard structure:
- A dedicated testing campaign, separate from scaling campaigns. New creative dropped into a proven ad set inherits (and pollutes) the algorithm’s existing optimization. Winners graduate out of the test campaign into scaling; losers die where they stand.
- Cells of 3–5 variants that share everything except the variable under test. Testing angles? Same format, same offer, different angle per variant. Testing hooks? Same body, different first three seconds.
- Broad targeting inside tests. Narrow test audiences make results unrepresentative of how the winner will behave at scale - the most common reason “test winners” collapse when promoted.
- A control. Your current best performer runs in the cell. Beating zero is not a bar; beating the champion is.
Test angles before hooks, and hooks before formats. An angle is a reason to buy; in a category where compliant angles are a constrained list (we mapped them in the creative rejection playbook), knowing which three angles carry your account is worth more than any single ad.
How much should you spend per variant before deciding?
Roughly 1–1.5x your target CPA per variant before any conversion-based judgment - decided before launch. The most expensive habit in creative testing is killing ads on $30 of feelings; the second most expensive is running losers for three weeks out of hope, and a pre-committed budget is the escape from both.
The rule in practice: Selling a $60 product with a $50 target CPA? A variant hasn’t been tested until it’s spent $50–75. Anything less and a single random conversion - or its absence - decides the ad’s fate, which is a coin flip wearing a spreadsheet.
Work backwards from that and you get the honest constraint most teams dodge: testing budget determines testing velocity. At a $50 target CPA, testing 20 variants a month costs $1,000–1,500 in dedicated test spend. We hold test spend at 10–20% of account budget for scaling accounts. If your budget only supports properly testing six variants a month, test six - twenty half-tested variants produce zero decisions, and a decision is the unit of output.
Higher-ticket telehealth complicates this: at a $200+ CPA, fully powering every test on purchases gets expensive fast. That’s where leading indicators earn their keep.
What are the leading indicators, and when do you trust them over ROAS?
ROAS is the decision metric, but it’s slow and expensive per data point. Leading indicators are cheap and fast - they can’t tell you an ad will convert, but they reliably tell you an ad won’t.
| Metric | What it tells you | Trust it to… |
|---|---|---|
| Hook rate (3-sec views / impressions) | Does the opener stop the scroll? | Kill early - a dead hook is unfixable by waiting |
| Hold rate (ThruPlay or 50% views) | Does the middle sustain attention? | Diagnose where creative loses people |
| CPM | Does the platform think the creative is quality? | Spot creative the auction is penalizing |
| Outbound CTR | Does the ad generate intent? | Rank variants before conversions accrue |
| CPA / ROAS | Does it actually make money? | Make the scale decision - nothing else does |
The operating rule: kill on leading indicators, scale on ROAS. A variant with a hook rate far below your account median after a few hundred impressions is dead - no purchase data required, and waiting for purchase data is just paying to confirm it. But the reverse move is the classic self-deception: scaling on a great hook rate and CTR before conversion data exists. Curiosity clicks are not buying intent, and health has a particular version of this trap - provocative-but-vague hooks that pull huge engagement from people who will never buy a $70 supplement.
Fatigue also announces itself in leading indicators first. Rising CPM and frequency on a stable winner is your two-week warning before ROAS sags - that’s the signal to queue the next iteration, not the day ROAS finally breaks.
What kill criteria stop a team from fooling itself?
Written kill criteria exist because every human watching an ad account will otherwise negotiate with the data. “It’s about to turn around” has burned more test budget than any algorithm change. Ours, roughly:
- Kill at ~25–30% of test budget if hook rate and CTR are both well below account median. The creative can’t earn attention; conversions won’t save it.
- Kill at 100% of test budget with zero conversions. No debate, no extension, no “but the comments are good.”
- Extend one budget cycle only for a specific, pre-named reason - e.g., CPA near target on low volume, or strong add-to-carts with a checkout-side explanation. “I like this ad” is not a reason.
- A winner must beat the control on CPA, not merely achieve profitability. The bar is the champion.
One more anti-self-deception rule: read results at the cell level in a fixed weekly review, not ad-by-ad whenever anxiety strikes. Peeking daily at small-sample data is how teams see patterns in noise and kill future winners.
Iterate on winners or test new concepts?
Both, on a fixed ratio - this is where the angle library becomes the actual asset. Every test result gets tagged by angle, format, and hook, so the account accumulates knowledge instead of anecdotes. From that library, our standard split within the 20+ monthly variants we produce per partner: roughly 70% iterations on proven winners (new hooks on a winning body, new formats of a winning angle, new creators delivering a winning script) and 30% genuinely new concepts.
Iterations keep this month’s CAC alive; new concepts are the only insurance against the day a whole angle fatigues. Teams that only iterate ride one angle into the ground and hit a wall with no successor - usually right when they’re trying to scale spend, which is the worst possible moment (more on that dynamic in the supplement CAC playbook). Teams that only chase new concepts abandon compounding learnings and reset to zero every month.
If your account can’t currently answer “what’s our hit rate, what does a test cost, and which angle is carrying us” - that’s not a talent gap, it’s a missing framework. We’ll map your last 90 days of creative spend against this structure on a free 30-minute strategy call and hand you the written testing plan either way.