Skip to main content

7 posts tagged with "Analytics"

View all tags

When Should You Stop a 1688 Ad Campaign? Run Five Checks First

Β· 14 min read

TL;DR​

Stopping campaigns is where marketplace ad budgets leak fastest: waiting one extra week on a keeper costs little, while keeping a real underperformer one week costs real money. Five checks β€” runtime, spend, continuity, learning period, true zero-inquiry β€” are not folklore. They are five rules that actually run inside a production judging codebase, each with an explicit numeric threshold: 16-day settlement, Β₯100 weekly spend, 25% cost gap, 5-record samples, 3Γ— spend. Four say "wait," one says "stop now." Every case below comes from one store's 40-week ad ledger, judged on settled data only.

The situation: half the "bad campaigns" on Monday's report were wronged​

Monday review: a campaign spent real money last week with zero inquiries, and your hand is already on the pause button. Hold it β€” in one industrial-supplies store's 40-week ledger (anonymized), most campaigns that "looked bad" were simply not done proving themselves: some hadn't run a full cycle, some had spent nothing at all, some had a delivery gap in the middle. A wrong stop costs twice: the killed campaign's future inquiries, and the money a real underperformer keeps burning while you hesitate. The cases below come from a campaign-verdict audit while building AI Operations.

Why: four checks say wait, one says stop​

Check 1: Has it run long enough? β€” only settled data counts​

Rule first: the judgment consumes settled weeks only β€” a calendar week clears the settlement line 16 days after it ends, and the line is strict: day 16 itself is not settled, the next day is. Nothing unsettled enters a verdict. Measurement shows 16 days is a floor, not the tail: one collection run wrote 27 new region rows into a week that had ended 29 days earlier β€” see Is 16 Days Enough for Marketplace Ad Data?.

In plain terms: take the new program below β€” across its 7 settled weeks, weekly cost per inquiry swung from Β₯46 to Β₯65. The eighth week, which has not cleared the line yet, currently reads Β₯99 β€” it settles on September 8, and this article draws no conclusion from it. Judge the program on unsettled data, or on any single week you happen to pick, and the verdict is wrong either way.

(Technical note: inquiry attribution back-fills over time; the settlement line exists precisely to fence off "draft" data. A day-29 back-fill means the fence needs a watching period behind it, not just the line.)

Check 2: Has it spent enough? β€” thin spend is not judged​

Rule first: the rules turn "hasn't spent enough" into a hard gate β€” campaigns whose settled-week average spend sits below Β₯100 are filed as "test" and receive no performance verdict; a first settled week with fewer than 5 inquiries plus leads is protected, not judged.

In plain terms: one "precision targeting" campaign appeared in the ledger twice β€” April and July β€” for 6 weekly rows and a lifetime spend of Β₯0. It held a name on the report without spending a cent, never earning the right to be judged. Another veteran campaign (the store-wide self-serve program) fell to Β₯90 a week in June β€” under a tenth of its Β₯1,197 peak. At that scale the rules file it under "test" too; the word "stop" never comes up. When thin spend looks bad, the finding is "spent too little," not "performs badly."

(Technical note: a Β₯0 weekly row means the campaign never passed the platform's delivery checks or is budget-throttled; a thin-spend sample is all noise, and noise drowns every ratio computed on it β€” the gate exists to block exactly that false signal.)

Check 3: Is delivery continuous? β€” interrupted campaigns get two paths​

Rule first: continuity is measured as delivery density β€” weeks actually delivered Γ· weeks spanned β€” and under 80% counts as interrupted. An interrupted campaign gets two paths only: if effective cost-per-acquisition runs more than 25% above benchmark, stop, with the reason stated as "the bidding model keeps re-learning"; below the threshold, the file is marked "optimize" and the only action is: restore continuity.

In plain terms: one measured 3-week gap dropped store-wide weekly spend from Β₯1,960 to Β₯281 β†’ Β₯90 β†’ Β₯281, with inquiries hitting zero in the middle. Once continuous delivery resumed, store-wide cost per inquiry jumped from Β₯33 before the gap to Β₯46 in the restart week β€” restarting after a gap is not starting from where you left off, which is exactly why the rule restores continuity before judging.

(Technical note: the traffic mix before and after a gap can differ, so grading the restart against the pre-gap baseline runs systematically optimistic; "keeps re-learning" refers to the bidding model falling back to cold start after every interruption.)

Check 4: Was the learning period honored? β€” protection lasts one extra week at most​

Rule first: the learning-period protection here is not a fixed number of weeks. It fires only when the first settled week is continuous but sample-poor β€” fewer than 5 inquiries plus leads β€” and grants at most one more week; from the second settled week on, there is no protection at all. From there, every week is measured by the same ruler: effective CPA above target by more than 25%, combined with a thin inquiry share (under 15%, or under 70% of the store's own level) β€” or spend still trending up (up more than 12% over two weeks with cost above target, or a rising 3-week slope) β€” means the stop tier; gaps past 50% with cost above target are judged even faster.

In plain terms: the same store's new program β€” the "Merchant growth" plan β€” launched in late June, and its human operators kept waiting: by press time it had 8 delivery weeks on the books, 7 of them settled. Weekly cost per inquiry ran Β₯46–65 across the settled weeks β€” not one week back inside the store's own normal band of Β₯25–31, with the cheapest week still nearly 50% above the band's ceiling. Over the same stretch, two other new programs in the same store, aimed at the same products β€” the "potential-customer harvest pack" and the "cross-border express program" β€” bought inquiries at Β₯30 and Β₯35 across the same 7 settled weeks. So neither "the market got expensive" nor "it hadn't started yet" holds. The 8 weeks of patience came from people, not from the rules: under the judging logic, from the second settled week on, this program had no protection and should have been measured every single week.

Eight weeks of learning period, not one week back in the band

(Technical note: effective CPA = spend Γ· (quality inquiries + plain inquiries Γ— 0.6 + raw leads Γ— 0.1) β€” raw leads are worth little, they cannot prop up the denominator, and piles of junk leads cannot buy a cheap cost. Across the 7 settled weeks the program read Β₯15,188 Γ· 236 inquiries = Β₯64.4, 2.3Γ— the store's historical median of Β₯28.)

Check 5: Is it truly zero-inquiry? β€” the only stop-now tier​

Rule first: accumulated spend above 3Γ— the target cost-per-acquisition with zero total inquiries is a hard stop; when no target is configured, the threshold degrades to 3Γ— the campaign's own settled-week average spend. One tier fires even earlier: a closed-but-unsettled calendar week (Sunday passed, still inside the attribution window) spending past max(Β₯300, 3Γ— the weekly average) with zero inquiries is an early hard stop β€” it does not wait for settlement. And the target is never hand-set β€” it is the median across the store's last 12 settled, computable weeks, updated automatically.

In plain terms: this is the one check you never hesitate on. Its real battlefield is the keyword layer: across 46 weeks, 78% of the same store's 1,641 keyword-week records produced no inquiry while absorbing 27% of keyword spend β€” see 78% of Keywords Never Brought an Inquiry.

(Technical note: the early stop dares to skip settlement because spend is real-time billing β€” fixed once written, zero drift measured on settled weeks β€” while inquiries are a conversion field that back-fills from zero; the pair of conditions, a high bar and a closed week, is what bounds the false-kill risk.)

The experiment and the data​

The five checks' thresholds at a glance (values live in the production judging code):

CheckProduction thresholdVerdict
RuntimeA week clears the settlement line 16 days after it ends (day 16 itself not settled, next day counts)Unsettled data enters no verdict
SpendSettled-week average spend < Β₯100; or first settled week inquiries + leads < 5Filed "test": not judged / one extra week at most
ContinuityWeeks delivered Γ· weeks spanned < 80% = interrupted; interrupted and gap > 25%Stop; below the line β†’ restore continuity first
Learning periodFires only on a sample-poor, continuously delivered first settled week; never from the second settled week onOne extra week at most
Stop tierCost above target and gap > 50% β†’ stop; gap > 25% plus (inquiry share < 15% or under 70% of store level, or spend trending up) β†’ stopStop
True zero-inquiryLifetime spend > 3Γ— target with zero inquiries; early tier: closed-but-unsettled week > max(Β₯300, 3Γ— weekly average) with zero inquiriesStop now
Keep tierGap ≀ 5% and cost ≀ target and continuousKeep
Target costMedian of the store's last 12 settled, computable weeks, auto-updatedNo hand-setting
  • Sample: one industrial B2B store (anonymized), campaign-by-week ad ledger from Nov 2025 to Aug 2026 β€” 40 weeks; the keyword layer covers 46 weeks and 1,641 keyword-week records of the same store.
  • Calibers: inquiry cost = weekly spend Γ· weekly inquiries (cumulative uses total spend Γ· total inquiries); effective CPA = spend Γ· (quality inquiries + plain inquiries Γ— 0.6 + leads Γ— 0.1). The "normal band" is the store's own median weekly cost across the 28 normal weeks before the new programs launched (3 spring-festival weeks and 1 zero-inquiry gap week excluded) β€” Β₯28 Β±10%, i.e. Β₯25–31.
  • Settlement boundary: every verdict in this article reads settled weeks only. As of press time (Sep 4, 2026) the newest settled week is the week of Aug 10; the program's 8th week (week of Aug 17) settles on Sep 8 and appears here as an unsettled observation only.
  • Judging code: every rule and threshold cited here was verified against the production judge (the thin-spend gate, the interrupted-delivery branch, the learning-period sample gate, and the two-tier stop plus hard stop); file- and function-level provenance is an internal record and stays out of the article.
  • Anonymization: no store or campaign IDs appear; campaigns are referred to by their public platform program names.

What it's worth: two accounts​

The account of stopping late. Those 7 settled weeks of the "Merchant growth" program: Β₯15,188 spent for 236 inquiries. At the store's own median of Β₯28 across 28 normal weeks, the same 236 inquiries should have cost about Β₯6,600 β€” 7 settled weeks of overpaying, roughly Β₯8,600. Under the rules, protection lapsed at the second settled week and the program should have been measured weekly β€” every extra week of human patience was real money.

The account of not stopping. The 78% zero-inquiry keyword records carried Β₯8,962 of real spend β€” 27% of the store's Β₯33,417 keyword budget β€” without producing a single inquiry. Cutting them touches no campaign structure, and the money returns the same week.

For operators​

  1. Runtime: judge on settled weeks only (+16 days after week end); when single weeks swing hard, only cumulative numbers count.
  2. Spend: below a Β₯100 settled-week average, fund it before judging it; a Β₯0 weekly row is a delivery question, not a performance question.
  3. Continuity: under 80% delivery density, restore continuous delivery first; after a gap, reset the baseline β€” never grade the restart against gap weeks.
  4. Learning period: protection belongs to a sample-poor first settled week, one extra week at most; "give the new program time" stops being an argument at the second settled week.
  5. True zero-inquiry: past 3Γ— your target cost-per-acquisition with still zero inquiries β€” stop now, at campaign level and keyword level alike.

For developers​

  1. Persist both granularities: campaign-by-week and keyword-by-week are separate tables β€” the keyword layer is where the stoppage money lives, and campaign-level views never see it.
  2. Keep collection audit fields: re-collection rewrites historical weeks (measured: new rows arrived on day 29), so your pipeline must distinguish "what was visible then" from "settled data."
  3. Keep thresholds in one place: gather every gate into a single configuration and keep the judging logic free of scattered magic numbers β€” tune thresholds without touching logic, and version every logic change.
  4. Make the short-circuit order explicit: the five checks are not parallel options but a short-circuit chain β€” who judges first and what short-circuits what decides the verdict. The production judge's actual order:
1 Early stop  : closed-but-unsettled week spend > max(Β₯300, 3Γ— weekly avg), zero inquiries β†’ stop (runs first, skips settlement)
2 Hard stop : lifetime spend > 3Γ— target cost, zero inquiries β†’ stop
3 Test gates : settled weeks < 1 β†’ test; weekly average spend < Β₯100 β†’ test
4 Gap branch : delivery density < 80% β†’ gap > 25% stops; no benchmark or gap below the line β†’ optimize (restore continuity)
5 Learning : first settled week only, continuous, inquiries + leads < 5 β†’ test (one extra week at most)
6 Stop tier : settled β‰₯ 2 (β‰₯ 3 without a target): cost above target and gap > 50% β†’ stop; gap > 25% with (thin inquiry share or no improvement) β†’ stop
7 Keep tier : gap ≀ 5% and cost ≀ target and continuous β†’ keep
8 Fallback : everything else β†’ optimize
9 Store guard : if everything gets stopped, the biggest non-hard-stop spender is downgraded to optimize (or a sample-poor plan is kept when none qualifies)

How to use the five checks

Data not past the settlement line (+16 days after week end) β†’ wait; weekly average spend under Β₯100 β†’ fund it first; delivery interrupted β†’ restore it first; sample-poor first settled week β†’ one extra week at most; spend past 3Γ— target cost with zero inquiries β†’ stop now. Four "waits," one "stop."

FAQ​

How long should a B2B ad campaign run before judging it?​

Judgment reads settled weeks only: a calendar week clears the settlement line 16 days after it ends, and measured back-fill has arrived as late as day 29. The first settled week is judgeable, but a first week with fewer than 5 inquiries plus leads is protected for one extra week at most.

Which failure justifies stopping a campaign immediately?​

Lifetime spend past 3Γ— the target cost-per-acquisition with zero total inquiries β€” the hard-stop tier, with the target auto-set to the median of the last 12 settled weeks. An earlier tier fires too: a closed-but-unsettled week spending past max(Β₯300, 3Γ— its weekly average) with zero inquiries stops without waiting for settlement.

How do I judge a campaign after a delivery gap?​

Delivery density under 80% counts as interrupted: stop if cost runs more than 25% above benchmark (the bidding model keeps re-learning), otherwise restore continuity first. One measured 3-week gap pushed store-wide inquiry cost from Β₯33 to Β₯46.

All five checks are built into AI Operations β€” LLM-powered analysis that reads market trends, buyer behavior, and sales data to ground your operating decisions in numbers. It waits when waiting is right, and flags the stop a week early.

CCLEE

Independent developer, 24 years in e-commerce, focused on grounding AI in real business scenarios.

Work with me

78% of Keywords Never Produced an Inquiry: Your Cost per Inquiry Is Understated 28%

Β· 4 min read

TL;DR​

One store's keyword ledger: of 1,641 keyword-week records, 1,277 (78%) produced zero inquiries while consuming Β₯8,962 β€” 27% of total spend. The average over "keywords that did produce" comes to Β₯32; the blended cost that counts every yuan of spend is Β₯41 β€” a 28% understatement. The only correct formula is total spend Γ· total inquiries: not one yuan of zero-inquiry spend may vanish from the denominator.

The situation: the cost you see is the survivors' cost​

Keyword reports are organized per keyword, so your eyes land on keywords that produced inquiries β€” those have a cost to show. Zero-inquiry keywords display no cost, and so they quietly exit your field of view.

Their spend, however, left the account in full. This audit came out of a keyword-level check while building AI Operations.

The data: the missing 27%​

MetricValue
Keyword-week records1,641
…with zero inquiries1,277 (78%)
Spend on zero-inquiry recordsΒ₯8,962 (27% of total)
Total spend / total inquiriesΒ₯33,417 / 824
Naive average over inquiring keywordsΒ₯32
Blended cost (total Γ· total)Β₯41

The naive average only bills the survivors β€” (technical note: this is textbook survivorship bias in ad data. Counting only producing samples donates the non-producing samples' spend for free; with 27% of the money missing from the denominator, the cost "improves" by two to three tenths.)

What it's worth: what 28% understatement does​

The understatement is not cosmetic β€” it cascades:

  • Acquisition budget: budget set at Β₯32 while reality is Β₯41 leaves a Β₯9-per-inquiry hole β€” across 824 inquiries, about Β₯7,400
  • Product go/no-go: a product line judged against an understated keyword cost reads "still viable" while truly underwater
  • Pricing and margin: acquisition cost is the hidden floor of B2B quotes; a floor 28% too low cannot carry real deal prices

Disciplines for operators​

  1. One base formula: cost = total spend Γ· total inquiries. Any "average" that excludes zero-inquiry samples is void on sight.
  2. Keep a separate zero-inquiry watchlist, sorted by accumulated spend β€” this is the main battlefield of "check five: truly zero inquiries"; words that burn past a reasonable cost with nothing to show get stopped without ceremony.
  3. Quality weighting is layer two: once the base is right, weight purchase-ready inquiries (pricing asks, sample requests, volume) above casual ones. No universal weights exist β€” derive them from your own deal path and freeze them, so months stay comparable.
  4. The target line comes from your own history: normal months (holidays excluded) define the band β€” the method is in A Real Store-Wide Efficiency Alert, From a 40-Week Ledger.

One line to remember

Cost = total spend Γ· total inquiries. Zero-inquiry spend never disappears from the denominator; quality weighting is always layer two.

FAQ​

What is the right way to calculate B2B ad inquiry cost?​

Total spend Γ· total inquiries β€” no exceptions. Leave zero-inquiry spend out of the denominator and the cost reads two to three tenths too low.

Should zero-inquiry keywords be paused?​

Check accumulated spend and observation window first: spend clearly above a reasonable cost with still zero inquiries means stop; freshly added keywords deserve a full settlement cycle.

Is quality-weighting inquiries still worth doing?​

Yes β€” as a second layer. Fix the base formula first (zero-inquiry spend must not vanish from the denominator), then grade purchase-ready vs casual inquiries.

That base-formula discipline is built into AI Operations β€” LLM-powered analysis that automatically surfaces market trends, user behavior, and sales data to drive strategy. One notch wrong on the cost formula, and everything downstream is wrong.

CCLEE

Independent developer, 24 years in e-commerce, focused on grounding AI in real business scenarios.

Work with me

4,111 Leads vs 824 Inquiries: Two Acquisition Costs in One B2B Ad Ledger

Β· 4 min read

TL;DR​

One and the same ad ledger, two "acquisition costs": lead cost Β₯8, inquiry cost Β₯41 β€” 5Γ— apart. One store's keyword-week records measured 4,111 leads against 824 inquiries. Leads are the wide net (favorites, add-to-carts, coupons, contact clicks β€” all counted per action); inquiries are one thing only: buyer-initiated quote requests. Read the buzz from leads, build budgets from inquiries β€” plan on Β₯8 and the shortfall exists from day one.

The situation: Β₯8 per customer, uplifting enough to scale on​

The monthly report's "lead cost" line is beautiful: Β₯8 per interested customer. Configuring next quarter's acquisition budget on that number feels like clean logic.

Until "inquiry cost" sits down next to it: Β₯41. Same account, same spend, same month β€” two numbers, 5Γ— apart. This metric audit came out of a definitions check while building AI Operations.

The data: a 5Γ— gap, and a one-way containment​

MetricValue
Total leads4,111
Total inquiries824
Ratio5.0Γ—
Lead cost (Β₯33,417 Γ· 4,111)Β₯8.1
Inquiry cost (Β₯33,417 Γ· 824)Β₯40.6
Keyword-weeks with leads, zero inquiries269
Keyword-weeks with inquiries, zero leads0

The last two rows carry the argument: inquiries always bring leads; leads almost never guarantee an inquiry. Leads are a superset β€” (technical note: the lead definition counts favorites, add-to-carts, coupons, contact clicks and similar interactions, accumulated per action; one buyer clicking "contact" three times logs three leads. An inquiry is exactly one thing: a buyer-initiated quote request. These are not "two metrics" β€” they are "all interactions" versus "the most valuable kind.") β€” and those 269 lead-without-inquiry records are the distance between buzz and business: engagement happened, the quote request never did.

What it's worth: a 5Γ— budget hole​

Set the budget floor on a Β₯8 lead cost and the market spend gets configured as if Β₯8 buys a customer; the real denominator is Β₯41, so the gap is 5Γ— from the first day. That is not optimism β€” it is systematic misallocation: every downstream decision (quotes, margins, scaling pace) sits on an inflated denominator. The lead metric still has a job, and only one: telling you whether content and campaigns moved the buzz.

Disciplines for operators​

  1. Book the two metrics separately, names included: "lead cost Β₯8," "inquiry cost Β₯41" β€” any report that says just "acquisition cost" for both is an accident waiting.
  2. Budgets, repricing, product P&L run on inquiries only: quote requests cannot be inflated, sit closest to orders, and are the only denominator that counts.
  3. Leads are the buzz thermometer: a lead spike sends you to check creative and campaigns; it was never a scorecard.
  4. Same rule for audience repricing: audiences are graded on inquiry cost β€” the full logic is in Inquiry Costs 10% Apart, ROI 4Γ— Apart.

One line to remember

Leads count per action, wide net, for buzz; inquiries one per request, close to orders, for budgets. Β₯8 tells the win; Β₯41 is the truth.

FAQ​

What is the difference between leads and inquiries in marketplace reports?​

An inquiry is a buyer-initiated quote request β€” one action, one record. Leads bundle favorites, add-to-carts, coupons, and contact clicks, counted per action. Same account measured: 4,111 leads vs 824 inquiries β€” 5Γ— apart.

Which cost should reports and budgets use?​

Lead cost (Β₯8) is fine for external wins; budget decisions run on inquiry cost (Β₯41). Planning on Β₯8 builds a 5Γ— hole into the quarter.

What does 'leads but no inquiries' mean?​

269 measured keyword-weeks had leads with zero inquiries β€” engagement that never reached the quote request. Fine for reading buzz; not a result.

That "metrics booked by definition" practice is built into AI Operations β€” LLM-powered analysis that automatically surfaces market trends, user behavior, and sales data to drive strategy. Two metrics 5Γ— apart should never share a name.

CCLEE

Independent developer, 24 years in e-commerce, focused on grounding AI in real business scenarios.

Work with me

One Week at ROI 61.6, Eight Weeks Under 2: How a Lucky Order Ruins Ad Judgment

Β· 4 min read

TL;DR​

Nine weeks of real ledger for one product in one campaign: 8 weeks at ROI between 0 and 2.1, and one week at 61.6 β€” 6 orders carrying Β₯15,143. If you happened to open the report that week and scaled budget, the next month walked it straight back down. A big order is a surprise, not a baseline: budgets follow the normal pace of inquiries, not the luck of orders.

The situation: the weekly report that looked too good to question​

The weekly report pulls up: one product at ROI 61.6 on Β₯246 of spend, Β₯15,143 in orders. Every operator's pulse quickens β€” a ten-x signal, worth budgeting, worth replicating.

Lay out nine weeks before touching anything. This drill-down came out of a product-ledger audit while building AI Operations.

The data: one needle in nine weeks​

Same product, same campaign (whole-store promotion), nine consecutive weekly rows:

WeekSpendOrdersOrder valueROI
1Β₯1983Β₯240.1
2Β₯2494Β₯2130.9
3Β₯1960Β₯00.0
4Β₯2466Β₯15,14361.6
5Β₯2338Β₯4902.1
6Β₯1563Β₯2281.5
7Β₯1802Β₯2081.2
8Β₯1803Β₯520.3
9Β₯2340Β₯00.0

Spend held steady at Β₯156–249 all nine weeks. The only variable that moved in week four was order value. The running norm is ROI around 1; the 61.6 fell out of the sky.

Why it deceives​

(Technical note: weekly granularity cannot see inside the orders. Week four's Β₯15,143 across 6 orders averages Β₯2,524 per order β€” dozens of times the neighboring weeks' per-order value. Whether that was one large order or several mid-size ones is only answerable at daily or order level; the weekly report can't say β€” but it says enough that "something unusual happened; conclude nothing yet.")

The big-order week creates three illusions at once: it inflates perceived acquisition ability (inquiry volume never moved), it promises repeatability (big orders are low-probability draws), and it aims your budget at the wrong place (the norm was ROI β‰ˆ 1 β€” scaling a norm-negative setup scales the loss).

What it's worth: the misallocation ledger​

Budgeting on week four's 61.6 treats the setup as a ten-x machine. Two calculations, two worlds: the 9-week blended ROI is Β₯16,358 Γ· Β₯1,872 = 8.7; excluding the big-order week, the 8-week norm is Β₯1,215 Γ· Β₯1,626 = 0.7. The average lies on the big order's behalf β€” one number, two lives. A norm of 0.7 means seventy cents back per yuan spent: this setup's ~Β₯200 weekly burn was already a net loss, and scaling it only scales the loss.

Disciplines for operators​

  1. Read windows in segments, never as one average: cut the observation period into 4-week chunks β€” the average blends "once was good" and "now is not" into a fictitious "okay."
  2. Inquiries are the thermometer: flat inquiries across the spike mean acquisition ability never changed, only luck did; inquiries shrinking alongside means real decay β€” entirely different treatment.
  3. Book big orders as surprises: budget decisions run on the no-big-order norm; for setups propped by one, extend observation and look at a cycle without the luck.
  4. Extreme weeks trigger drill-downs, not decisions: seeing 61.6 or 0.0, step one is always the daily-level distribution β€” never the budget slider.

One line to remember

Attribution clustered in one week, falling back after, inquiries unchanged = big-order illusion. Budget on the norm of inquiries, not the luck of orders; extreme weeks trigger drill-downs, not decisions.

FAQ​

How do I tell if ROI is propped up by a big order?​

Split the window and look: attribution clustered in one week, falling back after, with inquiry volume unchanged β€” the acquisition ability never changed; that week's luck did.

Why can't I budget on a big-order week's ROI?​

It doesn't repeat. Measured case: one week at ROI 61.6, the other eight between 0 and 2.1 β€” scale budgets on 61.6 and the next cycle arrives before the luck does.

What if attributed orders suddenly drop to zero?​

Check inquiries first: unchanged inquiries mean the luck receded β€” decide on normal efficiency; shrinking inquiries mean real decay β€” entirely different treatment.

That "segmented windows + inquiry thermometer + daily drill-down" method is built into AI Operations β€” LLM-powered analysis that automatically surfaces market trends, user behavior, and sales data to drive strategy. Extremely beautiful numbers deserve verification before belief.

CCLEE

Independent developer, 24 years in e-commerce, focused on grounding AI in real business scenarios.

Work with me

Is 16 Days Enough for Marketplace Ad Data? We Re-Collected 5 Weeks to Find Out

Β· 8 min read

TL;DR​

"Wait 16 days before judging" is not enough. We re-collected five long-settled weeks of B2B marketplace ad data and diffed every row: finished weeks keep growing new data as late as day 29 (keywords added late reach back and pull their history in), yet past day 40, even a full re-collection changed not one existing number; and the harsher finding is retention β€” the platform deletes detail reports after ~6–7 weeks, and past 54 days none of our 6 campaigns could be recovered. Collect weekly, archive locally β€” what that buys is twofold: no more mis-calls on irreversible decisions, and ad history that actually belongs to you.

The trigger: 27 new rows on day 29​

During a routine collection on August 10, one industrial-goods B2B storefront (anonymized) wrote 27 new region-detail rows into a week that had ended 29 days earlier β€” five campaigns, a +22% row increase. By the folklore timeline, that week had "settled at 16 days" roughly two weeks prior.

Is 16 days platform fact or folklore? The test is ready-made: re-collect older weeks. If long-settled weeks grow new rows or rewrite old values, 16 days is nowhere near the end β€” and the same re-collection measures how long the platform keeps detail at all. On August 21 we re-pulled five older weeks, which became the experiment below.

The experiment itself came out of a data-semantics check while building AI Operations.

The design: re-collect five "settled" weeks and diff​

The method is plain: baseline β†’ re-collection β†’ row-level diff.

  1. Pick 5 consecutive weeks (2026-06-08 ~ 07-06), already 40–74 days old β€” settled under any definition;
  2. Export a baseline across 4 detail tables (overview / products / keywords / areas β€” 1,199 rows);
  3. Trigger a full platform re-collection over an explicit date range, export again;
  4. Diff row by row: additions, value changes, deletions. Re-checked on 2026-09-02; conclusions unchanged.

The result:

TableBaseline rowsAfterAddedValue changesRemoved
Campaign overview1818000
Product detail682682000
Region detail405405000
Keyword detail94105+1100

The net result of the row-level diff across all 1,199 rows: not a single historical number was revised (zero value changes, zero removals). The only difference is 11 new keyword rows β€” not revised values but newly grown entities, a second back-fill mechanism covered in Finding 2; which campaigns' detail the platform re-returned at all is Finding 3.

Lifecycle of 1688 P4P data: from routine writes to back-fill events to deletion

Finding 1: values blow past 16 days β€” and truly settle by 40​

The boundary of the claim: "16 days" is both right and wrong.

  • The platform's stated attribution window is ~15 days β€” the folk "settle at 16" comes from there;
  • The trigger event was exactly that counterexample: +27 region rows on day 29, nearly double the window;
  • The experiment supplies the stopping point: re-collection at 40–47 days showed zero value changes β€” back-fill does stop, just much later than day 16.

For operators: 16 days works as a "probably stable" heuristic, not as a definition of final. For irreversible calls like pausing a campaign, wait the full 4–5 weeks so the decision lands past the measured stopping point. The full pre-pause checklist is in Five checks before you pause a marketplace ad campaign.

Finding 2: keywords you add later reach back and pull their history in​

Plain version first: how much "history" you can recover depends on which keywords you are running now.

The test store added keywords to one campaign mid-flight. On re-collection, the platform returned the campaign's current keyword list together with past weekly performance β€” one recovered historical week carried 11,179 impressions, Β₯342.7 spend, 10 inquiries, and 11 orders. Those numbers sat on the platform's side all along; they only appear when you come collect.

(Technical note: the mechanism is entity-level back-fill β€” the API returns history for your current keyword list. Row counts rise; existing values never change. So when re-collected data grows, first tell revised values from newly grown rows: the former is a warning, the latter is a gift.)

But the gift has a precondition: the platform must still retain the detail. For how long? That is Finding 3.

Finding 3: the real cliff is deletion, not settlement​

Slow back-fill costs waiting; retention costs everything. We checked, per campaign, whether the platform could still return detail during re-collection:

  • Weeks aged 40–47 days: from the same batch of six campaigns, only 3 still returned complete detail;
  • Weeks aged 54–74 days: the same six again β€” 0 of 6; not stale, deleted from the platform's side, unrecoverable by any means;
  • The 3 campaigns already returning nothing at 40 days run a shorter, per-campaign retention β€” they hit the cliff earlier; one sample so far.

The retention cliff: 3/6 campaigns retrievable at 40–47 days, 0/6 after 54

Put bluntly: any detail not in your own database within ~6 weeks has been deleted on your behalf. "I'll export it later" is not procrastination β€” it is deletion.

What these findings are worth: two ledgers​

The mis-kill ledger. The biggest cost of the 16-day folklore is reading "the data hasn't arrived" as "the campaign doesn't work". The keyword whose history grew back in Finding 2 carries 10 inquiries and 11 orders in a single week β€” pause a ramping campaign early, and that is the weekly loss, before counting the extra settlement weeks it needs to prove itself again. Waiting the full 4–5 weeks buys "no wrong irreversible calls".

The loss ledger. The platform deletes detail after ~6–7 weeks, so "I'll back-fill later" is a promise with an expiry date. Which regions fed you inquiries last quarter, which keywords' costs were quietly climbing β€” only people whose data landed in their own database can answer. A weekly collection run buys "history, always queryable".

Three disciplines for operators​

  1. Watch trends anytime; wait 4–5 weeks for irreversible calls. 16 days is the reference line, 29 the safety line.
  2. Collect or export detail weekly and archive locally. Platform-side detail is far shorter-lived than assumed; the archive is the asset you own.
  3. Close the "settlement tail" (the data each finished period is still quietly back-filling) before any retro analysis. Period ROI without its back-fill is systematically understated β€” and when you cannot tell revised values from newly grown rows, use the experiment's move: diff at row level.

For data engineers: leave at least 5 weeks of overlap​

If your pipeline collects marketplace ad reports weekly (or models attribution for a similar platform): the stated 15-day window actually back-fills as late as +29 days. A strict weekly cadence with a 4-week overlap performs its last rewrite at week-end +22 days β€” which fails to cover the measured day-29 event (ours survived only because a collection gap happened to stretch). Raise the overlap from 4 weeks to 5: ~20–25% more requests each week β€” every extra overlap week is one more week of paginated detail to pull β€” in exchange for never losing rows.

Two weekly habits

  • Collect: land this week's detail in your own database, with a 5+ week overlap β€” archiving is the only defense against deletion;
  • Wait: make keep-or-stop calls only on data settled 4–5 weeks β€” settlement is the only defense against misjudgment.

FAQ​

How long until marketplace ad data is final?​

The platform states a ~15-day attribution window, but the latest back-fill we observed landed on day 29. Wait 4–5 weeks before irreversible keep-or-stop calls.

Can I back-fill ad history I forgot to export?​

Usually no. Detail reports survive roughly 6–7 weeks: at 40–47 days only 3 of 6 campaigns still returned data; past 54 days, none. After that it is gone.

Why do new keyword rows appear after re-collection?​

The API returns history for your current keyword list. Keywords added late grow historical rows β€” entity-level back-fill: row counts rise, existing values never change.

That baseline β†’ re-collection β†’ row-level diff method is built into AI Operations β€” LLM-powered analysis that automatically surfaces market trends, user behavior, and sales data to drive strategy. Ad data is the foundation of all of it β€” and a foundation deserves to be measured.

CCLEE

Independent developer, 24 years in e-commerce, focused on grounding AI in real business scenarios.

Work with me

A Real Store-Wide Efficiency Alert, From a 40-Week Ledger

Β· 5 min read

TL;DR​

A real ledger from one store: across 29 historical weeks, store-wide inquiry cost held a stable band of Β₯25–31. After new ad solutions launched on June 29, it ran at Β₯34–49 for 7 consecutive weeks β€” then hit Β₯69 in week eight. Store-wide decay never announces itself as an incident; it shows up as "every campaign looks fine." So you need a ruler for the sum: a band set by your own history, and an alert on 3 consecutive weeks above it. Had the alarm fired in week three, roughly Β₯7,400 of the overspend in the following five weeks would have been avoided.

The situation: every campaign "fine," the sum deteriorating​

Per-campaign reviews read as usual: this campaign's numbers match last month's, that one is stable, the new one is still ramping. All passable.

At the store level: every July week sat above Β₯40 per inquiry β€” in the previous 29 weeks, that line had been crossed exactly once (the Spring Festival week). This post-mortem came out of a store-ledger audit while building AI Operations.

Store-wide inquiry cost across 40 weeks: historical band and the breach

The method: set the band from your own history​

The baseline is not the industry and not a target β€” it is your own past:

  1. Take weekly store-wide inquiry cost (weekly spend Γ· weekly inquiries) for the past six months
  2. Exclude abnormal weeks: here, two kinds β€” the Spring Festival week (Β₯121, spend collapse masquerading as expensiveness) and a 3-week delivery gap in early June (weekly spend Β₯90–281, near-dark)
  3. The remaining 26 weeks land between Β₯20–41, concentrated in Β₯25–31 β€” that band is "normal"
  4. Alert condition: 3+ consecutive weeks above the band's top, with spend not shrinking (cost rising because spend is collapsing is a different problem)

(Technical note: why "consecutive 3 weeks" rather than any single week β€” in the 29 historical weeks, single-week breaches happened 4 times, all noise; consecutive breaches happened zero times. The baseline tells you how strict the threshold should be.)

The full post-mortem​

  • June 8–22: a 3-week gap. Weekly spend fell from ~Β₯1,900 to Β₯90–281 β€” near-dark
  • June 29: new ad solutions launched (the solution-switch details and data are in Same Product, 4Γ— the Inquiry Cost)
  • From June 29: 7 consecutive weeks above the band β€” Β₯46 / 43 / 40 / 43 / 40 / 49 / 34, every one above the historical top of Β₯31
  • Week of August 17: Β₯69, as spend spiked to Β₯5,210 without inquiries following

A gap-and-restart is not a return to the old normal: the environment changed and the solutions changed β€” the old cost level no longer applies. That is exactly the kind of account-level shift single-campaign views cannot see, and only the store line exposes.

What it's worth: the overspend ledger​

The 7 breached weeks (Jun 29 – Aug 10) spent Β₯24,699 for 592 inquiries β€” Β₯41.7 each. At the historical level (Β₯28), the same inquiries would have cost Β₯16,576: about Β₯8,100 of overspend in 7 weeks. Had the alarm fired in week three and intervention started in week four, roughly Β₯7,400 of the last five weeks' overspend was avoidable. That is the price of the ruler: set it once, watch one number.

Disciplines for operators​

  1. The band must come from your own history: at least six months of weekly data, abnormal weeks excluded. A three-week average as a ruler is worse than no ruler.
  2. 3 consecutive weeks above the top, with spend holding, = alert: single weeks are noise; consecutive weeks are structure.
  3. After an alert, hunt the common cause first: solution switches, gap-and-restarts, category-wide competition shifts are account-level events β€” fixing campaigns one by one treats symptoms.
  4. The cost numbers themselves must settle first: weekly data carries a settlement tail; the discipline is in Is 16 Days Enough for Marketplace Ad Data? We Re-Collected 5 Weeks to Find Out.

One line to remember

Set the band from history (six months, abnormal weeks out); 3 consecutive weeks above it = alert. Alerts trigger a hunt for account-level causes; actions stay at campaign level.

FAQ​

What signals a store-wide ads efficiency decline?​

Weekly inquiry cost above your own historical band for 3+ consecutive weeks while spend holds β€” each campaign can look passable alone while the sum is sinking.

How do I set the 'normal band'?​

Take your own half-year of weekly inquiry costs, exclude abnormal weeks (holidays, delivery gaps), and use the median Β±10% as the band. Stores with thin history should accumulate first.

What is the first move after an alert?​

Hunt for an account-level cause before touching campaigns: solution switches, gap-and-restart episodes β€” store-wide breaches are usually account-level events.

That "band from history, watch the store" alerting logic is built into AI Operations β€” LLM-powered analysis that automatically surfaces market trends, user behavior, and sales data to drive strategy. Efficiency decay should not wait for a quarterly review to be discovered.

CCLEE

Independent developer, 24 years in e-commerce, focused on grounding AI in real business scenarios.

Work with me