Skip to main content

4,111 Leads vs 824 Inquiries: Two Acquisition Costs in One B2B Ad Ledger

· 4 min read

TL;DR​

One and the same ad ledger, two "acquisition costs": lead cost ¥8, inquiry cost ¥41 — 5× apart. One store's keyword-week records measured 4,111 leads against 824 inquiries. Leads are the wide net (favorites, add-to-carts, coupons, contact clicks — all counted per action); inquiries are one thing only: buyer-initiated quote requests. Read the buzz from leads, build budgets from inquiries — plan on ¥8 and the shortfall exists from day one.

The situation: ¥8 per customer, uplifting enough to scale on​

The monthly report's "lead cost" line is beautiful: ¥8 per interested customer. Configuring next quarter's acquisition budget on that number feels like clean logic.

Until "inquiry cost" sits down next to it: ¥41. Same account, same spend, same month — two numbers, 5× apart. This metric audit came out of a definitions check while building AI Operations.

The data: a 5× gap, and a one-way containment​

MetricValue
Total leads4,111
Total inquiries824
Ratio5.0×
Lead cost (¥33,417 ÷ 4,111)¥8.1
Inquiry cost (¥33,417 ÷ 824)¥40.6
Keyword-weeks with leads, zero inquiries269
Keyword-weeks with inquiries, zero leads0

The last two rows carry the argument: inquiries always bring leads; leads almost never guarantee an inquiry. Leads are a superset — (technical note: the lead definition counts favorites, add-to-carts, coupons, contact clicks and similar interactions, accumulated per action; one buyer clicking "contact" three times logs three leads. An inquiry is exactly one thing: a buyer-initiated quote request. These are not "two metrics" — they are "all interactions" versus "the most valuable kind.") — and those 269 lead-without-inquiry records are the distance between buzz and business: engagement happened, the quote request never did.

What it's worth: a 5× budget hole​

Set the budget floor on a ¥8 lead cost and the market spend gets configured as if ¥8 buys a customer; the real denominator is ¥41, so the gap is 5× from the first day. That is not optimism — it is systematic misallocation: every downstream decision (quotes, margins, scaling pace) sits on an inflated denominator. The lead metric still has a job, and only one: telling you whether content and campaigns moved the buzz.

Disciplines for operators​

  1. Book the two metrics separately, names included: "lead cost ¥8," "inquiry cost ¥41" — any report that says just "acquisition cost" for both is an accident waiting.
  2. Budgets, repricing, product P&L run on inquiries only: quote requests cannot be inflated, sit closest to orders, and are the only denominator that counts.
  3. Leads are the buzz thermometer: a lead spike sends you to check creative and campaigns; it was never a scorecard.
  4. Same rule for audience repricing: audiences are graded on inquiry cost — the full logic is in Inquiry Costs 10% Apart, ROI 4× Apart.

One line to remember

Leads count per action, wide net, for buzz; inquiries one per request, close to orders, for budgets. ¥8 tells the win; ¥41 is the truth.

FAQ​

What is the difference between leads and inquiries in marketplace reports?​

An inquiry is a buyer-initiated quote request — one action, one record. Leads bundle favorites, add-to-carts, coupons, and contact clicks, counted per action. Same account measured: 4,111 leads vs 824 inquiries — 5× apart.

Which cost should reports and budgets use?​

Lead cost (¥8) is fine for external wins; budget decisions run on inquiry cost (¥41). Planning on ¥8 builds a 5× hole into the quarter.

What does 'leads but no inquiries' mean?​

269 measured keyword-weeks had leads with zero inquiries — engagement that never reached the quote request. Fine for reading buzz; not a result.

That "metrics booked by definition" practice is built into AI Operations — LLM-powered analysis that automatically surfaces market trends, user behavior, and sales data to drive strategy. Two metrics 5× apart should never share a name.

CCLEE

Independent developer, 24 years in e-commerce, focused on grounding AI in real business scenarios.

Work with me

One Week at ROI 61.6, Eight Weeks Under 2: How a Lucky Order Ruins Ad Judgment

· 4 min read

TL;DR​

Nine weeks of real ledger for one product in one campaign: 8 weeks at ROI between 0 and 2.1, and one week at 61.6 — 6 orders carrying ¥15,143. If you happened to open the report that week and scaled budget, the next month walked it straight back down. A big order is a surprise, not a baseline: budgets follow the normal pace of inquiries, not the luck of orders.

The situation: the weekly report that looked too good to question​

The weekly report pulls up: one product at ROI 61.6 on ¥246 of spend, ¥15,143 in orders. Every operator's pulse quickens — a ten-x signal, worth budgeting, worth replicating.

Lay out nine weeks before touching anything. This drill-down came out of a product-ledger audit while building AI Operations.

The data: one needle in nine weeks​

Same product, same campaign (whole-store promotion), nine consecutive weekly rows:

WeekSpendOrdersOrder valueROI
1¥1983¥240.1
2¥2494¥2130.9
3¥1960¥00.0
4¥2466¥15,14361.6
5¥2338¥4902.1
6¥1563¥2281.5
7¥1802¥2081.2
8¥1803¥520.3
9¥2340¥00.0

Spend held steady at ¥156–249 all nine weeks. The only variable that moved in week four was order value. The running norm is ROI around 1; the 61.6 fell out of the sky.

Why it deceives​

(Technical note: weekly granularity cannot see inside the orders. Week four's ¥15,143 across 6 orders averages ¥2,524 per order — dozens of times the neighboring weeks' per-order value. Whether that was one large order or several mid-size ones is only answerable at daily or order level; the weekly report can't say — but it says enough that "something unusual happened; conclude nothing yet.")

The big-order week creates three illusions at once: it inflates perceived acquisition ability (inquiry volume never moved), it promises repeatability (big orders are low-probability draws), and it aims your budget at the wrong place (the norm was ROI ≈ 1 — scaling a norm-negative setup scales the loss).

What it's worth: the misallocation ledger​

Budgeting on week four's 61.6 treats the setup as a ten-x machine. Two calculations, two worlds: the 9-week blended ROI is ¥16,358 ÷ ¥1,872 = 8.7; excluding the big-order week, the 8-week norm is ¥1,215 ÷ ¥1,626 = 0.7. The average lies on the big order's behalf — one number, two lives. A norm of 0.7 means seventy cents back per yuan spent: this setup's ~¥200 weekly burn was already a net loss, and scaling it only scales the loss.

Disciplines for operators​

  1. Read windows in segments, never as one average: cut the observation period into 4-week chunks — the average blends "once was good" and "now is not" into a fictitious "okay."
  2. Inquiries are the thermometer: flat inquiries across the spike mean acquisition ability never changed, only luck did; inquiries shrinking alongside means real decay — entirely different treatment.
  3. Book big orders as surprises: budget decisions run on the no-big-order norm; for setups propped by one, extend observation and look at a cycle without the luck.
  4. Extreme weeks trigger drill-downs, not decisions: seeing 61.6 or 0.0, step one is always the daily-level distribution — never the budget slider.

One line to remember

Attribution clustered in one week, falling back after, inquiries unchanged = big-order illusion. Budget on the norm of inquiries, not the luck of orders; extreme weeks trigger drill-downs, not decisions.

FAQ​

How do I tell if ROI is propped up by a big order?​

Split the window and look: attribution clustered in one week, falling back after, with inquiry volume unchanged — the acquisition ability never changed; that week's luck did.

Why can't I budget on a big-order week's ROI?​

It doesn't repeat. Measured case: one week at ROI 61.6, the other eight between 0 and 2.1 — scale budgets on 61.6 and the next cycle arrives before the luck does.

What if attributed orders suddenly drop to zero?​

Check inquiries first: unchanged inquiries mean the luck receded — decide on normal efficiency; shrinking inquiries mean real decay — entirely different treatment.

That "segmented windows + inquiry thermometer + daily drill-down" method is built into AI Operations — LLM-powered analysis that automatically surfaces market trends, user behavior, and sales data to drive strategy. Extremely beautiful numbers deserve verification before belief.

CCLEE

Independent developer, 24 years in e-commerce, focused on grounding AI in real business scenarios.

Work with me

Is 16 Days Enough for Marketplace Ad Data? We Re-Collected 5 Weeks to Find Out

· 8 min read

TL;DR​

"Wait 16 days before judging" is not enough. We re-collected five long-settled weeks of B2B marketplace ad data and diffed every row: finished weeks keep growing new data as late as day 29 (keywords added late reach back and pull their history in), yet past day 40, even a full re-collection changed not one existing number; and the harsher finding is retention — the platform deletes detail reports after ~6–7 weeks, and past 54 days none of our 6 campaigns could be recovered. Collect weekly, archive locally — what that buys is twofold: no more mis-calls on irreversible decisions, and ad history that actually belongs to you.

The trigger: 27 new rows on day 29​

During a routine collection on August 10, one industrial-goods B2B storefront (anonymized) wrote 27 new region-detail rows into a week that had ended 29 days earlier — five campaigns, a +22% row increase. By the folklore timeline, that week had "settled at 16 days" roughly two weeks prior.

Is 16 days platform fact or folklore? The test is ready-made: re-collect older weeks. If long-settled weeks grow new rows or rewrite old values, 16 days is nowhere near the end — and the same re-collection measures how long the platform keeps detail at all. On August 21 we re-pulled five older weeks, which became the experiment below.

The experiment itself came out of a data-semantics check while building AI Operations.

The design: re-collect five "settled" weeks and diff​

The method is plain: baseline → re-collection → row-level diff.

  1. Pick 5 consecutive weeks (2026-06-08 ~ 07-06), already 40–74 days old — settled under any definition;
  2. Export a baseline across 4 detail tables (overview / products / keywords / areas — 1,199 rows);
  3. Trigger a full platform re-collection over an explicit date range, export again;
  4. Diff row by row: additions, value changes, deletions. Re-checked on 2026-09-02; conclusions unchanged.

The result:

TableBaseline rowsAfterAddedValue changesRemoved
Campaign overview1818000
Product detail682682000
Region detail405405000
Keyword detail94105+1100

The net result of the row-level diff across all 1,199 rows: not a single historical number was revised (zero value changes, zero removals). The only difference is 11 new keyword rows — not revised values but newly grown entities, a second back-fill mechanism covered in Finding 2; which campaigns' detail the platform re-returned at all is Finding 3.

Lifecycle of 1688 P4P data: from routine writes to back-fill events to deletion

Finding 1: values blow past 16 days — and truly settle by 40​

The boundary of the claim: "16 days" is both right and wrong.

  • The platform's stated attribution window is ~15 days — the folk "settle at 16" comes from there;
  • The trigger event was exactly that counterexample: +27 region rows on day 29, nearly double the window;
  • The experiment supplies the stopping point: re-collection at 40–47 days showed zero value changes — back-fill does stop, just much later than day 16.

For operators: 16 days works as a "probably stable" heuristic, not as a definition of final. For irreversible calls like pausing a campaign, wait the full 4–5 weeks so the decision lands past the measured stopping point. The full pre-pause checklist is in Five checks before you pause a marketplace ad campaign.

Finding 2: keywords you add later reach back and pull their history in​

Plain version first: how much "history" you can recover depends on which keywords you are running now.

The test store added keywords to one campaign mid-flight. On re-collection, the platform returned the campaign's current keyword list together with past weekly performance — one recovered historical week carried 11,179 impressions, ¥342.7 spend, 10 inquiries, and 11 orders. Those numbers sat on the platform's side all along; they only appear when you come collect.

(Technical note: the mechanism is entity-level back-fill — the API returns history for your current keyword list. Row counts rise; existing values never change. So when re-collected data grows, first tell revised values from newly grown rows: the former is a warning, the latter is a gift.)

But the gift has a precondition: the platform must still retain the detail. For how long? That is Finding 3.

Finding 3: the real cliff is deletion, not settlement​

Slow back-fill costs waiting; retention costs everything. We checked, per campaign, whether the platform could still return detail during re-collection:

  • Weeks aged 40–47 days: from the same batch of six campaigns, only 3 still returned complete detail;
  • Weeks aged 54–74 days: the same six again — 0 of 6; not stale, deleted from the platform's side, unrecoverable by any means;
  • The 3 campaigns already returning nothing at 40 days run a shorter, per-campaign retention — they hit the cliff earlier; one sample so far.

The retention cliff: 3/6 campaigns retrievable at 40–47 days, 0/6 after 54

Put bluntly: any detail not in your own database within ~6 weeks has been deleted on your behalf. "I'll export it later" is not procrastination — it is deletion.

What these findings are worth: two ledgers​

The mis-kill ledger. The biggest cost of the 16-day folklore is reading "the data hasn't arrived" as "the campaign doesn't work". The keyword whose history grew back in Finding 2 carries 10 inquiries and 11 orders in a single week — pause a ramping campaign early, and that is the weekly loss, before counting the extra settlement weeks it needs to prove itself again. Waiting the full 4–5 weeks buys "no wrong irreversible calls".

The loss ledger. The platform deletes detail after ~6–7 weeks, so "I'll back-fill later" is a promise with an expiry date. Which regions fed you inquiries last quarter, which keywords' costs were quietly climbing — only people whose data landed in their own database can answer. A weekly collection run buys "history, always queryable".

Three disciplines for operators​

  1. Watch trends anytime; wait 4–5 weeks for irreversible calls. 16 days is the reference line, 29 the safety line.
  2. Collect or export detail weekly and archive locally. Platform-side detail is far shorter-lived than assumed; the archive is the asset you own.
  3. Close the "settlement tail" (the data each finished period is still quietly back-filling) before any retro analysis. Period ROI without its back-fill is systematically understated — and when you cannot tell revised values from newly grown rows, use the experiment's move: diff at row level.

For data engineers: leave at least 5 weeks of overlap​

If your pipeline collects marketplace ad reports weekly (or models attribution for a similar platform): the stated 15-day window actually back-fills as late as +29 days. A strict weekly cadence with a 4-week overlap performs its last rewrite at week-end +22 days — which fails to cover the measured day-29 event (ours survived only because a collection gap happened to stretch). Raise the overlap from 4 weeks to 5: ~20–25% more requests each week — every extra overlap week is one more week of paginated detail to pull — in exchange for never losing rows.

Two weekly habits

  • Collect: land this week's detail in your own database, with a 5+ week overlap — archiving is the only defense against deletion;
  • Wait: make keep-or-stop calls only on data settled 4–5 weeks — settlement is the only defense against misjudgment.

FAQ​

How long until marketplace ad data is final?​

The platform states a ~15-day attribution window, but the latest back-fill we observed landed on day 29. Wait 4–5 weeks before irreversible keep-or-stop calls.

Can I back-fill ad history I forgot to export?​

Usually no. Detail reports survive roughly 6–7 weeks: at 40–47 days only 3 of 6 campaigns still returned data; past 54 days, none. After that it is gone.

Why do new keyword rows appear after re-collection?​

The API returns history for your current keyword list. Keywords added late grow historical rows — entity-level back-fill: row counts rise, existing values never change.

That baseline → re-collection → row-level diff method is built into AI Operations — LLM-powered analysis that automatically surfaces market trends, user behavior, and sales data to drive strategy. Ad data is the foundation of all of it — and a foundation deserves to be measured.

CCLEE

Independent developer, 24 years in e-commerce, focused on grounding AI in real business scenarios.

Work with me

Same Product, 4× the Inquiry Cost: It's the Ad Solution, Not the Product

· 6 min read

TL;DR​

The same product, listed in different marketplace ad solutions, can cost 4× more per inquiry. Weekly ad records for three key products in one store show the same product ranging from ¥17 to ¥78 — and the ordering is dictated by the solution: the merchant growth program is the most expensive for all three products, and switching solutions moves one product's cost by up to 4×. Three rules fall out: grade products on the blended number, grade solutions on same-product same-period comparisons, and never compare solution numbers across non-overlapping time windows.

The situation: the product you're about to pause was wronged by its channel​

Monday review: a product's inquiry cost in your flagship solution looks terrible, and you're considering pausing it. Hold on — the real ledgers of three key products in one industrial-goods store (anonymized) show the same product ranging from ¥20 to ¥78 depending on the channel it enters through. Same product, same page, same price. This reconciliation came out of a weekly-ad-ledger audit while building AI Operations.

Why: for the same product, the solution sets the price​

Line up the weekly records of all ad solutions for three key products. Full-history view first (inquiry cost = cumulative spend ÷ cumulative inquiries):

Ad solutionDelivery windowProduct AProduct BProduct C
Whole-store promotion2024-04 ~ 2026-06 (68–116 wks)¥30¥37¥26
Site-wide, shop-boosting2025-11 ~ 2026-06 (31–33 wks)¥25¥25¥17
Merchant growth programsince 2026-06-29 (8 wks)¥77¥78¥49
New-customer crowdsame period (8 wks)¥50¥20¥29
Cross-border expresssame period (7–8 wks)¥42¥22¥28

Three layers of structure, each more useful than the last:

1. Full history: priciest vs cheapest is 4×+ — ¥78 against ¥17.

2. The ordering is dictated by the solution. The "Merchant growth program" is the most expensive for all three products (¥49–78) — three completely different products, uniformly expensive in this one solution. The dominant factor is the solution (what traffic it buys), not the product (what it sells).

3. Within the same period: 1.8–3.9×. The last three solutions share one time window (8 weeks from 2026-06-29), so their comparison is clean: 1.8× for Product A, 3.9× for Product B, 1.8× for Product C.

Same product, three ad solutions: inquiry cost comparison

(Technical note: the first two solutions' data ends on 2026-06-29 and the last three start that very day — the windows don't overlap. So "old ¥25 vs new ¥77" mixes two factors: solution differences and market seasonality; concluding directly misleads. Statistically this is kin to Simpson's paradox — conclusions consistent per layer can flip once merged. Every "n×" claim in this article comes from the same-period window only.)

One counter-intuitive detail: the "Merchant growth program" isn't cold-start expensive — it keeps getting more expensive. Across its 8 weeks, inquiry cost climbed from ¥26 to ¥107. That retires the "give the new solution time" excuse; money dictated by traffic structure does not arrive with waiting.

What it's worth: two ledgers​

The mis-kill ledger. Product B runs at ¥20 per inquiry in "New-customer crowd," about a dozen-plus inquiries a month. Pause the product because it shows ¥78 in the growth program, and what you discard is not a bad product — it's a cheap channel still delivering steadily.

The true-cost ledger. Which of Product B's five numbers (¥37 / ¥25 / ¥78 / ¥20 / ¥22) is real? All of them, and none. Its actual acquisition cost is the blended one: ¥32,265 total spend ÷ 911 inquiries = ¥35. A single-solution number can overstate or understate a product; only the blend is the product's real price tag — and the stable anchor for budget allocation.

Disciplines for operators​

  1. Grade products on the blend: total spend ÷ total inquiries. Per-solution numbers answer "is this channel expensive," never "is this product good."
  2. Grade solutions on same-product, same-period comparisons: fix a basket of products and a time window; only then does the ordering mean anything.
  3. Never compare across non-overlapping windows: solution handover periods are the danger zone — an old solution's historical cost is not the new one's ruler.
  4. A persistently worsening solution isn't worth waiting for: cut budget after 4+ weeks of climbing costs; make keep-or-stop calls with the settlement discipline from Is 16 Days Enough for Marketplace Ad Data? and the full pre-pause checklist in Five checks before you pause.

One line to remember

Products get the blend; solutions get the same period. Before comparing costs across windows, align the time.

FAQ​

The same product shows very different inquiry costs across ad solutions — is that normal?​

Yes. Across three key products we measured ¥17 to ¥78 for the same product, and the ranking followed the solution, not the product — each solution buys different traffic.

Which inquiry cost should I use to judge a product?​

The blended one: total spend across all solutions ÷ total inquiries. A single-solution number only says what that channel pays for this traffic — it cannot grade the product.

Can I compare costs across different ad solutions directly?​

Only within overlapping time windows. When old and new solutions don't share dates, the gap mixes solution differences with market seasonality — comparing directly misleads.

That "line up every channel for the same product" reconciliation is built into AI Operations — LLM-powered analysis that automatically surfaces market trends, user behavior, and sales data to drive strategy. Every product's real acquisition cost deserves to be computed once, fully.

CCLEE

Independent developer, 24 years in e-commerce, focused on grounding AI in real business scenarios.

Work with me

A Real Store-Wide Efficiency Alert, From a 40-Week Ledger

· 5 min read

TL;DR​

A real ledger from one store: across 29 historical weeks, store-wide inquiry cost held a stable band of ¥25–31. After new ad solutions launched on June 29, it ran at ¥34–49 for 7 consecutive weeks — then hit ¥69 in week eight. Store-wide decay never announces itself as an incident; it shows up as "every campaign looks fine." So you need a ruler for the sum: a band set by your own history, and an alert on 3 consecutive weeks above it. Had the alarm fired in week three, roughly ¥7,400 of the overspend in the following five weeks would have been avoided.

The situation: every campaign "fine," the sum deteriorating​

Per-campaign reviews read as usual: this campaign's numbers match last month's, that one is stable, the new one is still ramping. All passable.

At the store level: every July week sat above ¥40 per inquiry — in the previous 29 weeks, that line had been crossed exactly once (the Spring Festival week). This post-mortem came out of a store-ledger audit while building AI Operations.

Store-wide inquiry cost across 40 weeks: historical band and the breach

The method: set the band from your own history​

The baseline is not the industry and not a target — it is your own past:

  1. Take weekly store-wide inquiry cost (weekly spend ÷ weekly inquiries) for the past six months
  2. Exclude abnormal weeks: here, two kinds — the Spring Festival week (¥121, spend collapse masquerading as expensiveness) and a 3-week delivery gap in early June (weekly spend ¥90–281, near-dark)
  3. The remaining 26 weeks land between ¥20–41, concentrated in ¥25–31 — that band is "normal"
  4. Alert condition: 3+ consecutive weeks above the band's top, with spend not shrinking (cost rising because spend is collapsing is a different problem)

(Technical note: why "consecutive 3 weeks" rather than any single week — in the 29 historical weeks, single-week breaches happened 4 times, all noise; consecutive breaches happened zero times. The baseline tells you how strict the threshold should be.)

The full post-mortem​

  • June 8–22: a 3-week gap. Weekly spend fell from ~¥1,900 to ¥90–281 — near-dark
  • June 29: new ad solutions launched (the solution-switch details and data are in Same Product, 4× the Inquiry Cost)
  • From June 29: 7 consecutive weeks above the band — ¥46 / 43 / 40 / 43 / 40 / 49 / 34, every one above the historical top of ¥31
  • Week of August 17: ¥69, as spend spiked to ¥5,210 without inquiries following

A gap-and-restart is not a return to the old normal: the environment changed and the solutions changed — the old cost level no longer applies. That is exactly the kind of account-level shift single-campaign views cannot see, and only the store line exposes.

What it's worth: the overspend ledger​

The 7 breached weeks (Jun 29 – Aug 10) spent ¥24,699 for 592 inquiries — ¥41.7 each. At the historical level (¥28), the same inquiries would have cost ¥16,576: about ¥8,100 of overspend in 7 weeks. Had the alarm fired in week three and intervention started in week four, roughly ¥7,400 of the last five weeks' overspend was avoidable. That is the price of the ruler: set it once, watch one number.

Disciplines for operators​

  1. The band must come from your own history: at least six months of weekly data, abnormal weeks excluded. A three-week average as a ruler is worse than no ruler.
  2. 3 consecutive weeks above the top, with spend holding, = alert: single weeks are noise; consecutive weeks are structure.
  3. After an alert, hunt the common cause first: solution switches, gap-and-restarts, category-wide competition shifts are account-level events — fixing campaigns one by one treats symptoms.
  4. The cost numbers themselves must settle first: weekly data carries a settlement tail; the discipline is in Is 16 Days Enough for Marketplace Ad Data? We Re-Collected 5 Weeks to Find Out.

One line to remember

Set the band from history (six months, abnormal weeks out); 3 consecutive weeks above it = alert. Alerts trigger a hunt for account-level causes; actions stay at campaign level.

FAQ​

What signals a store-wide ads efficiency decline?​

Weekly inquiry cost above your own historical band for 3+ consecutive weeks while spend holds — each campaign can look passable alone while the sum is sinking.

How do I set the 'normal band'?​

Take your own half-year of weekly inquiry costs, exclude abnormal weeks (holidays, delivery gaps), and use the median ±10% as the band. Stores with thin history should accumulate first.

What is the first move after an alert?​

Hunt for an account-level cause before touching campaigns: solution switches, gap-and-restart episodes — store-wide breaches are usually account-level events.

That "band from history, watch the store" alerting logic is built into AI Operations — LLM-powered analysis that automatically surfaces market trends, user behavior, and sales data to drive strategy. Efficiency decay should not wait for a quarterly review to be discovered.

CCLEE

Independent developer, 24 years in e-commerce, focused on grounding AI in real business scenarios.

Work with me

Life Enters Limited Beta: A Money & Mood Journal You Just Talk To

· 3 min read

Anyone who has kept a journal of expenses knows: the hard part isn't logging, it's sticking with it. Open the app, find the form, pick a category, type the amount, save — every step talks you out of it. Mood and medication notes are even more scattered: by the time you remember, the moment has passed.

Life's answer: just say it, like a text message.

  • "lunch cost me 32 today" → a Food expense with amount, date, and category already in place
  • "lent Li 2,000 last month, they paid back 500 today" → a ledger entry, outstanding 1,500 computed for you
  • "how much did I spend on food this month?" → no list-digging, just ask

👉 Try Life: life.ccleeai.com | Documentation: aidevhub.ai/docs/life

Clicking One Row Highlights Many in Ant Design Table? Your rowKey Isn't Unique

· 5 min read

Clicking a table row on a reporting page to inspect details, the clicked row — plus several other rows — highlighted at the same time, while the console flooded with React duplicate key warnings.

Encountered this while building AI Ops — LLM-powered analysis that surfaces market trends, user behavior, and sales insights to drive precise operations strategy. The ad weekly-report page contains several detail tables (campaign × keyword, campaign × area), each supporting click-to-highlight so users can drill into a row's delivery details. After launch, clicking any row "lit up" every row under the same campaign.

TL;DR​

Ant Design's Table uses the return value of rowKey as each row's React key. When that field isn't the data's real business key — guessed by naming, or missing a dimension that participates in uniqueness — multiple rows generate identical keys: every key-matched row interaction (selection, highlight, expansion) hits multiple rows at once, and React throws duplicate key warnings. The fix: query information_schema for the table's actual columns, pick a truly unique column or composite columns for rowKey, and verify with GROUP BY HAVING.

The Symptom​

A minimal reproduction (Ant Design 5 + React 18):

import { Table } from 'antd';
import { useState } from 'react';

// Data granularity: keyword × product — one keyword splits into multiple rows per promoted product
const data = [
{ keyword_id: 88, keyword: 'summer dress', offer_id: 101, clicks: 12 },
{ keyword_id: 88, keyword: 'summer dress', offer_id: 102, clicks: 7 },
{ keyword_id: 90, keyword: 'maxi skirt', offer_id: 103, clicks: 5 },
];

export default function WeeklyKeywords() {
const [selected, setSelected] = useState<string[]>([]);
return (
<Table
rowKey={(r) => String(r.keyword_id)} // Pitfall: keyword_id is not unique at this granularity
columns={[
{ title: 'Keyword', dataIndex: 'keyword' },
{ title: 'Product', dataIndex: 'offer_id' },
{ title: 'Clicks', dataIndex: 'clicks' },
]}
dataSource={data}
rowSelection={{ selectedRowKeys: selected, onChange: setSelected }}
onRow={(r) => ({ onClick: () => setSelected([String(r.keyword_id)]) })}
/>
);
}

Two symptoms: clicking the first row highlights both rows with keyword_id 88, and the console repeatedly prints:

Warning: Encountered two children with the same key, `88`.
Keys should be unique so that components maintain their identity across updates.

Root Cause​

An antd Table row's identity is exactly the return value of rowKey. It becomes the React key of that row's element. When keys duplicate, React's diff treats multiple rows as the same element: rendering can go wrong and controlled state bleeds between rows.

Every key-matched row interaction gets amplified. rowSelection's selectedRowKeys, onRow clicks, and expandedRowKeys all match by key — with duplicate keys, one match hits multiple rows. That's the direct cause of "click one, highlight many."

The wrong key column usually comes from the data side. The actual root cause here: key columns were guessed from naming — we assumed the area table had region_id, but its real business key column was area_name; we assumed the keyword table's granularity was keyword, but it was actually keyword × product (offer_id also participates in the unique key). Miss one dimension and every row in a group shares the same rowKey.

The Fix​

Step 1: Query the table's actual columns — don't guess from names​

SELECT column_name, data_type
FROM information_schema.columns
WHERE table_name = 'ad_weekly_keywords'
ORDER BY ordinal_position;

Confirm which columns actually form the business key, and whether the column you assumed even exists.

Step 2: Verify the key (or key combination) is unique​

SELECT keyword, offer_id, COUNT(*)
FROM ad_weekly_keywords
GROUP BY keyword, offer_id
HAVING COUNT(*) > 1;
-- 0 rows = unique; also confirm key columns contain no NULLs

Step 3: Configure rowKey with a composite key​

<Table
rowKey={(r) => `${r.keyword}::${r.offer_id}`}
// Or even safer: JSON.stringify([r.keyword, r.offer_id])
dataSource={data}
...
/>

When concatenating composite keys, use a separator that cannot appear in the field values (or just JSON.stringify the array) to avoid collisions between a + b and ab.

After the change, clicking a row highlights only that row, and duplicate key warnings drop to zero. This is the same family of problems as React list key duplicates causing DOM errors — only with unique keys can diffing and row interactions be correct.

Heads up

Before choosing key columns, query information_schema for the table's actual columns — don't guess from field names: business key columns can differ entirely from intuition (the table has only area_name, no region_id; keyword granularity is actually keyword × product).

The duplicate key warning is not "just a warning": it means React reconciliation is broken — row state bleeding, wrong highlights, and updates not taking effect can all follow. It must go to zero.

After changing rowKey, re-verify uniqueness with GROUP BY ... HAVING COUNT(*) > 1, and check key columns for NULLs — NULL keys create duplicates too.

FAQ​

How should I set rowKey on an Ant Design Table?​

Use the field — or combination of fields — that uniquely identifies a row: a single unique field works directly; if no single field is unique, build a composite key from multiple columns (with a collision-proof separator or JSON.stringify). Never use a non-unique business field, and don't take shortcuts with array index — after sorting, filtering, or pagination, state will bleed between rows.

How do I fix the React duplicate key warning?​

Duplicate keys make React treat multiple nodes as the same element, corrupting rendering and state. Locate the list rendering site producing the duplicate keys, switch to a truly unique key, then verify uniqueness at the data source with GROUP BY HAVING — silencing the warning without checking the data means the problem will resurface in another form.

CCLEE

Independent developer, 24 years in e-commerce, focused on grounding AI in real business scenarios.

Work with me

DeepSeek/Qwen Structured Calls Succeed but Return Empty? Disable Thinking Before It Burns Your Token Budget

· 6 min read

While running batch LLM semantic validation over production data, all 2,449 calls "succeeded" — yet every single result fell back to the default value, and the logs contained almost no failures.

Encountered this while building AI Ops — LLM-powered analysis that surfaces market trends, user behavior, and sales insights to drive precise operations strategy. The product title optimization pipeline asks an LLM to semantically validate "keyword × product" pairs: one call per keyword batch, returning a tiny JSON judgment. This should be the simplest kind of LLM call — yet on the first production run, all 2,449 pairs silently fell back, and the validation layer produced zero effective LLM judgments.

TL;DR​

Thinking-capable models like DeepSeek and Qwen share a single max_tokens budget between reasoning text and final content. If structured small-output calls don't explicitly disable thinking, the reasoning chain eats the entire budget on its own: finish_reason becomes length, content comes back empty — and since the parse-retry loop only logs exceptions, the silent retries exhaust and degrade the whole batch without a single error. Two things to do: disable thinking explicitly for structured calls, and log finish_reason on parse failure instead of only catching exceptions.

The Symptom​

This minimal reproduction shows the whole process (requires pip install openai and a thinking-capable model):

import os
from openai import OpenAI

client = OpenAI(
api_key=os.environ["DEEPSEEK_API_KEY"],
base_url="https://api.deepseek.com",
)

resp = client.chat.completions.create(
model="deepseek-reasoner",
messages=[
{
"role": "user",
"content": (
'判断下面的关键词是否适合写入商品标题,'
'只返回 JSON:{"suitable": true} 或 {"suitable": false}。\n'
"关键词:summer women dress\n"
"商品:floral midi dress for women"
),
}
],
max_tokens=2000, # reasoning and content share this budget
)

print("finish_reason:", resp.choices[0].finish_reason)
print("content:", repr(resp.choices[0].message.content))

Typical output with thinking enabled:

finish_reason: length
content: ''

Everything looks fine at the HTTP layer: no timeout, no 5xx, no SDK exception. If the outer layer is a "retry on parse failure, fall back after retries exhaust" loop, the logs end up with a single "retries exhausted" warning — which reads exactly like an intermittent network problem.

Root Cause​

Thinking models don't budget reasoning separately. In DeepSeek, Qwen and similar models, reasoning content and final content share the same max_tokens ceiling — there is no independent "reasoning budget" field.

Structured small-output calls tend to use small budgets. A boolean judgment is expected to produce a few dozen tokens, so max_tokens=2000 looks generous — but the reasoning chain's length is completely uncontrolled. Once it consumes all 2,000 tokens, the model never gets a chance to emit the actual answer: finish_reason returns length, content is an empty string, yet the API returns 200 normally.

There's also an amplifier on the engineering side. An empty string is not valid JSON, but many retry loops only catch network and API exceptions, treating parse failure as "no result this round" and retrying silently. This kind of exception-swallowing silent failure is notoriously hard to diagnose inside retry loops: all N retries fail for the same root cause, yet the log shows only the final fallback warning — easy to misread as network flakiness.

The Fix​

Step 1: Explicitly disable thinking for structured-output calls​

Boolean judgments, JSON extraction, and classification calls don't need multi-step reasoning. Disable thinking per provider:

def thinking_disabled_extra_body(provider: str) -> dict:
"""Structured small-output calls: disable thinking explicitly per provider."""
if provider == "deepseek":
return {"thinking": {"type": "disabled"}}
if provider == "qwen":
return {"enable_thinking": False}
return {}


resp = client.chat.completions.create(
model=model_name,
messages=messages,
max_tokens=2000,
extra_body=thinking_disabled_extra_body("deepseek"),
)

With thinking off, the entire 2,000-token budget goes to the JSON judgment itself. After the fix, rerunning the batch returned valid judgments for all 2,449 keyword×product pairs — before the fix, the whole batch silently degraded, leaving just 3 "retries exhausted" warnings in the logs.

Step 2: Don't let parse failures go silent​

Even with thinking disabled, turn "parse failure" into an evidence-bearing log line, so next time an empty output occurs — for any reason — a single log line locates it:

import json
import logging

logger = logging.getLogger(__name__)


def parse_judgment(resp) -> dict | None:
content = resp.choices[0].message.content
try:
return json.loads(content)
except (TypeError, json.JSONDecodeError):
# Key: log finish_reason and raw content, not just exceptions
logger.warning(
"LLM output parse failed: finish_reason=%s content=%r",
resp.choices[0].finish_reason,
content,
)
return None

When debugging LLM degradation, check finish_reason first: length means the output budget was exhausted (most likely by reasoning), while stop means normal completion. This is far more effective than scrolling exception logs.

If your returned JSON also passes through schema validation and you use Zod in a TypeScript project, watch out for this related pitfall: Zod schema validation silently dropping LLM output.

Heads up

The parameter to disable thinking is not standardized across providers: DeepSeek uses {"thinking": {"type": "disabled"}}, Qwen uses {"enable_thinking": False}, and OpenAI's o-series uses reasoning_effort-style parameters. Check each provider's docs before integrating — don't assume the parameter is universal.

If a provider's thinking cannot be disabled, you must raise max_tokens based on measured reasoning length — otherwise the same silent degradation will happen again.

FAQ​

How do I disable thinking output in DeepSeek?​

On the OpenAI-compatible API, pass {"thinking": {"type": "disabled"}} via extra_body; Qwen uses {"enable_thinking": False}. Structured small-output calls like boolean judgments and JSON extraction don't need a reasoning chain — turn it off by default so the entire budget goes to the actual answer.

Why does my DeepSeek API call succeed but return empty content?​

Thinking models share max_tokens between reasoning and content. When reasoning exhausts the budget, finish_reason returns length and content is empty — with no exception thrown. Disable thinking or raise max_tokens based on measured reasoning length, and check finish_reason before digging through exception logs.

CCLEE

Independent developer, 24 years in e-commerce, focused on grounding AI in real business scenarios.

Work with me

Docker Volume Mounted but Empty? etcd Data Lived in the Writable Layer, Lost on Recreate

· 6 min read

During a disk-cleanup pass, docker ps -as showed the etcd container's writable layer at 385MB — while its mounted data volume held just 8K. We didn't dare touch it during cleanup: one recreate, and every piece of Milvus metadata would evaporate with the writable layer.

Encountered this while building AI Customer Service — a 24/7 AI support agent answering product questions; its knowledge-base retrieval runs on Milvus, and all of Milvus's metadata lives in this etcd.

TL;DR​

compose's volumes: only guarantees "the volume is mounted at this path" — where the app writes is decided by its own parameters. etcd without a data-dir defaults to default.etcd under its working directory: it never used the /etcd mount, and all 385MB sat in the writable layer, one docker compose up -d rebuild away from oblivion. The fix follows etcd's native migration path: online snapshot → restore → swap data during a brief stop → add ETCD_DATA_DIR → recreate with the volume. Zero data loss, under two minutes of downtime.

Symptoms​

The etcd service in compose looks "correct" — the volume is declared:

  etcd:
image: quay.io/coreos/etcd:v3.5.5
environment:
- ETCD_AUTO_COMPACTION_MODE=revision
- ETCD_LISTEN_CLIENT_URLS=http://0.0.0.0:2379
# ... nothing about data-dir
volumes:
- etcd_data:/etcd

But two numbers disagree:

$ docker ps -as --format "{{.Names}}\t{{.Size}}" | grep etcd
rag-service-etcd-1 385MB (virtual 199MB) ← writable layer: 385MB

$ docker exec rag-service-etcd-1 du -sh /etcd
8K /etcd ← the mounted volume: empty

$ docker exec rag-service-etcd-1 ls -la / | grep etcd
drwx------ 3 root root 4096 default.etcd ← the data is here: container root

Root Cause​

Mounted ≠ used. volumes: etcd_data:/etcd only mounts the volume at the path /etcd; where etcd writes is decided by its --data-dir parameter. etcd's default data-dir is default.etcd under the working directory — this compose set neither the ETCD_DATA_DIR env var nor a --data-dir flag, so etcd created default.etcd at the container root and wrote everything into the writable layer. The /etcd mount point had been empty since day one.

The nastiest property of a "decorative volume" is that it's completely symptom-free: the service runs, reads and writes work, dashboards stay green. It only bites at the moment you run docker compose up -d --force-recreate (config change, writable-layer recycling, host migration) — the layer is discarded wholesale and the data vanishes, precisely when you're doing urgent ops work and can least afford a second incident.

Solution​

Use etcd's native snapshot migration: the online snapshot guarantees consistency, and downtime only happens at the final data swap.

Step 1: Consistent online snapshot​

docker exec rag-service-etcd-1 sh -c \
'ETCDCTL_API=3 etcdctl --endpoints=http://127.0.0.1:2379 snapshot save /tmp/etcd-snap.db'
docker cp rag-service-etcd-1:/tmp/etcd-snap.db /root/etcd-snap.db

snapshot save is safe against a running etcd (it goes through the Raft backend, not file copying) — no stop, no write lock.

Step 2: Restore into the target data-dir structure​

docker run --rm -v /root:/host quay.io/coreos/etcd:v3.5.5 \
etcdctl snapshot restore /host/etcd-snap.db --data-dir /host/etcd-restored

Restore produces a full member/ data directory (a raw snapshot file cannot be used as a data-dir directly).

Step 3: Brief stop, move data into the volume​

docker stop rag-service-etcd-1
rm -rf /var/lib/docker/volumes/etcd_data/_data/* # volume is empty; clear mount residue
cp -a /root/etcd-restored/. /var/lib/docker/volumes/etcd_data/_data/

Step 4: Add the missing config, recreate with the volume​

  etcd:
environment:
- ETCD_DATA_DIR=/etcd # ← the missing line
# ...
volumes:
- etcd_data:/etcd
docker compose up -d etcd    # config change triggers recreate; new container uses the volume

Step 5: Verify​

docker exec rag-service-etcd-1 etcdctl --endpoints=http://127.0.0.1:2379 endpoint health
du -sh /var/lib/docker/volumes/etcd_data/_data # data should be in the volume
docker ps -as | grep etcd # writable layer should drop to KB scale

After this fix: writable layer 385MB → 8KB, 123MB of data living on the volume, Milvus reconnected automatically with metadata reads and writes healthy.

Notes

  • When auditing stateful containers (etcd/postgres/redis/minio), make "writable layer size vs volume size" a standing check: an inflated SIZE in docker ps -as with an empty volume almost always means data isn't on the volume.
  • To find where data actually lives, trace the path the process really reads and writes (docker top for args, look for data dirs inside the container) — never trust the mere presence of a volumes: line; declared is not used.
  • Single-node etcd snapshot restore regenerates member metadata and is only valid for single-node setups; multi-node cluster migrations go through member change procedures instead.
  • The same inspection method appears in Container Logs Filling Your Server Disk? docker system df 'Reclaimable' Lies — docker ps -as writable-layer watching is the same knife; and for the other flavor of mount surprise, see Docker Volume Override Bind Mount.

FAQ​

Why is my Docker volume mount empty?​

Two usual causes: an empty volume shadows whatever the image had at that path (documented Docker behavior); or the application's data-directory setting never pointed at the mount, so data went to the container's writable layer — the etcd case in this post. du on the volume vs the writable layer tells them apart instantly.

How do I safely migrate etcd to a new data-dir?​

etcdctl snapshot save for a consistent online snapshot (no downtime), etcdctl snapshot restore --data-dir to build a directory with proper member structure, stop the container, place the data, start with the new data-dir, then verify with endpoint health plus upstream reconnection. Downtime is only the swap itself.

I declared a volume in compose — why isn't it used?​

volumes: mounts the volume at a container path; the application decides where to write from its own config — etcd's data-dir, postgres's PGDATA, redis's dir. When the mount point and the app config disagree, the volume is an empty directory and all data sits in the writable layer, lost on recreate.

CCLEE

Independent developer, 24 years in e-commerce, focused on grounding AI in real business scenarios.

Work with me