Kill bad ideas in days, not months
For founders about to spend months building something nobody asked for. Describe your idea, pre-register what validated means, set a hard budget cap. An agent tests audience x message x offer cells against real strangers and returns a verdict with evidence, not vibes.
14-day free trial, no card required. Ad spend runs on your own accounts at cost, no markup.
| cell | audience x message x offer | spend | evidence | ev / $ | status |
|---|---|---|---|---|---|
| C-04 | solo agencies · "books closed in 20 min" · survey + call | $61.20 | 212 | 3.46 | leading |
| C-02 | solo agencies · "never touch a spreadsheet" · waitlist | $38.75 | 74 | 1.91 | running |
| C-07 | etsy sellers · "books closed in 20 min" · survey | $29.90 | 43 | 1.44 | running |
| C-05 | bookkeepers · "10 clients, one dashboard" · call | $22.40 | 21 | 0.94 | defunded |
| C-01 | freelance devs · "stop dreading taxes" · waitlist | $24.10 | 19 | 0.79 | killed d3 |
| C-03 | freelance devs · "never touch a spreadsheet" · survey | $11.05 | 6 | 0.54 | killed d2 |
Signups are not demand
Email lists lie, because people sign up for anything free. So every conversion is weighted by how much it actually proves. A raw signup counts 1x. A verified email, 3x. A qualifying survey match, 5x. A booked call or deposit, 10-25x.
The optimizer maximizes evidence per dollar, so it cannot cheat by farming junk signups from cheap traffic. When a segment shows promise, the agent escalates the ask, survey, call, deposit, instead of just scaling spend.
ladder for cell C-04, day 6:
2 booked calls outweigh
24 raw signups. that is the point.
| rung | weight | count | weighted | share of score |
|---|---|---|---|---|
| email signup | 1x | 41 | 41 | 19% |
| verified email (double opt-in) | 3x | 29 | 87 | 41% |
| qualifying survey match | 5x | 12 | 60 | 28% |
| booked call | 12x | 2 | 24 | 11% |
| deposit | 25x | 0 | 0 | escalation queued for approval |
d1 09:12 deploy 6 pages live on try.ledgerpilot.dev, disclosure + double opt-in on all
d1 09:14 queue 4 community post drafts -> awaiting your approval
d2 18:40 kill C-03: 0 verified emails at 210 visitors, ceiling < floor
d3 08:05 kill C-01: ev/$ 0.79 after $24.10, 3 consecutive losing days
d3 08:05 shift reallocate $9.60/day C-01 -> C-04 (posterior best arm, p=0.81)
d4 11:32 escalate C-04: add booked-call ask for survey matches, spend unchanged
d5 07:00 guard daily cap check: $38.20 of $40.00, within cap
d5 14:11 queue google search campaign draft ($12/day) -> awaiting your approval
d6 07:00 guard tracking heartbeat ok, dead-man switch armed, 12h trip wire
d6 09:30 shift C-05 defunded to floor spend, C-04 now 46% of daily budget
A bandit defunds losers cell by cell
Your idea becomes a grid of cells, each one an audience, a message, and an offer. A bandit reallocates budget toward whatever earns the most evidence per dollar, so weak hypotheses lose funding in days instead of quietly eating months.
Autonomy is on a leash. The agent deploys pages, mutates variants and shifts budget within your caps on its own. Anything outbound, community posts or ads, and anything spend-increasing waits in your approvals queue. Hard daily and total caps are enforced in code, and a dead-man switch pauses all paid campaigns if signup tracking goes dark for 12 hours.
traffic order is fixed:
free channels first, community
drafts and SEO pages, then
meta and google ads from
your own accounts at cost.
It learns which feature earns the signup
Every test page is assembled from feature blocks, and each block is tracked per audience: dwell time, expansions, and an explicit "this is the part I need" vote. Qualified signups are then asked to rank features outright.
Below the sample floor, the map says insufficient data instead of faking certainty. You will know what to build first before you write a line of code, and you will know what nobody voted for.
interest index = dwell +
expansions + explicit votes,
normalized per segment.
sample floor: n ≥ 25.
| feature block | solo agencies n=214 | etsy sellers n=96 | bookkeepers n=18 |
|---|---|---|---|
| bank feed auto-categorization | 0.91 | 0.44 | insufficient data |
| one-click month-end close | 0.78 | 0.31 | insufficient data |
| quarterly tax estimate | 0.52 | 0.83 | insufficient data |
| receipt inbox forwarding | 0.47 | 0.61 | insufficient data |
| client invoicing | 0.12 | 0.09 | insufficient data |
Two outcomes. Both are the product.
Either a named segment hits your pre-registered threshold and you get a validation report plus an MVP Blueprint with every line citing its evidence. Or every segment's ceiling falls below your floor, the run stops, explains why, and hands your remaining budget back for the next idea.
segment
solo agency owners, 2-5 seats, doing their own books
threshold, pre-registered d0, locked
30 qualified signups · verified email + survey match · ≤$15.00 blended CAC · ≥2 independent sources
result
31 qualified · $13.70 blended · reddit organic + meta ads [log d6-d9]
mvp blueprint excerpt
- build first: bank feed auto-categorization [ranked #1 by 71% of n=31]
- build second: one-click month-end close [interest 0.78, 2 booked calls cited it]
- skip: client invoicing [ranked last, 9% of votes]
- pricing start: $39/mo [median stated WTP $40, n=31]
- day one: launch email to the 78-person waitlist you just built [verified, double opt-in]
threshold, pre-registered d0, locked
25 qualified signups · ≤$12.00 blended CAC · ≥2 sources
why it died
- all 6 cells: projected ceiling below floor at 90% confidence [best cell: 9 qualified at $31.40]
- message failure: "save hours restocking" drew clicks, zero survey matches [n=112 visitors]
- offer failure: 0 of 14 verified signups accepted a call [escalation d5]
closeout
$212.60 unspent budget returned to pool · 14 signups thanked, data deletion offered, one click [log d7]
what you did not spend
roughly 3 months of building for nobody
How a trial runs
The order matters. The threshold is locked before traffic, free channels run before paid, and escalation comes before scale.
Pre-register the trial
Describe the idea, set a hard budget cap, and define validated before any traffic runs, for example 30 qualified signups with verified email plus a matching survey answer, at or under $15 blended acquisition cost, across at least two independent traffic sources. The threshold is locked so nobody can move the goalposts later.
The agent builds the courtroom
It researches the space, proposes ICP hypotheses as audience x message x offer cells, and builds an honest early-access page for each from feature blocks, deployed on your own domain with disclosure and double opt-in.
Traffic flows, evidence accumulates
Free channels first: community post drafts and SEO pages sit in your approvals queue until you say go. Then Meta and Google ads run from your own ad accounts at cost. The bandit shifts budget toward cells earning the most evidence per dollar, and when a segment shows promise the agent escalates the ask, survey, call, deposit, instead of just scaling spend.
Verdict with receipts
If a segment hits your threshold, you get a validation report plus an MVP Blueprint: what to build first ranked by measured feature interest, what to skip, a pricing starting point anchored to stated willingness to pay, and a day-one launch plan for the waitlist you just built, every line citing its evidence. If every segment's ceiling falls below your floor, the run stops, explains why, and hands your budget back. Killing an idea one-click thanks signups and offers to delete their data.
Cheaper than one month of building the wrong thing
The subscription covers the platform and the agent. Ad spend runs on your own Meta and Google accounts at cost, with no markup, ever.
$49 /mo
One idea on trial at a time. Everything you need to reach a verdict: pages, the bandit, the evidence ladder, and the report.
- 1 active product, 4 concurrent experiments
- Meta ads connector plus community draft-and-queue
- Feature-interest tracking per audience
- Validation report and MVP Blueprint
$99 /mo
For serial builders with a pipeline. Run three ideas in parallel and let the losers die fast.
- 3 active products, 8 concurrent experiments each
- Meta plus Google ads connectors
- Feature x segment heatmaps and 3x AI generation credits
- Learnings library carried across ideas
$249 /mo
For studios and teams validating a portfolio. Ten ideas on trial at once, all channels open.
- 10 active products, 8 concurrent experiments each
- All traffic channels
- 10x AI generation credits
- Priority support
Questions a skeptic should ask
Three ways. First, you pre-register what validated means before traffic runs, so you cannot rationalize weak results afterward. Second, conversions are weighted on an evidence ladder, a booked call or deposit counts 10-25x more than a raw email signup, so cheap junk traffic cannot fake a win. Third, a bandit tests multiple audience x message x offer cells at once and moves budget to what converts, instead of you betting everything on one page and one guess.
Because email lists lie. People sign up for anything free, so a raw signup counts 1x, a verified email 3x, a qualifying survey match 5x, and a booked call or deposit 10-25x. The optimizer maximizes evidence per dollar, and when a segment shows promise the agent escalates the ask rather than just scaling spend.
No. Hard daily and total budget caps are enforced in code, and the agent can reallocate within your budget but can never increase it. Anything outbound, community posts or ads, and anything spend-increasing waits in your approvals queue. A dead-man switch pauses all paid campaigns if signup tracking goes dark for 12 hours, and every decision is written to a log with its rationale.
Yours. Pages deploy to your own domain, and Meta and Google ads run from your own ad accounts at cost with no markup. The subscription covers the platform and the agent, and you keep the accounts, the pixels, and the waitlist.
Every test page carries an early-access disclosure and uses double opt-in, so nobody is misled about what stage the product is at. If you kill the idea, one click thanks your signups and offers to delete their data.
A documented kill, which is the second valuable outcome. When every segment's ceiling falls below your pre-registered floor, the run stops, explains exactly why each hypothesis failed, and hands your remaining budget back for the next idea. That is weeks of building you did not waste.
The cheapest thing you will ever learn is that you were wrong
The default path is to build for months on a hunch, launch to silence, and only then learn the ICP was wrong. The alternative is days, a budget cap you set, and a verdict either way: a validated segment with receipts, or a documented kill and your money back for the next idea.
14-day free trial · no card required · ad spend on your own accounts at cost, no markup