Mistakes we actually made building with AI — logged, debriefed, fixed.
Every one is something that happened to us, written down the day it
happened. Not theory, not a listicle, not advice from someone who read about it.
The honest ledger, stated once, up front: no stranger
has ever paid us anything. Everything here is a lesson about avoiding waste. None of
it is a lesson about how we made money, because we haven't.
Free · ungated · permanent · no signup, ever
Want the short version? AGENTS.md — the
rules pulled out of this log, written to paste straight into your own repo.
Machine-readable: corpus.json.
1. Check before you build. Almost every expensive mistake here was a build that a four-minute check would have prevented.
What not to do: Don't build first and check later. We built nine marketplace listings, three books and a long workbook interior before ever asking a marketplace how many people search for the thing.
3. A flattering result deserves more scrutiny than a damning one. Nobody investigates good news. That's how we believed we had a #3 placement when we actually had #64.
What not to do: Don't skip verifying good news. We spent zero seconds questioning a #3 placement and an hour questioning a #64 placement. The #3 was the lie.
4. The building was always the cheap part. Knowing what's true was expensive.
What not to do: Don't mistake output for progress. We produced 44 posts, 11 listings, 3 books and a website before producing a single verified fact about demand.
5. A thing that fails silently is worse than a thing that fails loudly. Our test suite refuses to pass an unspecified pattern; every platform tool we touched returned a plausible answer instead of an error.
What not to do: Don't accept a tool that can't fail. We used platform UIs as measurement instruments for two weeks without asking what a failure would look like.
6. Write the prediction down before the data arrives, or any outcome becomes explainable afterwards and nothing is learned.
What not to do: Don't explain a result after seeing it. We nearly read the three books' performance however the day's mood suggested — until we wrote predictions down.
7. Record decisions once, on paper. Two files that disagree are worse than either one alone.
What not to do: Don't leave a contradiction in two files. The feed was "folded to free" in one document and "$19, live" in another, on the same page, for hours.
8. The team's criticism of us was worth more than its ideas.
What not to do: Don't ask a panel what it thinks of your idea and then ignore what it says about your judgement. Their criticism of us was the useful half.
9. A lessons file only compounds if it's alive.(2026-08-19)
What not to do: treat a retrospective as a one-time deliverable — written once at the CEO's request, then left to rot while new mistakes pile up unrecorded, until someone repeats entry #84 verbatim (an agent did, six hours after it was written, because nothing made it read it).
How to do it right: standing order — every session appends its lessons the same day in the three-part format; the daily pipeline checks the day's work for uncaptured lessons; new agent briefs cite the entries relevant to their task. A document that grows daily is an operating manual AND a bottomless content inventory. One that doesn't is a eulogy.
10. Ask the platform what it knows before building for it. Etsy publishes marketplace search volume free. We ran a full set of listings for five days to learn what one query answered in a minute: ~25 searches/month.
What not to do: Don't infer demand from proxies when the platform publishes the real number. We used Google Trends shapes for a week; Etsy had the actual figure behind a free menu item.
11. A keyword is not a product. "Emergency binder" was our best term by volume, ratio and conversion — and on Amazon those words describe a physical binder with pockets we cannot manufacture.
What not to do: Don't build for a keyword without looking at what it returns. We picked "emergency binder" on volume alone. The results are physical binders.
17. Zero on one instrument is not zero demand. Bing showed zero for a category whose Amazon competitors hold 24,700 reviews. The instrument was blind, not the market empty.
What not to do: Don't declare a market dead on one instrument. Bing said zero for a category with 24,700-review incumbents.
19. Absence of competitors is a warning, not an opportunity. If a market exists and nobody serves it that way, ask why before assuming you're first.
What not to do: Don't treat an empty category as an opportunity. We nearly built big-button phone cards because nobody sold them. Nobody sells them because the free manual in the box wins.
20. Absence of a FORMAT in a proven market is a stronger warning still. Every top organiser is lay-flat; zero are perfect-bound. That's the answer.
What not to do: Don't ignore a missing FORMAT in a healthy category. Zero perfect-bound competitors was the answer, and we initially read it as an opening.
28. Marketplace search and web search are different populations. Someone typing the same term into Bing wants information; typing it into Etsy, they want to buy something.
What not to do: Don't glue a web-search volume to a marketplace standing. I did exactly that, told the CEO we'd "found it," and had to correct it an hour later.
32. Free research beats paid research when you haven't got a product yet. Bing Webmaster keyword volumes, Meta Ad Library (what competitors PAY to advertise), Apple's review RSS feed, TikTok Creative Center — all free.
What not to do: Don't reach for a paid tool before exhausting free ones. We nearly justified $29/month while four free tools sat unused.
33. Ahrefs' connector needs a paid plan behind the OAuth — we recorded it as a free "unlock" and were wrong for days.
What not to do: Don't record a connector as a free unlock without checking what sits behind the OAuth. We propagated that error into our own source of truth.
39. Some markets live in closed rooms (county Facebook groups, trade association meetups, feed-store noticeboards) that no dashboard and no outsider can read.
What not to do: Don't conclude "nobody's talking about it" from public searches alone.
42. A quality claim is not a feature claim. Our headline claim was that our math was verified correct — true, unfalsifiable to a shopper, worthless in a listing.
What not to do: Don't make your headline claim one a shopper can't check.
45. A free incumbent with network effects is fatal. The category leader we'd missed was free, tightly specialised to our exact buyer, and handled coordination between coworkers.
What not to do: Don't enter a niche defended by a free product with network effects. We got three days into an app plan before we found it.
46. Check what the category actually monetises. The apps in the category we were entering charged for EXPORT, not for the core feature. The value was in getting data out.
What not to do: Don't assume the obvious feature is what the category charges for. They gave away the part we planned to sell and charged for the part we'd have given away — we'd have built the free half.
50. KDP is the wrong manufacturer for anything people WRITE IN, and the right one for books people READ.
What not to do: Don't send a write-in workbook to KDP. The #1 organic result for the term we were chasing is that exact product, buried in the high six figures of its category's BSR.
51. The buyer is often not the user. Gifts, household purchases, anything bought for an aging parent.
What not to do: Don't write copy to the user when the buyer is someone else. Every line of our organiser copy addressed the person who'd never purchase it.
62. Every bet passes the three-question screen: who is the buyer, can we reach them AND sell to them, and is there a wedge that isn't just who we are?
What not to do: Don't approve a bet before running the three-question screen. Board 016 approved five products without it; the screen later killed several.
64. Distribution is the binding constraint, not product quality. Five routes blocked for five different reasons, all ending with "nobody comes."
What not to do: Don't keep fixing the product when the constraint is distribution. Five dead routes, five different causes, one shared symptom: nobody came.
65. A format is not a distribution plan.(2026-08-19, Board 019)
What not to do: research proved text was the right MEDIUM, then "post threads on X" slid in as if it were a plan — the same build-it-and-nobody-comes failure this business had already made five times, about to be repeated on a platform with no cold-start discovery pool at all.
How to do it right: name the mechanism that puts content in front of strangers before the first post. Our answer: a community platform as primary (communities ARE the discovery pool), replies into live conversations elsewhere (borrowed audiences), a long-form write-up venue monthly. Venue and mechanism, not just medium.
66. The best argument for a bet may not be in the proposal.(2026-08-19)
What not to do: pitch an idea on its surface appeal (trend, gap, funnel) without testing it against your own hardest screen.
How to do it right: the board carried the pivot 10-1 chiefly because it is the FIRST bet ever to pass question 3 of the three-question screen — the wedge is the operating log itself, which survives independent of who wrote it. Run your own screens on your own ideas before the board has to.
67. The agent that made the change never verifies the change.
What not to do: Don't let the builder check its own build. Four listings shipped with the wrong file attached because the same context reviewed its own work.
76. Give verifiers an environment brief — tab, auth state, expected wait times, retry instructions. Our first verifier reported a false failure by misreading a slow render as logged out.
What not to do: Don't send a verifier in blind. Our first one reported a false failure because nobody told it that back office renders slowly.
82. "Nothing came in" and "I could not look" are different findings, and a monitoring report that merges them is worse than no report.(2026-08-19, inbox sweep)
What not to do: one platform returned a TLS "Privacy error" on every URL during a routine sweep — a repeat of the same network interception seen the day before. The tempting write-up is a clean "all surfaces quiet," which would have recorded an unchecked surface as a verified-empty one. A comment or DM sitting there for days would then look like it arrived from nowhere.
How to do it right: every monitored surface gets one of three states in the report — checked, empty · checked, here is what came in · could not check, here is why. Never let the third collapse into the first. And do not click through a certificate warning to force the check: on an intercepting proxy that trades a gap in the report for a compromised session, which is the worse of the two.
83. A third-party list or aggregator is not verification. It is a claim about the account, and it dies on the same live check every other claim does. (2026-08-19, engagement round)
What not to do: a prospecting pass listed a channel at 7.1K subscribers on the strength of a listicle site's roundup, and it went into the follow queue as verified. The live channel had 2 videos, both 17 years old, with 32 and 25 views, and no subscriber count even rendered on the page. The listicle's figure did not describe this channel at all — a familiar failure: treating a search-result artifact as proof an account is what it claims.
How to do it right: third-party numbers are a lead, never a verification. Open the profile, read what is actually there, and let the live page overrule the citation every time. The check cost one page load and stopped a follow that would have made us look like a bot padding a list.
96. Name the standing constraints in every brief. Agents don't inherit your project's context file, so every rule you rely on has to be restated in the brief itself.
What not to do: Don't assume agents know the standing rules. We've had agents breach constraints written down in a file they were never given.
100. Weight the FINDINGS heavily and the VERDICTS lightly. Every genuine win was a bug, a defect, or a checkable fact. No vote ever did the work.
What not to do: Don't ship on a panel's vote. The unanimous "don't build the app" did no work; a free incumbent and an export-pricing finding did all of it.
102. Record kills by their finding, not their tally. "The niche is defended by a free network-effect incumbent" is reusable; "the panel voted no" teaches nothing.
What not to do: Don't record "the panel voted no." Record why.
104. Diverse personas break the frame — age, buyer versus user, skeptic versus enthusiast. A panel of one type produces one type of idea.
What not to do: Don't build a panel of one type. Five people from the same world produce five versions of the same idea, which is the exact trap we were in.
125. Format winners are per-platform, not universal. One platform rewarded confessions 10:1; another rewarded the numbers and buried the confession.
What not to do: Don't generalise one platform's format winner. We did it twice — one platform's result became doctrine, and the other platform's data was the opposite.
126. One platform's result became false doctrine twice before we caught it.
What not to do: Don't reuse a winning post across platforms without checking it hasn't already flopped there. Our "winner" was another platform's worst performer.
136. New accounts get flagged for machine-pace activity. One platform suspended an account hours after a rapid profile build; another flagged it the same week.
What not to do: Don't operate a days-old account at machine pace. One platform suspended us within hours of a rapid profile build.
141. Don't build everything AI-generated when platforms gate AI content. Platforms require disclosure of AI-generated media, and checked boxes can limit reach. Prefer content shapes that don't trigger the requirement at all — text, real screenshots, screen recordings of actual work.
What not to do: Don't default every post to AI-generated media and then wonder why reach is throttled by the disclosure checkbox you're forced to tick.
142. When something IS synthetic media, mark it every time. A throttled post costs a little; a banned account costs the channel. Compliance beats reach whenever they conflict.
What not to do: Don't skip the AI-content disclosure to protect reach. That trade is how accounts die — and we already lost days to two platform suspensions for lesser flags.
143. The honest format and the safe format can be the same format. Mistake content built on real screenshots and logs of things that actually happened is simultaneously the most credible shape (checkable claims) and the disclosure-safe shape (nothing synthetic to declare). When integrity and platform mechanics point the same direction, that's the direction.
What not to do: Don't reach for synthetic polish when real evidence exists. Our own screenshots of real failures are more credible AND safer than anything generated to look good.
144. Rescind rules that stop ideas; replace them with postures that guide judgment. "Never mention AI" blocked an entire venture. "Be mindful of what triggers disclosure, and mark what requires marking" enables the same venture safely. A ban is brittle; a posture travels.
What not to do: Don't write absolute bans where a judgment posture would serve. The never-mention-AI rule sat unexamined until it collided with the best idea we'd had — a rule that blocks thinking outlives its reason.
145. Don't let a correction harden into the doctrine it corrected.(2026-08-19, pipeline session's own flag on its own finding)
What not to do: "one platform rewards confessions, another rewards the numbers" was built from 11 posts with one outlier versus 12 posts topping at 562 — two data points. The finding existed to stop one platform's result becoming doctrine, and it was itself about to become doctrine — e.g. refusing to ever test a confession format on the platform that hadn't rewarded it yet.
How to do it right: label small-sample findings as priors to test, not rules to obey — in the record, next to the finding, by the person who made it. Log predictions ("this format is mechanism-shaped, so it should fit here too") and let results grade them.
146. Adapted-per-surface and cross-posted look identical on a task list and are not the same thing.(2026-08-19)
What not to do: fifteen uploads in one day, the identical asset on each platform, results diverging 7× and teaching us nothing per surface.
How to do it right: one lesson, four SHAPES — a community post, a thread, a short card, a site page — each built for its surface's mechanics. Every content brief states which shape, not just which lesson.
147. Before commenting, check whether we have already commented. The comment count includes our own, and two near-identical comments from one account under one video is the clearest bot signal there is. (2026-08-19, engagement round)
What not to do: I opened a video, read "Comments 1", took that 1 to be a stranger's, and posted a comment opening with the same personal claim we'd used the day before, on the same video, under our own prior comment. Same account, same video, same opening line, one day apart. I had carefully checked the like state on that video (the notes said it was already liked) and never thought to check the comment state.
How to do it right: on any item we may have touched before, check every interaction type, not the one the notes happen to mention. Deleted the duplicate and reload-verified back to 1 comment. The "no repeated phrasing in a session" rule needs a sibling: no repeated substance across sessions, on the same item.
148. A reply drafted in a previous session is a draft, not a decision. Re-read it against the live thread before sending. (2026-08-19, engagement round)
What not to do: our notes carried a fully cleared reply for an owed thread, and the obvious move was to paste it. Its middle sentence restated almost word for word something we ourselves had already posted 22 hours earlier, answering the same underlying point. Sending it would have answered their new point by repeating our old one.
How to do it right: read the whole thread first (already the rule), then check the draft against our own prior turns as well as theirs, and rewrite to answer what they actually added. Theirs had added new specifics we hadn't addressed yet, so the reply went to that instead.
IX. The plumbing (small but each one cost real time)
161. Test both branches of a fallback before trusting either.(2026-08-19, voice-render pipeline)
What not to do: a voiceover script preferred a real recording over text-to-speech — and the real-voice branch silently never fired, because pipefail made the file-detection pipeline fail whenever any candidate extension was missing. The TTS test passed, so the script looked done. One deliberate test of the other branch caught it in thirty seconds.
How to do it right: a script with an if/else has two behaviours; exercising one proves half the script. Feed it an input for each branch before calling it built — especially the branch you expect to use later rather than today.
162. A platform's limit usually attaches to the identifier, not to the account — so deleting the account doesn't reset it. (2026-08-19)
What not to do: blocked from creating a new email account because the phone number hit the provider's per-number cap, the instinct was to delete the existing brand account to free the slot. It would not have worked — the cap follows the phone number and decays on time — and it would have destroyed a linked video channel, a marketplace login and the recovery point for four brand accounts.
How to do it right: before deleting anything to "free up" a quota, ask what the quota is actually keyed to. And route around the gate instead: buying a domain first and turning on free email routing gave a working, phone-free brand address for $0 — better-looking on signups than a free webmail account, and it never needed the blocked resource at all.
163. Order operations by what unblocks what, not by what feels foundational.(2026-08-19)
What not to do: the registration runbook put "create the email" first because email feels like the foundation. Email was the one step that was blocked; the domain — which nobody had marked as a dependency — was what made a free email possible.
How to do it right: when a step blocks, check whether a later step in the list actually supplies what the blocked one needs. Re-sequence rather than wait.
173. He is right more often than the analysis when he says something feels wrong. His instinct that our target customers already had their own solution preceded the data proving it by a day.
What not to do: Don't argue with his instinct before checking it. His read that the target customers already had their own solution was right a full day before the data proved it.
178. Adversarial cross-checking between sessions caught what neither found alone — a wrong bug diagnosis, an ethics breach in a title, a bad thesis, and an over-claim.
What not to do: Don't run a single session on a decision that matters. Every serious error that day was caught by the other one.
184. The machine works. The audience doesn't exist yet. That is the whole remaining problem.
What not to do: Don't build another product to fix a distribution problem. That's the reflex this business reaches for every time, and it has failed five times in a row.