Rules of thumb learned from practice. These claims may not be provable in all cases, but they have been observed to work reliably. Heuristics evolve into rules as evidence accumulates.
Unresolvable errors: log and stop. No automatic retry.
Why: Retries on unresolvable errors create noise.
Failure mode: Agent retries 100 times, consumes rate limits.
Scope every future Wiki version against the AIBP Wiki 3.0 Feature-to-Milestone Map. Each requested feature must be mapped to a named milestone (or explicitly marked 'Future version, not yet scoped') before it enters delivery. Map: Weekly Case Studies Email / Monthly DeepResearch Newsletter / Advanced News Filtering / Vendor & Tool Refresh / Monthly Webinars -> Wiki 3.0 general content automation; Better Interactivity -> [3.0] UX Improvements and Deferred Features; Messageboard -> [3.0] Message Board Community; Dynamic ROI Graphs -> [3.0] ROI Quadrant Graphs; AI Chatbot -> [3.0] AI Chatbot; Graphic Org Chart -> [2.0] UI Standardization & Cleanup; User Organization Team Functionality -> Future version, not yet scoped; Research Tools for Sales Team -> [3.0] Lead Generation Analytics Dashboard; Audit Integration -> [3.0] Whitelabel Audit; Industry Landing Pages -> [3.0] Industry Landing Pages; AI Image Generation -> [3.0] AI Image Generation.
Why: A fixed feature-to-milestone map prevents scope creep and grooming churn, keeps requirements traceable to a named deliverable, and gives every future Wiki version a consistent scoping reference instead of re-litigating scope each time.
Failure mode: Feature requests for the AIBP Wiki were being scoped ad hoc without a consistent mapping to named milestones, causing scope ambiguity (e.g. UI Standardization ballooned 16 to 69+ tickets).
Scope every future Wiki version against the AIBP Wiki 3.0 Feature-to-Milestone Map. Each requested feature must be mapped to a named milestone (or explicitly marked 'Future version, not yet scoped') before it enters delivery. Map: Weekly Case Studies Email / Monthly DeepResearch Newsletter / Advanced News Filtering / Vendor & Tool Refresh / Monthly Webinars -> Wiki 3.0 general content automation; Better Interactivity -> [3.0] UX Improvements and Deferred Features; Messageboard -> [3.0] Message Board Community; Dynamic ROI Graphs -> [3.0] ROI Quadrant Graphs; AI Chatbot -> [3.0] AI Chatbot; Graphic Org Chart -> [2.0] UI Standardization & Cleanup; User Organization Team Functionality -> Future version, not yet scoped; Research Tools for Sales Team -> [3.0] Lead Generation Analytics Dashboard; Audit Integration -> [3.0] Whitelabel Audit; Industry Landing Pages -> [3.0] Industry Landing Pages; AI Image Generation -> [3.0] AI Image Generation.
Why: A fixed feature-to-milestone map prevents scope creep and grooming churn, keeps requirements traceable to a named deliverable, and gives every future Wiki version a consistent scoping reference instead of re-litigating scope each time.
Failure mode: Feature requests for the AIBP Wiki were being scoped ad hoc without a consistent mapping to named milestones, causing scope ambiguity (e.g. UI Standardization ballooned 16 to 69+ tickets).
Treat all figures beyond the current quarter as directional planning targets, not fixed commitments. Re-scope every forward KPI, pipeline, and ramp figure at each Quarterly Planning session against real actuals.
Why: The growth ramps (visitors, subscribers, pipeline, engagements, use cases) are forecasts built before actuals exist. Re-forecasting against real results each quarter keeps the plan honest and avoids managing to stale numbers.
Failure mode: Multi-quarter KPI and pipeline figures risked being treated as fixed commitments rather than planning targets, creating false precision and accountability against numbers that were only ever directional.
Scope every future Wiki version against the AIBP Wiki 3.0 Feature-to-Milestone Map. Each requested feature must be mapped to a named milestone (or explicitly marked 'Future version, not yet scoped') before it enters delivery. Map: Weekly Case Studies Email / Monthly DeepResearch Newsletter / Advanced News Filtering / Vendor & Tool Refresh / Monthly Webinars -> Wiki 3.0 general content automation; Better Interactivity -> [3.0] UX Improvements and Deferred Features; Messageboard -> [3.0] Message Board Community; Dynamic ROI Graphs -> [3.0] ROI Quadrant Graphs; AI Chatbot -> [3.0] AI Chatbot; Graphic Org Chart -> [2.0] UI Standardization & Cleanup; User Organization Team Functionality -> Future version, not yet scoped; Research Tools for Sales Team -> [3.0] Lead Generation Analytics Dashboard; Audit Integration -> [3.0] Whitelabel Audit; Industry Landing Pages -> [3.0] Industry Landing Pages; AI Image Generation -> [3.0] AI Image Generation.
Why: A fixed feature-to-milestone map prevents scope creep and grooming churn, keeps requirements traceable to a named deliverable, and gives every future Wiki version a consistent scoping reference instead of re-litigating scope each time.
Failure mode: Feature requests for the AIBP Wiki were being scoped ad hoc without a consistent mapping to named milestones, causing scope ambiguity (e.g. UI Standardization ballooned 16 to 69+ tickets).
Scope every future Wiki version against the AIBP Wiki 3.0 Feature-to-Milestone Map. Each requested feature must be mapped to a named milestone (or explicitly marked 'Future version, not yet scoped') before it enters delivery. Map: Weekly Case Studies Email / Monthly DeepResearch Newsletter / Advanced News Filtering / Vendor & Tool Refresh / Monthly Webinars -> Wiki 3.0 general content automation; Better Interactivity -> [3.0] UX Improvements and Deferred Features; Messageboard -> [3.0] Message Board Community; Dynamic ROI Graphs -> [3.0] ROI Quadrant Graphs; AI Chatbot -> [3.0] AI Chatbot; Graphic Org Chart -> [2.0] UI Standardization & Cleanup; User Organization Team Functionality -> Future version, not yet scoped; Research Tools for Sales Team -> [3.0] Lead Generation Analytics Dashboard; Audit Integration -> [3.0] Whitelabel Audit; Industry Landing Pages -> [3.0] Industry Landing Pages; AI Image Generation -> [3.0] AI Image Generation.
Why: A fixed feature-to-milestone map prevents scope creep and grooming churn, keeps requirements traceable to a named deliverable, and gives every future Wiki version a consistent scoping reference instead of re-litigating scope each time.
Failure mode: Feature requests for the AIBP Wiki were being scoped ad hoc without a consistent mapping to named milestones, causing scope ambiguity (e.g. UI Standardization ballooned 16 to 69+ tickets).
Treat all figures beyond the current quarter as directional planning targets, not fixed commitments. Re-scope every forward KPI, pipeline, and ramp figure at each Quarterly Planning session against real actuals.
Why: The growth ramps (visitors, subscribers, pipeline, engagements, use cases) are forecasts built before actuals exist. Re-forecasting against real results each quarter keeps the plan honest and avoids managing to stale numbers.
Failure mode: Multi-quarter KPI and pipeline figures risked being treated as fixed commitments rather than planning targets, creating false precision and accountability against numbers that were only ever directional.
GPT creative output undergoes a mandatory 2-hour hold before entering the review queue. No same-session generation and approval.
Why: Pattern: account managers reviewing GPT output immediately after generation approved it at a 94% rate. When we added a 2-hour hold (the AM reviews the batch later in the day or the next morning), the approval rate dropped to 71%. The 23% gap was copy that "sounded good in the moment" but had issues visible with fresh eyes -- subtle tone mismatches, claims that were technically true but misleading, and formatting that didn't match the client's brand voice.
Failure mode: Same-session review creates familiarity bias. The reviewer just saw the brief, the context is loaded, and the output feels like a natural continuation. Distance improves judgment.
Track and report the cross-model error rate separately from single-model error rates. Any task that involves both Claude and GPT is measured as a distinct category.
Why: Our overall agent error rate was 3.2%. When we segmented, single-model tasks (Claude-only or GPT-only) had a 1.8% error rate. Cross-model tasks had an 8.7% error rate -- nearly 5x. The errors were concentrated at handoff points: incomplete context transfer, schema mismatches, and misinterpretation of structured fields. Without segmenting, the 3.2% blended rate masked a systemic handoff problem.
Failure mode: Blended error rates hide that cross-model handoffs are the primary failure point. Resources are allocated to improving individual agents when the real problem is the integration layer.
Every agent must include its model name and version in the metadata of its shared state file output.
Why: When debugging an analysis discrepancy, we couldn't tell which model version produced a particular output. The ad monitor had been running on Claude 3 Opus while the pacing agent had been upgraded to Claude 3.5 Sonnet. Their outputs used different rounding conventions, making numbers mismatch by $1-3 per metric. Model version in metadata would have identified the discrepancy source in minutes instead of the 2 hours it took.
Failure mode: Without model version tracking, debugging cross-agent discrepancies requires testing each agent individually. Root cause identification is slow when version differences aren't visible.
Ad performance is evaluated at the location level with a minimum 14-day window. No optimization decisions are made on less than 14 days of data per location.
Why: Small-market locations (Tampa, Phoenix suburbs) have low daily lead volume. Day-to-day variance is enormous. A "bad day" is meaningless; a bad 2-week trend is actionable.
Failure mode: The ads agent paused a Google Ads campaign for the Tampa location after 5 days of zero leads. The campaign had averaged 2.1 leads/day over the prior month. The 5-day drought was within normal variance. Restarting the campaign lost 3 days of learning and reset the algorithm.
Lead quality scoring weights location-specific conversion history over network averages. A "hot" lead in Chicago (where average ticket is $149/mo) has different characteristics than a "hot" lead in Tampa (where average ticket is $89/mo).
Why: Network-wide lead scoring models produce false positives in lower-ticket markets and false negatives in higher-ticket markets.
Failure mode: The lead distribution agent prioritized Tampa leads using the Chicago-trained scoring model. It flagged members interested in premium personal training as "hot." Tampa doesn't offer premium PT. Staff called 30 "hot" leads pitching a service that didn't exist at their location.
Seasonal patterns are tracked per-location, not network-wide. January surges vary by 40-60% across locations. Summer dips range from 10% to 35% depending on climate and demographics.
Why: Using network-average seasonality for budget planning over- or under-allocates at the extremes.
Failure mode: The ads agent applied a uniform 30% January budget increase across all locations. Chicago needed 50% (cold weather drives indoor fitness demand). Tampa needed only 15% (year-round outdoor fitness options). Tampa overspent by $1,200 in January. Chicago underspent and missed 40+ leads.
When a client requests more than 3 rounds of revisions on a single deliverable, flag it for Marcus as a potential scope issue before logging revision 4.
Why: Unlimited revisions is the silent killer of creative agency margins. The flag forces a conversation about whether to charge for additional rounds.
Failure mode: One project went through 7 revision rounds without anyone noticing the pattern. The client was happy but the project margin was -12%. Marcus didn't realize until quarterly review.
Shot list suggestions must include at least 2 reference links from the client's existing brand content or stated references. Never suggest shots without grounding in the client's visual language.
Why: Generic shot suggestions feel like templates. Grounded suggestions feel like someone studied the brand.
Failure mode: Agent suggested a "drone reveal shot" for a brand that exclusively uses intimate, handheld footage. Designer flagged it as "clearly from a machine that doesn't understand the brand." Technically correct suggestion, totally wrong for the client.
Invoice reminders to clients use the same casual tone Marcus uses in his own emails. No formal language, no "please remit payment."
Why: Artifact's brand is approachable and creative. Formal invoice language feels like it's coming from a different company.
Failure mode: First automated invoice reminder used "Please find attached your outstanding invoice for services rendered." Client replied to Marcus: "Did you hire an accountant? Lol." Minor, but it broke the illusion of a small, personal shop.
Document assembly for a standard estate plan (trust, will, POA, advance directive) takes the agent 12 minutes. Priya's review takes 45 minutes per client package. Beth's formatting and preparation for signing takes 30 minutes. Total pipeline: 87 minutes per client versus the pre-agent baseline of 3.5 hours.
Why: Knowing the pipeline timing lets Priya schedule accurately. Before agents, she routinely underestimated document prep time and fell behind. Now she schedules 90-minute blocks per client and consistently hits the mark.
Failure mode: Without accurate pipeline timing, Priya overschedules. Four signing appointments in one day when she can only prepare for three. Last client's documents rushed. Error rate increases when Priya is behind schedule.
The scheduling agent sends three follow-up touches for annual trust reviews: 60 days before anniversary, 30 days, and 7 days. After the third touch with no response, it flags the client as DORMANT and stops. No more than 3 touches per year.
Why: Over-following-up annoys clients and feels desperate. Under-following-up loses annual review revenue ($750-$1,200 per review). Three touches at 60/30/7 days produced a 62% review booking rate, up from 38% when Beth was sending manual reminders on an ad hoc schedule.
Failure mode: Without a cap, the agent sends monthly reminders. Client perceives the firm as aggressive. Leaves a negative review mentioning "constant harassment." Priya loses a client and a referral source. In a solo practice, every client lost is felt in the revenue.
The assembly agent generates documents in the firm's standard formatting: 12pt Times New Roman, 1-inch margins, numbered paragraphs, firm letterhead. It never uses alternate fonts, creative layouts, or formatting that deviates from the template.
Why: Estate planning documents are read by courts, trustees, financial institutions, and opposing counsel. Non-standard formatting signals carelessness. A bank once questioned a trust because the formatting looked "different from what we usually see" and requested a letter of opinion confirming its validity. That letter cost Priya 2 hours.
Failure mode: Agent assembles a trust with a different font because the template's font metadata was corrupted. Bank receiving the trust as part of an account titling process flags it as potentially invalid. Client calls Priya. Priya spends 2 hours writing a letter of opinion. $600 in non-billable time because of a font.
Progress reports are generated biweekly, not weekly. Weekly reports create parent anxiety without providing meaningful new information.
Why: Tutoring progress is nonlinear. A bad week followed by a good week looks like a crisis and a recovery in weekly reports. Biweekly smooths the noise.
Failure mode: Weekly reports caused 3 parents to request "emergency conferences" in a single month because their child had one below-average session. Keisha spent 6 hours in unnecessary meetings. Moving to biweekly reduced parent escalations by 80%.
Tutor session notes must be submitted within 24 hours of the session. The scheduling agent flags missing notes at the 24-hour mark.
Why: Tutors forget details after 24 hours. Late notes are less accurate, which contaminates progress reports.
Failure mode: One tutor submitted 3 weeks of notes in a single batch. The notes were vague ("worked on math") and unusable for progress reports. Keisha had to contact 8 families to apologize for the delayed report. She now pays tutors a $5 bonus for same-day notes.
When a student misses 2 consecutive sessions without parent communication, the parent communication agent drafts a check-in message for Keisha.
Why: Missed sessions without communication often signal a family considering leaving. Early outreach retains 60% of at-risk families.
Failure mode: Before this rule, a family missed 4 sessions over a month. Keisha assumed they were on vacation. They had actually switched to a competitor. She found out when the mother mentioned it casually at a school event. $4,800/year lost with zero warning.
The release notes agent generates from merged PR titles and descriptions only. It does not read code diffs to infer what changed.
Why: The release notes agent read a diff that renamed an internal function and generated a changelog entry: "Breaking: API endpoint /auth/refresh renamed." The function rename was internal only -- the API contract hadn't changed. A user on the changelog RSS feed opened a GitHub issue asking about the "breaking change" and whether they needed to update their integration. The founder spent 30 minutes clarifying that nothing had changed for users.
Failure mode: Agents infer user-facing changes from internal code diffs. Internal refactors are misrepresented as breaking changes. Users react to non-existent breaking changes.
The support triage agent checks for duplicate issues before categorizing. Duplicates are linked, not re-triaged independently.
Why: The same bug was reported 4 times across GitHub issues and Slack over a weekend. The triage agent created 4 separate P1 entries. The founder's Monday morning check showed 4 P1 issues and he panicked, thinking 4 different critical bugs had emerged. It was 1 bug reported 4 ways. After implementing deduplication (matching by error message, stack trace similarity, and affected endpoint), false P1 volume dropped 35% over the next month.
Failure mode: Duplicate reports are triaged independently, inflating priority counts. The single reviewer overestimates severity based on volume. Panic replaces triage.
The code review agent includes a "confidence" indicator on each finding: HIGH (definite bug or security issue), MEDIUM (likely problem, needs human judgment), LOW (style preference, could go either way).
Why: Without confidence labels, the founder treated all code review findings equally. He either fixed everything (30 minutes on style nits) or skipped everything (missed real bugs). Confidence labels let him triage: fix all HIGHs immediately, review MEDIUMs during dedicated review time, batch LOWs for monthly style cleanup. Effective review time dropped from 25 minutes/day to 8 minutes/day with zero increase in bugs reaching production.
Failure mode: Uniform presentation of findings forces binary processing: everything or nothing. Confidence labels enable triage. Without them, the solo founder's limited review time is misallocated.
Proposal drafts from Archer should be 60-70% complete, not 95%. Leave strategic positioning, pricing rationale, and the executive summary for the consultant to write.
Why: Over-polished agent drafts create a false sense of completion. Consultants rubber-stamp instead of thinking critically. The best proposals have the consultant's genuine strategic voice in the sections that matter most.
Failure mode: Archer produced a near-perfect 28-page proposal. The partner skimmed it, approved it, and sent it. The pricing section included a 15% discount that Archer inferred from a prior similar engagement but that was not appropriate for this client's scope. Cost us $31,500 in margin.
Vault's knowledge base must be refreshed quarterly. Methodologies, case studies, and templates older than 18 months must be flagged for review and either updated or archived.
Why: Stale methodologies in proposals and deliverables make the firm look dated. Clients in fast-moving industries notice when frameworks reference pre-pandemic market conditions.
Failure mode: Archer pulled a "Digital Transformation Readiness Assessment" template from Vault that referenced 2023 technology benchmarks. The client's CTO pointed out the benchmarks were 3 years old during the proposal review call.
For engagements involving direct competitors, assign different consultant teams and ensure no agent holds context for both engagements simultaneously. Rotate agent context between engagements, never run them in parallel.
Why: Even with firewalls, simultaneous context is the highest-risk vector for information leakage. Sequential processing with context clearing is safer than parallel processing with access controls.
Failure mode: This is the architectural response to the Haldane/Orion incident (C001). No second incident has occurred since implementing sequential processing.
Lead nurture sequences pause automatically when a prospect books a trial class. Resume only if they no-show or don't convert within 7 days.
Why: Continuing to nurture someone who already booked feels tone-deaf and spammy.
Failure mode: A prospect booked a trial, received 3 more "Book your free class!" emails before attending. Replied "I already did, is anyone actually reading these?" Trial converted but trust was damaged from the start.
Trainer performance metrics use a 4-week rolling average, not single-week snapshots. Seasonal patterns (New Year surge, summer dip) are normalized against the same period last year.
Why: Single-week data is noisy. A trainer with one bad week due to illness shouldn't be flagged. Seasonal patterns create false positives.
Failure mode: January metrics showed every trainer "improving" dramatically. It was just the New Year resolution surge. Jamie almost gave bonuses based on phantom performance gains.
Location-specific context must be attached to every agent action. No agent operates in a "generic CoreFit" mode. Each location has different peak hours, demographics, and class preferences.
Why: The downtown location skews young professionals (25-35). The suburban location skews parents (35-50). Messaging that works for one alienates the other.
Failure mode: Lead nurture sent "Bring the kids to our Saturday Family Fitness!" to downtown prospects. Downtown has no kids' classes. 4 confused replies.
Member retention risk scoring uses a weighted model: visit frequency (40%), class variety (20%), social engagement (15%), billing consistency (15%), tenure (10%). No single factor triggers an at-risk flag alone.
Why: Single-factor triggers produce too many false positives. A long-tenured member who drops visit frequency for 2 weeks might be on vacation, not churning.
Failure mode: Early model used visit frequency alone. Flagged 23 members as at-risk in one week. 19 were on a local school spring break vacation. Jamie wasted 4 hours reviewing false flags.
Issue triage prioritizes by: (1) enterprise customer reports, (2) security issues, (3) issues with reproduction steps, (4) feature requests with 5+ thumbs-up, (5) everything else.
Why: Enterprise customers pay. Security issues are existential. Reproducible issues get fixed faster. Community-validated features should ship. Everything else can wait.
Failure mode: Before prioritization, issues were triaged by recency. An enterprise customer's critical bug sat at position #14 in the queue behind 13 minor feature requests. The customer escalated via email after 5 days. Kai fixed it in 30 minutes but the delayed response nearly cost the $2,400/year contract.
Release notes are published within 24 hours of a release. If Kai hasn't reviewed the draft within 12 hours, the agent sends a reminder to his private Slack channel.
Why: Enterprise customers monitor releases. A release without notes triggers "what changed?" emails that cost more time than writing the notes.
Failure mode: Kai shipped v2.6.0 on a Friday and forgot to publish release notes. By Monday, 4 enterprise customers had emailed asking what changed. One customer's security team flagged the update as "unreviewed" and blocked their team from upgrading. It took 2 weeks to get through their security review after the late notes were published.
Discord monitoring tracks sentiment, not just questions. A shift from positive to negative sentiment in any channel triggers a summary to Kai's private Slack within 1 hour.
Why: Developer communities turn fast. A frustrating bug or a perceived lack of responsiveness can shift tone from supportive to hostile in a single day.
Failure mode: A breaking change in v2.3.0 caused issues for 15+ users over a weekend. Discord #help went from 2 messages/day to 30 messages/day, all negative. Kai was offline and didn't see it until Monday. By then, the narrative had solidified: "DevForge ships breaking changes without warning." A community member forked the project as a "stable alternative." The fork got 200 stars before Kai could respond.
Support ticket priority is determined by financial impact. Tickets mentioning failed payments, incorrect balances, or unauthorized transactions are auto-escalated to P1 regardless of the user's tone or language.
Why: A polite user reporting a $500 balance discrepancy is more urgent than an angry user complaining about the UI. Financial impact trumps sentiment.
Failure mode: The triage agent initially used sentiment analysis for priority. An angry user complaining about a font change was prioritized over a calm user reporting that $1,200 appeared to be missing from their account. The calm user waited 18 hours for a response. The "missing" money was a categorization display bug, but the delay eroded trust.
API health checks run every 30 seconds for Plaid and Stripe. Degraded performance (response time >2x baseline) triggers a P2 alert. Full outage (no response for 3 consecutive checks) triggers P1.
Why: Degraded performance often precedes full outages. Early warning gives the engineering team time to activate failover or notify users before the situation becomes critical.
Failure mode: Before the degradation detection, a Stripe partial outage caused payment processing to slow from 200ms to 8 seconds. Users experienced "spinning" payment screens. 14 users abandoned mid-payment. The monitoring agent only alerted when Stripe went fully down 40 minutes later.
The categorization QA agent flags systematic drift when any category's miscategorization rate exceeds 5% over a 7-day window. Drift below 5% is logged but not alerted.
Why: Individual miscategorizations are normal (merchants change names, new merchants appear). Systematic drift indicates a model problem that affects many users simultaneously.
Failure mode: A merchant data provider changed their taxonomy, causing "Groceries" to be classified as "General Merchandise" for 380 users. The drift wasn't flagged for 12 days because the old threshold was 10%. Users noticed before Greenline did. 7 support tickets in one day about "wrong categories."
Demand letter drafts use the firm's established template structure and tone. The agent does not experiment with novel legal arguments or creative formatting. If the fact pattern suggests a non-standard approach, the agent flags it and defers to the attorney.
Why: Insurance adjusters read thousands of demand letters. They recognize standard, professionally structured letters as coming from competent firms. A creatively formatted letter or an experimental legal argument can signal inexperience, even if the content is sound.
Failure mode: Agent drafts a demand letter using a narrative style instead of the firm's established format. Adjuster perceives the firm as inexperienced. Initial counteroffer is 40% lower than expected. Attorney spends two additional months negotiating to reach the same number a standard letter would have achieved.
Client update calls are scheduled between 10 AM and 3 PM on Tuesday through Thursday. Mondays are for internal case review. Fridays are for court appearances and depositions. The comms agent never schedules outside this window without attorney override.
Why: Attorneys need uninterrupted blocks for case preparation and court appearances. Early implementation allowed the comms agent to schedule calls at 8 AM and 4:30 PM. Attorneys were arriving at court unprepared because they had been on a client call until 15 minutes before their appearance.
Failure mode: Comms agent schedules a client call at 8:15 AM. Attorney's deposition starts at 9:00 AM. Call runs long. Attorney arrives at deposition flustered and underprepared. Opposing counsel notices and pushes harder on key points.
The intake agent asks 7 standardized screening questions before classification. If a potential client cannot answer 3 or more questions, the case is classified as INCOMPLETE rather than WEAK. Incomplete cases get a 48-hour follow-up, not a decline.
Why: Trauma patients often cannot recall details during the first call. A car accident victim called 2 days post-accident and could not provide the other driver's insurance info, the police report number, or the exact location. Agent classified the case as WEAK. Attorney overrode it. Case settled for $195K.
Failure mode: Traumatized potential client provides incomplete information. Agent classifies as WEAK or DECLINE. Firm turns away a strong case because the agent prioritized data completeness over human context.
Listing descriptions are between 150 and 300 words. Under 150 feels thin and suggests the agent did not visit the property. Over 300 gets truncated on Zillow's mobile display. Every description includes: location context (without superlatives), key features (square footage, bedrooms, bathrooms, lot size), notable upgrades, and a neutral call to action.
Why: Analysis of 200 Denver MLS listings showed that descriptions between 150-300 words received 23% more saves than those outside this range. Zillow's mobile truncation at approximately 280 characters for the preview means the first two sentences must contain the most important information.
Failure mode: Description runs to 450 words. Zillow mobile preview shows only the first two sentences, which happen to be generic neighborhood context. Buyer scrolling on their phone never sees the renovated kitchen or the mountain views. Listing gets fewer saves and fewer showing requests.
The qualifier responds to new web leads within 15 minutes during business hours (8 AM-7 PM) and within 2 hours outside business hours. Response includes a personalized acknowledgment referencing the specific property or search criteria the lead expressed interest in.
Why: NAR research shows that leads contacted within 5 minutes are 21x more likely to convert than those contacted after 30 minutes. Our 15-minute target balances speed with personalization quality. Before automation, average response time was 4.7 hours. After: 11 minutes during business hours.
Failure mode: Lead submits interest on a Zillow listing at 10 AM. Without automated qualification, the inquiry sits in an agent's email until she checks between appointments at 2 PM. By then, the lead has received responses from three other brokerages. Lead goes with the fastest responder.
Seller reports are generated every Thursday at 4 PM and delivered by 6 PM. This timing allows sellers to review over the weekend and come to Monday meetings with questions. Reports include week-over-week showing trends, not just raw numbers.
Why: Sellers who receive reports on Monday morning feel blindsided going into the work week. Friday delivery felt end-of-week. Thursday gives sellers 72 hours to process the information before their next conversation with their agent. Seller satisfaction scores improved from 7.2 to 8.4 (out of 10) after the timing change.
Failure mode: Report delivered Monday morning shows a 40% drop in showings. Seller panics, calls agent before the agent has had coffee. Reactive conversation instead of strategic one. Agent spends 45 minutes calming the seller instead of 15 minutes discussing next steps.
Recommendation-first behavior is preferred over question-first behavior when enough context exists to make a strong next-step proposal.
Why: The organization is explicitly designed to reduce user bottleneck load and decision fatigue.
Failure mode: If agents default to interrogation instead of recommendation, the principal becomes the routing and reasoning layer, defeating the purpose of the system.
Agents should ask at most one focused clarifying question when a missing detail blocks action, rather than opening broad discovery loops.
Why: The system values execution momentum and low-friction interaction.
Failure mode: Over-questioning slows progress, increases user effort, and creates the sense that AI is adding coordination overhead rather than removing it.
No-show prediction models must use only day-of-week, time-of-day, appointment type, weather, and visit number in sequence (first visit, second visit, etc.). Patient demographics, diagnosis, and insurance type must not be used as predictive features even if they improve model accuracy.
Why: Using diagnosis or insurance type to predict no-shows creates discriminatory scheduling practices. If the model learns that Medicaid patients no-show more frequently and double-books those slots, it's implementing economic discrimination in healthcare access. This violates both ethical standards and potentially the Civil Rights Act.
Failure mode: Hypothetical, but the constraint was added proactively after a published case study from another practice showed that using insurance type as a no-show predictor resulted in Medicaid patients being systematically double-booked, reducing their available appointment times. Kinwell's practice manager read the case study and preemptively restricted Flow's feature set.
Beacon must include accurate clinical information in all patient education content. Every health claim must be sourced from peer-reviewed literature or professional PT association guidelines. Beacon must not generate exercise recommendations, recovery timelines, or treatment expectations without clinical review.
Why: Healthcare marketing content that includes inaccurate clinical information is a liability risk. A blog post that says "most ACL recoveries take 8-12 weeks" when the actual clinical range is 6-9 months creates patient expectations that the practice cannot meet. It also exposes the practice to malpractice claims if a patient cites the content as the basis for their treatment expectations.
Failure mode: Beacon drafted a blog post stating "heel spurs typically resolve within 4-6 weeks of physical therapy." The actual clinical consensus is that plantar fasciitis (the condition causing heel spurs) typically requires 6-12 months of conservative treatment including PT. The clinical reviewer caught it. Had the post been published, patients beginning PT for heel spurs would have expected resolution in 4-6 weeks and been dissatisfied when it took longer.
Maintenance requests received between 10 PM and 6 AM are held for triage until 6 AM unless the tenant explicitly states emergency language (flooding, gas smell, fire, no heat, sparking). Non-emergency overnight requests are acknowledged immediately ("received, will be reviewed at 8 AM") but not triaged or dispatched until morning.
Why: 68% of after-hours requests in our first 6 weeks were ROUTINE. Triaging them overnight triggered unnecessary overnight vendor dispatch planning. The hold-until-6-AM rule reduced after-hours vendor contacts by 71% without any increase in property damage from delayed responses.
Failure mode: Without the overnight hold, every after-hours request triggers full triage. 68% are routine but still generate 2 AM Slack notifications to Corinne. Corinne sleeps poorly. Decision quality degrades during the day. She misses a rent payment pattern that would have flagged a tenant at risk of default.
Tenant communications use a warm but professional tone. No exclamation points. No emojis. No slang. No first-name-only greetings. Every message includes a ticket reference number and Corinne's direct phone number for urgent follow-up.
Why: Tenants pay $1,350/month. They expect professional management, not casual text messages. Early comms agent output included "Hey Marcus! Got your request -- we're on it!" Marcus was a 62-year-old retired teacher who found the tone disrespectful. He called Mark directly to complain. Mark agreed.
Failure mode: Casual tone alienates older or more formal tenants. A 62-year-old tenant paying $16,200/year in rent does not want to receive a text that reads like it came from a college intern. Tone mismatch erodes trust. Tenant does not renew. Lost lifetime value over 4 years: $64,800.
The vendor agent tracks response times for all dispatched work orders. Vendors who exceed their committed response time by more than 4 hours on 3 or more occasions are flagged for review. The flag goes to Corinne, not the vendor. Vendor relationship management is human-only.
Why: Our preferred HVAC vendor was consistently 6-8 hours late on non-emergency calls. The agent flagged the pattern after 5 late responses. Corinne renegotiated the response time SLA and secured a $50 discount per late response. Without the tracking, the pattern would have gone unnoticed because each individual delay seemed reasonable.
Failure mode: Without vendor performance tracking, patterns hide in individual incidents. Each late response seems like a one-off. Over a year, the same vendor is late 15 times. Tenants associate slow repairs with poor management. Satisfaction drops. Renewal rates drop. Mark never knows the root cause is one vendor.
Support response time target: 4 hours during semester, 24 hours during breaks. The triage agent auto-escalates any ticket older than 3 hours during semester to the #support-urgent Slack channel.
Why: Students study on deadlines. A support ticket filed at 10 PM before an exam needs a response before midnight, not the next morning.
Failure mode: A student filed a ticket at 11 PM about not being able to access a study guide. The exam was at 8 AM. Support responded at 9 AM. The student had already failed to study the material. She left a 1-star app store review: "Platform broke the night before my exam and nobody helped." The review stayed up for 6 months and was cited by 2 prospective teachers who decided not to adopt.
Content QA prioritizes study guides that align with upcoming exam dates. Guides for exams within 2 weeks get priority review.
Why: A factual error in a guide nobody's using is low risk. A factual error in a guide that 500 students will use for tomorrow's exam is catastrophic.
Failure mode: Before priority-based QA, all content was reviewed in creation order. A new chemistry guide (exam in 3 days, 280 students) sat behind 15 older guides in the QA queue. It contained an incorrect molecular weight. Caught 6 hours before the exam by a student who filed a support ticket.
Stripe billing alerts (failed payments, subscription cancellations) are routed to Priya within 1 hour. The support agent drafts a personal "we miss you" email for cancellations, but only sends if Priya approves.
Why: Most cancellations are recoverable within 48 hours. After 48 hours, the student has found an alternative and the recovery rate drops from 35% to 8%.
Failure mode: 12 students cancelled during a billing system migration. The cancellation alerts were batched and delivered 3 days later. By then, 10 of 12 had switched to Quizlet. Recovery emails were ignored. $1,440/year in lost revenue.
New agents run in shadow mode for 14 days before their output reaches anyone outside the team.
Why: Our auto-send experiment on day 3 of deploying the client update agent. It sent a "your CPL increased 34% this week" email to 12 clients at 6:47 AM on a Saturday. Three clients called within the hour. The CPL increase was real but within normal weekly variance and the email lacked context about seasonality. We turned off auto-send and haven't re-enabled it.
Failure mode: Agents send raw metrics without context to clients. Clients panic. Founder spends Saturday morning on damage control calls.
Stale data flags must include hours-since-update, not just "stale."
Why: "Stale" means nothing. Is it 2 hours old or 2 days old? The briefing said "Dash data is stale" for our ad monitoring output. The founder assumed it was a few hours old and made decisions accordingly. It was actually 3 days old because the Meta API token had expired Friday evening and nobody noticed until Tuesday morning.
Failure mode: Ambiguous staleness labels lead to decisions based on data that's far older than assumed.
Weekly agent review must include a false positive rate for each alerting agent. Target: below 15%.
Why: After the 47-alerts-in-one-day incident (see C002), we started tracking false positive rates. Our Google Ads monitor was at 62% false positives -- nearly two-thirds of its alerts required no action. We recalibrated thresholds and got it to 11% over the next two weeks. The Meta monitor was already at 8%. Without tracking the rate, we wouldn't have known which agent needed calibration.
Failure mode: Without measurement, alert quality degrades silently. Teams compensate by ignoring alerts rather than fixing thresholds.
Feasibility spikes at Step 4 resolve technology uncertainty before committing. AI generates a small, throwaway proof-of-concept for each uncertain technology choice. If the spike fails, the technology is eliminated.
Why: Technology selection based on documentation and AI recommendation alone is unreliable. A 2-hour spike reveals integration problems that no amount of research can predict.
Failure mode: Team selects a database technology based on AI analysis of documentation. The technology has an undocumented limitation that blocks a core use case. Discovered at Step 10. Migration required.
Implementation proceeds one stakeholder slice at a time (Step 10). Each slice is reviewed by the human practitioner before the next slice begins.
Why: Large-batch implementation accumulates errors that compound. Per-slice review catches integration issues, requirement misunderstandings, and scope drift before they propagate.
Failure mode: Team implements three stakeholder slices without intermediate review. The first slice has a data model error. Slices two and three build on the error. All three require rework.
The scientific method applies to business purpose validation. Hypotheses are stated, predictions are made, tests are designed, and results are evaluated against pass/fail criteria. AI generates the test infrastructure. The human defines the hypotheses and evaluates the results.
Why: Without explicit pass/fail criteria, validation becomes subjective. "It seems to work" is not validation. "These three metrics exceeded these three thresholds" is validation.
Failure mode: Team "validates" by showing the prototype to stakeholders and asking "Does this look right?" Stakeholders approve politely. The system fails in production because polite approval is not the same as validated business purpose.
The banned phrases list must be reviewed monthly. Add phrases that appear in client feedback as "generic," "consultant-speak," or "AI-sounding." Remove phrases that have been successfully avoided for 3+ months (they're internalized).
Why: Language drift is continuous. New cliches emerge. Old ones fade. A static banned list becomes irrelevant over time. The list is a living document that reflects current failure patterns.
Failure mode: The banned list went 4 months without update. During that period, Forge started using "unlock value" and "drive impact" heavily -- phrases not on the original list. A client's feedback form noted "the deliverable felt AI-generated." The founder added 8 new phrases to the banned list.
For new client engagements, Scout must produce a "day zero" research brief within 4 hours of the signed SOW. This brief establishes the baseline: industry context, competitive landscape, key players, and known risks. Forge and Prep both read this brief before producing their first outputs.
Why: The first 48 hours of a new engagement set the tone. If the founder walks into the kickoff meeting without solid research, the client questions whether they made the right choice. The day-zero brief ensures every agent starts with shared context.
Failure mode: New engagement kicked off without a day-zero brief. Prep created meeting talking points based on the SOW alone (no industry context). The founder asked a question in the kickoff that revealed unfamiliarity with a major regulatory change in the client's industry. The client's General Counsel raised an eyebrow. It took 3 meetings to rebuild confidence.
If founder has fewer than 3 OTP hours in a week, defer all non-build work.
Why: Low-availability weeks must protect build above everything.
Failure mode: Low-availability week spent on outreach delays timeline by 2 weeks.
Proposal pricing is calculated using a blended rate model: Mara's time at $175/hr, senior designer time at $150/hr, junior designer time at $90/hr, with a 15% agency margin. The proposal agent calculates project pricing using this model and presents Mara with a price range (low/expected/high based on scope uncertainty).
Why: Mara was chronically underpricing projects because she estimated from memory. The agent's pricing model ensures every proposal covers costs and maintains margin.
Failure mode: Before the pricing model, Mara quoted a $12K brand identity project based on "feel." Actual cost (tracked post-project): $15,800 in labor. The $3,800 loss on a small agency's margins was felt for 2 months.
Client feedback synthesis groups feedback into three categories: (1) Factual corrections ("The phone number is wrong"), (2) Preference statements ("I prefer the blue version"), (3) Strategic concerns ("This doesn't feel premium enough for our audience"). The creative team receives all three but is expected to address #1 immediately, consider #2, and discuss #3 with Mara before acting.
Why: Not all client feedback carries the same weight. Factual corrections are objective. Preferences are subjective. Strategic concerns may require a creative rationale, not a revision.
Failure mode: Without categorization, the creative team treated all feedback equally. A client's casual "I kind of like the blue better" (preference) was treated the same as "This doesn't match our brand positioning" (strategic concern). The team changed the color without discussing the strategic point. The client was happy with the color but dissatisfied with the positioning. Two additional revision rounds followed.
The competitive visual analysis agent uses GPT-4V to analyze visual trends in competitor work. Analysis covers layout patterns, typography trends, color usage, and design language. The agent produces structured reports, not creative direction.
Why: Visual analysis requires a model capable of interpreting images. Claude handles text; GPT-4V handles visual interpretation. The two platforms serve complementary roles.
Failure mode: An early attempt to describe competitor visuals using text-only (Claude) produced vague descriptions ("clean, modern aesthetic with blue tones"). GPT-4V analysis was specific: "Competitor uses a 12-column grid with 60/30/10 color ratio, Helvetica Neue at 3 type scales, 24px base unit." The specificity made the analysis actually useful to the design team.
Memory should be treated as a first-class operating asset, with separate components for event logging, consolidation, and retrieval.
Why: The system includes Scribe for event logging, Archivist for summary consolidation, Seeder for bootstrap summary creation, and CustomerOps memory tools for event logs, summaries, refreshes, and rebuilds.
Failure mode: Without staged memory management, later agents reprocess too much raw data, lose continuity across interactions, and make decisions on stale or fragmented context.
When a contact already has usable memory, the org prefers lightweight contextual refresh over full recomputation.
Why: Lens explicitly uses different behavior on memory hit vs. memory miss, and the broader memory architecture supports incremental consolidation rather than always rebuilding from scratch.
Failure mode: Always recomputing full context increases token cost, slows response time, and creates more opportunities for inconsistency between runs.
The org prefers narrow, structured outputs over open-ended prose for machine-to-machine handoffs.
Why: Agent descriptions repeatedly reference structured summaries, schema versions, typed fields, validator outputs, and table/output writers rather than free-text-only communication.
Failure mode: Unstructured handoffs increase ambiguity between steps, raise parsing risk, and make validators and downstream tools less effective.
For local service businesses, the agent must check search terms for geographic intent mismatches weekly, not just cost and conversion metrics.
Why: An HVAC client's campaigns looked great on paper: CPL $31, 38 leads/month. But 12 of those leads searched for "AC repair [neighboring city]" and were served ads because the radius targeting overlapped into the next town. The HVAC company doesn't service that area. They were paying $31 per useless lead for a third of their volume. The geographic search term check caught what the performance metrics missed.
Failure mode: Performance metrics look healthy while geographic targeting silently wastes budget. Local businesses serve defined areas that don't always align with radius targeting.
Client reports must include a "leads by city" breakdown for any local service business.
Why: After the plumber incident (C001) and the HVAC issue (C007), we added city-level lead breakdowns to every report. Two more geographic problems were caught in the first month: a family law firm getting leads from a state where they're not licensed, and a dentist attracting patients from 45 minutes away who never convert because the drive is too far.
Failure mode: Aggregate lead counts mask geographic distribution problems. Clients don't know their ad dollars are leaking into wrong territories until they see the breakdown.
Conversion tracking status checks run daily and flag any account where conversion actions have recorded zero conversions in 48+ hours (for accounts that typically convert daily).
Why: A dentist client's Google Tag stopped firing after a WordPress plugin update. The agent saw "zero leads today" and reported it as low performance. It took 4 days for the media buyer to realize it was a tracking issue, not a performance issue. During those 4 days, the actual leads were coming in (the phone was ringing) but nothing was being attributed. Bidding algorithms degraded because they thought nothing was converting.
Failure mode: Tracking failure is misdiagnosed as performance decline. Bidding algorithms lose signal. Agents report "bad performance" when the actual problem is measurement.
Reports generated before 7 AM use yesterday's final numbers, not partial today numbers. Never mix time windows in a single report.
Why: Partial-day data creates misleading trends. A report showing "spend is down 60%" at 6 AM because only 6 hours of data exist causes unnecessary panic every single time.
Failure mode: Client receives early morning report showing spend down 60%. Calls account manager in alarm. AM spends 30 minutes explaining that it is just early-morning partial data. Happens three times before we fix the rule.
Reports generated before 7 AM use yesterday's final numbers, not partial today numbers. Never mix time windows in a single report.
Why: Partial-day data creates misleading trends. A report showing "spend is down 60%" at 6 AM because only 6 hours of data exist causes unnecessary panic every single time.
Failure mode: Client receives early morning report showing spend down 60%. Calls account manager in alarm. AM spends 30 minutes explaining that it is just early-morning partial data. Happens three times before we fix the rule.
Reports generated before 7 AM use yesterday's final numbers, not partial today numbers. Never mix time windows in a single report.
Why: Partial-day data creates misleading trends. A report showing "spend is down 60%" at 6 AM because only 6 hours of data exist causes unnecessary panic every single time.
Failure mode: Client receives early morning report showing spend down 60%. Calls account manager in alarm. AM spends 30 minutes explaining that it is just early-morning partial data. Happens three times before we fix the rule.
Reports generated before 7 AM use yesterday's final numbers, not partial today numbers. Never mix time windows in a single report.
Why: Partial-day data creates misleading trends. A report showing "spend is down 60%" at 6 AM because only 6 hours of data exist causes unnecessary panic every single time.
Failure mode: Client receives early morning report showing spend down 60%. Calls account manager in alarm. AM spends 30 minutes explaining that it is just early-morning partial data. Happens three times before we fix the rule.
When a client has not been contacted in 14+ days, flag it in the briefing regardless of how well their campaigns are performing. Silence is a churn signal even when the numbers are good.
Why: Three of our churned clients in the past year had strong performance numbers at the time they left. They did not leave because of results. They left because they felt ignored and undervalued.
Failure mode: Client campaigns perform well for 6 straight weeks. No proactive outreach from the team. Client quietly signs with a competitor who calls them every week.
When a client has not been contacted in 14+ days, flag it in the briefing regardless of how well their campaigns are performing. Silence is a churn signal even when the numbers are good.
Why: Three of our churned clients in the past year had strong performance numbers at the time they left. They did not leave because of results. They left because they felt ignored and undervalued.
Failure mode: Client campaigns perform well for 6 straight weeks. No proactive outreach from the team. Client quietly signs with a competitor who calls them every week.
When a client has not been contacted in 14+ days, flag it in the briefing regardless of how well their campaigns are performing. Silence is a churn signal even when the numbers are good.
Why: Three of our churned clients in the past year had strong performance numbers at the time they left. They did not leave because of results. They left because they felt ignored and undervalued.
Failure mode: Client campaigns perform well for 6 straight weeks. No proactive outreach from the team. Client quietly signs with a competitor who calls them every week.
When a client has not been contacted in 14+ days, flag it in the briefing regardless of how well their campaigns are performing. Silence is a churn signal even when the numbers are good.
Why: Three of our churned clients in the past year had strong performance numbers at the time they left. They did not leave because of results. They left because they felt ignored and undervalued.
Failure mode: Client campaigns perform well for 6 straight weeks. No proactive outreach from the team. Client quietly signs with a competitor who calls them every week.
New agents start in shadow mode for 2 weeks minimum. They generate output that a human reviews but the team does not act on. After 2 weeks of consistently accurate output, they graduate to draft mode where output is used after human review.
Why: We deployed the prospecting agent directly into production without a shadow period. Its first batch of outreach emails included a company that was a current client's direct competitor. Two weeks of shadow mode would have caught that conflict on day 4.
Failure mode: New agent sends outreach to a prospect that has a direct conflict with an existing client relationship. Client hears about it through industry contacts. Trust damaged.
New agents start in shadow mode for 2 weeks minimum. They generate output that a human reviews but the team does not act on. After 2 weeks of consistently accurate output, they graduate to draft mode where output is used after human review.
Why: We deployed the prospecting agent directly into production without a shadow period. Its first batch of outreach emails included a company that was a current client's direct competitor. Two weeks of shadow mode would have caught that conflict on day 4.
Failure mode: New agent sends outreach to a prospect that has a direct conflict with an existing client relationship. Client hears about it through industry contacts. Trust damaged.
New agents start in shadow mode for 2 weeks minimum. They generate output that a human reviews but the team does not act on. After 2 weeks of consistently accurate output, they graduate to draft mode where output is used after human review.
Why: We deployed the prospecting agent directly into production without a shadow period. Its first batch of outreach emails included a company that was a current client's direct competitor. Two weeks of shadow mode would have caught that conflict on day 4.
Failure mode: New agent sends outreach to a prospect that has a direct conflict with an existing client relationship. Client hears about it through industry contacts. Trust damaged.
New agents start in shadow mode for 2 weeks minimum. They generate output that a human reviews but the team does not act on. After 2 weeks of consistently accurate output, they graduate to draft mode where output is used after human review.
Why: We deployed the prospecting agent directly into production without a shadow period. Its first batch of outreach emails included a company that was a current client's direct competitor. Two weeks of shadow mode would have caught that conflict on day 4.
Failure mode: New agent sends outreach to a prospect that has a direct conflict with an existing client relationship. Client hears about it through industry contacts. Trust damaged.
If data is stale, flag it visibly. Never silently present old information as current.
Why: Stale data presented as current causes wrong decisions. Visible staleness lets the consumer decide how to weight the information.
Failure mode: Briefing shows yesterday's ad spend as today's. Founder makes budget decisions on wrong numbers.
If 3+ tasks from one person are overdue, flag as capacity pattern, not motivation problem.
Why: Individual overdue tasks might be forgotten. A pattern of overdue tasks indicates workload exceeds capacity.
Failure mode: Manager assumes delegation is lazy. Actual problem is team member is overwhelmed. Problem worsens.
Before inviting external users, audit in pairs: (1) for every authed endpoint, find its sibling that was missed (preview vs execute); (2) for every guard recorded in memory, verify it exists in CODE not just data — data guards don't cover future cases; (3) pool.on('error') + idleTimeout/keepAlive on any node-postgres pool behind a proxy; (4) never derive test-isolation ports from process.pid under vitest threads — use VITEST_POOL_ID + bind-retry.
Why: Each find was a class, not an instance: the unauthenticated endpoint was the sibling of a correctly-gated one; the memory-vs-code drift would have silently emailed the first paying customer; the pool crash explained the recurring ETIMEDOUT-fixed-by-redeploy pattern. Pair-auditing catches what single-point review misses.
Failure mode: SUCCESS: Claude/Conatus ran a full security+site audit of orgtp.com and shipped all fixes same-day. Key finds: /api/v1/merge/preview had ZERO auth (cross-org claim leak); the "Victor hard-guard" existed only as DB skip-rows, not code; pg pool had no error listener (idle Railway-proxy socket death crashed the whole process); vitest's THREADS pool shares process.pid so per-pid test ports collided.
For any design/polish pass, do a thorough comparative audit: re-study ALL reference material, enumerate many ranked findings (visual AND UX/flow), and show a comprehensive before/after, not a single tweak. Match effort to the word 'polish' = broad, not one change.
Why: A single cosmetic tweak reads as low effort and misses the actual gap. Design quality is cumulative; the value is in catching the full set, especially the UX/flow items that separate a product that 'looks designed' from ours.
Failure mode: Asked for a dashboard design polish pass, I shipped one nitpick (recolor a button) and called it done. David: "you only chose one... have another pass please and try better with a higher effort."
Before presenting a Dash blind-spot, billing trigger, or any state-file alert as a current action item, confirm it hasn't already been resolved. Stale state files (Dash May 25 was ~5 weeks old) carry point-in-time alerts that may be closed by now. Trust confirmed/observed status over stale notes; flag the data's age and treat unverified alerts as 'verify' not 'urgent.'
Why: Re-surfacing already-resolved alerts as urgent erodes trust in the L10 briefing and spends David's attention during a low-push recovery window. Honesty about data staleness matters more than appearing comprehensive.
Failure mode: Dan surfaced the HiTone billing trigger ($43-49K/mo possibly un-invoiced) from Dash's stale May 25 state file as a live concern during the Jun 29 L10. David confirmed HiTone billing is correct and already handled, and asked to close it out.
When /billing-report's spend line shows Google $0.00 while Meta is non-zero, treat it as a broken pull, not a real zero. Test the Google Ads REST API directly against the MCC at v21/v22 before writing the sheet; if a newer version works, bump GA_API in billing_pull_spend.py (line ~26) AND the MCP server's API_VERSION. Never write a billing sheet on a total-zero Google pull.
Why: Billing accuracy is paramount — a silently-empty Google pull under-bills every Google-only/Google-heavy client (SSP, GLS, Invictus, dual-platform accounts). The version constant exists in two places and the error is swallowed, so the failure is invisible unless someone notices Google totaling $0.
Failure mode: SUCCESS: Billing-report caught a silent $0 Google-spend pull before writing invoices. /billing-report's spend puller (billing_pull_spend.py) pins its own Google Ads API version (GA_API), and ga_search() swallows HTTP errors — so a stale version returns zero accounts and $0 Google spend across the whole MCC silently. First run showed Google $0.00 / total $8,780 (real numbers Google $50,994 / total $9,850, a $1,070 underbill).
KPIs/scorecards must live as tiles in OTP (the source of truth), not in markdown files or meeting briefs. Every active agent/human seat — including Dan's strategic co-founder seat — must own at least one OTP KPI tile. When proposing measurables, verify against list_my_kpis and create the missing tiles via update_kpi (auto-creates), rather than just tabling them in a doc. A seat with no number is sitting on the sidelines.
Why: EOS requires every seat to have a measurable. Discussing KPIs in a brief while OTP shows none of them makes the scorecard fiction and undercuts OTP as the coordination source of truth. Dan as co-founder must be measurable like everyone else.
Failure mode: Dan presented a Sneeze It agent-team scorecard as a markdown table in the L10 brief and treated it as 'the scorecard,' when the source of truth is OTP. David caught that Dan (and Arin/Pulse/Dirk) have NO KPI tiles in OTP at all — Dan's own seat had zero measurables. A scorecard that only lives in a file or meeting brief does not exist.
tsc + ejs.compile prove it compiles, not that it LOOKS right. For any visual/UI change, get a rendered screenshot before declaring it done. When the live page is authed and can't be driven, ship to an opt-in lab and treat the user's screenshot as the verification gate, then iterate. Also: flex action buttons need whitespace-nowrap + shrink-0 or they collapse.
Why: Compile-clean UI can still be visually broken or alarming. Blind UI shipping burns trust and ships regressions the user has to catch.
Failure mode: Shipped clean-dashboard UI (KPI heatmap, etc.) verified only by tsc + ejs.compile. Live render showed a broken Headlines "Add" button (text collapsed to vertical) and a heatmap that read as an alarming pink wall. David: "I can see where you were going but look at what actually occurred."
Before any OTP deploy: (1) confirm which lineage prod actually runs by comparing origin/main against both worktrees (origin/main has matched the otp-audit-fixes lineage, NOT otp-platform's working branch); (2) deploy by pushing to origin main (Railway service is GitHub-connected to david-steel/OTP and auto-builds in ~2 min) instead of railway up when the CLI upload endpoint times out; (3) never railway up from a dirty otp-platform tree, since up snapshots the whole working directory including unrelated uncommitted changes
Why: railway up uploads the local tree as-is. Deploying from otp-platform (branch merge-execution-provenance + uncommitted edits) would have silently rolled back the 2026-06-10 security hardening and live-audit fixes and shipped unvalidated merge-execution code. The 7 consecutive upload timeouts were what prevented a prod regression; the GitHub push path is both safer (commits only) and faster (135s to live).
Failure mode: SUCCESS: Claude shipped OTP via GitHub push after railway up timed out 7x, and caught that prod runs the otp-audit-fixes worktree lineage, not otp-platform's checked-out branch
Root causes were independent: (1) the team-chart invite UI (dashboard-team.ejs sendInvite + create-tile path) sent only {email, claimedEntityId, role} and dropped the tile's name, so org_members.display_name was null -> "(no name)"; the /dashboard/members page did no Clerk enrichment (the team-page enrichment that exists was gated on !displayName AND !email, skipping rows that have an email but no name). (2) acceptInvite only runs on the tokenized /accept-invite link, so a user whose ?token= was lost across the Clerk sign-up round-trip gets a Clerk account, no org_members row, and an invite stuck 'pending' forever (stranded). Fix: pass displayName from the tile label in the frontend; fall back to Clerk profile name at accept time; heal null names on the members page from Clerk; and add acceptPendingInviteByEmail() called on the /dashboard no-org branch to auto-convert stranded users. Diagnostic heuristic: when invited members "come in odd," separate the INVITE step (what issueInvite stored: displayName, claimedEntityId) from the ACCEPT step (did acceptInvite run at all). Same invite + different outcomes => the divergence is in accept, not invite.
Why: Invite/membership is the highest-stakes onboarding path; nameless or stranded members erode trust the moment a customer invites their team. The invite vs accept decomposition turns an "unnerving/random" report into two deterministic, separately-fixable defects.
Failure mode: SUCCESS: OTP invite flow had two distinct gaps making chart-invited members come in "wrong" (one stuck pending+stranded, one a member but "(no name)").
When moving or extracting an EJS partial to a different directory depth, rewrite every relative include() path inside it for the new location (here: '../partials/X' -> '../X'). And do NOT trust ejs.compile as the gate for include correctness — it only catches parse errors, not include resolution. Verify nested includes by either rendering the partial, or asserting each include target resolves to a real file from the partial's new path (grep the include() calls, map to files, test -f). Add a render/path check to the verification step for any partial extraction.
Why: Caused a live, ungated 500 on all built-in L10 meetings (scorecard section) right after shipping the agenda-driven l8-leadership refactor. Compile-clean gave false confidence; the real gate for includes is render-time resolution.
Failure mode: Extracted EJS partials to a deeper directory (src/views/partials/meeting/) but left their nested include() paths relative to the OLD location: '../partials/ui-pill' and '../partials/rich-description-editor' were correct from src/views/pages/ but resolved one level too deep from the new home, 500ing every meeting with a scorecard ("Could not find the include file ../partials/ui-pill"). EJS compile (ejs.compile) PASSED because it does not resolve includes — the break only surfaced at render in prod.
An OTP KPI with teamId=NULL renders only on /dashboard/kpis, never on any L10 scorecard (meeting scorecards filter strictly by meeting.team_id). To make a KPI show on a specific L10, PATCH /api/v1/kpis/:id with the meeting's teamId. The 'Dan L10' meetings run on the 'ai-army' team (065d1d4b-c7da-4e80-b3ed-d6b101471d2c). Tally's auto-create now includes teamId from a 'team_id' field in the registry entry, so new agent-army KPIs land on the Dan L10 automatically instead of orphaned. Find team IDs via GET /api/v1/teams; meeting->team via GET /api/v1/meetings.
Why: A KPI nobody can see on their meeting scorecard is functionally not on the scorecard. The owner/title is necessary but not sufficient — team scoping is what makes it report. This is a recurring gotcha for any agent creating KPIs via the API.
Failure mode: SUCCESS: Tally — agent KPIs were invisible on the L10 because auto-create left teamId NULL. David flagged that the new KPIs weren't reporting on the Dan L10 or /dashboard/kpis as expected.
Before uploading a local file to Drive via create_drive_file with file://, copy it into /Users/dsteel/.workspace-mcp/attachments first, then point fileUrl there. Verify the upload returned a file ID and surface failures loudly.
Why: Broken path plus a fail-quietly rule means the delivery step fails on every run with no signal. Applies to all agents writing to Drive, not just Dash.
Failure mode: coach-report Drive upload failed silently every run. create_drive_file used a file:// URL at ~/.claude/, but the google-workspace MCP tool only reads files in /Users/dsteel/.workspace-mcp/attachments, so it errored and the spec hid the failure.
Root cause was the SearchAtlas OTTO pixel having an EMPTY src="" in the layout head (v7.ejs, onboarding.ejs, main.ejs). The OTTO tag must carry its base64 data-URI loader in src that appends dynamic_optimization.js with data-uuid; with src="" the runtime never loads, so OTTO injects/verifies nothing. When an OTTO/SearchAtlas audit reports 0/N across ALL on-page categories, suspect the pixel loader, not the actual tags — verify the sa-dynamic-optimization script's src is populated, not the page's own meta.
Why: A 0/16 across every category despite visibly correct meta is the signature of a non-loading optimization runtime, not missing tags. Checking the pixel first avoids a pointless rewrite of titles/descriptions that were never the problem.
Failure mode: SUCCESS: Beacon/SEO — orgtp.com OTTO on-page audit showed 0/16 (titles, meta descriptions, headings, meta keywords all failing) even though pages had perfectly good title tags and meta descriptions server-side.
When the CCM Speed To Lead sheet shows a negative STL (~-235 to -239 min), the sub-account's first-call timestamps are logged ~4 hours behind ET (Mountain Time CloudCRM sub-account config). Treat the lead as called and the STL value as unusable; never flag the project for slow lead response off a negative row. Currently affects all 6 Villa Sport Fitness projects.
Why: A naive STL scan either averages negative values (masking real slow responses portfolio-wide) or flags Villa for impossible call times. The -237 constant equals the 4-hour TZ offset, so it is diagnosable and filterable. Root-cause fix is the sub-account timezone setting, escalated to Dash 2026-06-12.
Failure mode: SUCCESS: Arin identified that negative Speed-To-Lead values in CCM-Stats are a timezone artifact, not bad calling
A brand battle cry needs a genuinely designed moment (confident display type, intentional line breaks, brand device, real whitespace), not a centered text block plopped in. And the VISIBLE battle cry copy is the short clause only: 'Unlocking the potential in every person through the partnership of people and AI' — drop 'so together we leave the world better than we found it' from the hero display (keep the full sentence only for formal/footer contexts).
Why: A mission line is a brand centerpiece. Long copy dilutes the punch, and an undesigned drop-in reads as filler. The payoff phrase 'partnership of people and AI' must land as the climax with design weight behind it.
Failure mode: Adding the OTP mission as a 'battle cry' on the landing page, I dropped the full sentence into a plain centered text band wedged between hero and Step 1. David called it 'a weak attempt to just throw it on the page' and said the full line is too long for the visible battle cry.
Distinguish agent vs tool when sequencing monetization. AGENTS (adoption-driven, sticky surfaces) can ship free to drive usage. TOOLS that ARE the upsell (AI-assist, metered features) must be PAID from first use — never free-first. Build them paywalled/gated behind the wallet (GHL 'turn them on' pattern: free users SEE the button + 'add credits' upsell, never get the output free), ready to monetize the instant billing is live. Giving away the upsell trains users to expect it free and destroys the price.
Why: Free-first on a paid tool kneecaps the revenue model it was built for. The value prop of AI-assist is that it's metered; free distribution undercuts the exact margin lever (markup multiple) the whole Phase 2 wallet exists to capture.
Failure mode: Proposed shipping OTP's Rock AI-assist FREE-first to gather usage/price signal while Stripe paperwork clears. David corrected: this is a TOOL that is the upsell, not an agent. Free-first is wrong for it.
Any agent doing client identification or contact audits must reconcile two sources of truth: pepper-clients.md (curated, human-approved) and Accelo list_companies with status=active (system of record). If an Accelo-active domain is not in pepper-clients.md, do NOT auto-add. Flag it to David as a PROMOTION CANDIDATE with context (company name, Accelo ID, why it's missing). David makes the call because some Accelo companies have domains intentionally excluded (e.g. goldsgym.com is a multi-franchise brand where only a few locations are Sneeze clients; qualitylearning.net is Cellebration Wellness's parent domain but unknown individual contacts still need human confirmation). Inverse check also matters: any domain in pepper-clients.md with no active Accelo company is a candidate for removal.
Why: pepper-clients.md is downstream of many agent decisions: Pepper buckets emails as CLIENT vs NOISE, Dirk suppresses cold outreach to existing customers, Dash scopes active-client analysis, Radar flags client comms. When it drifts, every downstream agent silently makes wrong calls - no error, just degraded signal. Specific risk is cold-emailing someone Sneeze is billing, which damages trust. Fix is cheap (30-second reconciliation during any client audit) and prevents a class of errors that is otherwise invisible until a client or teammate points it out.
Failure mode: Client-list audits silently misclassified active Sneeze It clients as cold prospects because pepper-clients.md (the canonical client-domain allowlist that Pepper, Dirk, Dash, and Radar all read) drifted out of sync with Accelo (the authoritative active-company record). In a 2026-04-23 audit of 1,790 contacts, 3 domains were active in Accelo but missing from pepper-clients.md: almarose.com (Alma Rose), delawaredigitalmedia.com (Delaware Digital Media white-label), studstillfirm.com (Studstill Firm). Contacts on those domains were being treated as cold prospects, which risks Dirk sending a cold email to someone Sneeze is actively billing.
Never infer a 'David <> [Company]' calendar title is a sales/prospect call from the title alone. Check the attendee email domains first. Law-firm domains (e.g. mdwcg.com = Marshall Dennehey, thompsoncoe.com = Thompson Coe) and insurer domains (e.g. usli.com = USLI, an E&O carrier) signal a LEGAL/litigation call, not sales. When attorney or insurance-carrier domains are present, classify as legal/sensitive and never frame it as pipeline or prep it as an outreach/sales call.
Why: Misclassifying active litigation as a sales opportunity is a serious tone failure — it could lead an agent to draft sales follow-up to opposing counsel or surface a lawsuit as a 'win' in the pipeline. Attendee domains are the reliable signal the calendar title hides.
Failure mode: Radar/good-morning labeled a calendar event titled 'David <> Jeremy Golds Gym South Texas' as a sales prospect call. It was actually an E&O litigation defense call — Gold's Gym South Texas is suing Sneeze It, and the attendees were defense counsel and the E&O insurance carrier, not a prospect.
Manifesto/mission pages must be written as movement recruitment, not product marketing: second-person address (the reader is the protagonist), "We believe" creed statements people can recite, a named enemy, stakes, and invitation CTAs ("Join the movement") instead of transactional ones ("Start free"). Product features appear only once, framed as how the movement fights, not what the product includes.
Why: People join movements because they believe what the movement believes (Sinek: start with why). Copy that sells the what on a page whose job is to recruit believers reads as generic SaaS and inspires no one, no matter how good the design is.
Failure mode: Redesigned the orgtp.com manifesto homepage with strong visual design but kept product-brochure copy (feature lists, "free meeting software", "Start free" CTAs). David: "the writing does not inspire an army of followers... this just looks the same as every other company... blah."
Judge conversion on the full path the visitor actually walks (page, door, day-one experience), not on the surface being edited. If the honest answer to "would you sign up" is "yes IF another surface delivers," the answer is no, and the work moves to that surface. Never write a promise on a button that the destination page cannot cash.
Why: Trust destroyed at the moment of verification is unrecoverable; a skeptical buyer who clicks "watch us run" and lands on a data page is gone forever. Copy that outruns proof is hype by definition, and the exact audience OTP needs (operators) is the audience that punishes it hardest.
Failure mode: After rewriting the OTP homepage, I declared the copy converts because skeptics would "click through to the live OOS page and sign up IF it delivers." David called it: I kicked the can to a page I know does not deliver, and called it a win. The button promises "Watch our company run, live" but the OOS page it links to is a list of published rules, not a running company.
On any rename/rebrand, grep every occurrence and fix ALL visible references in one pass (title, alt text, headings, signoff, sender from-name, comments). Don't propose a partial fix or leave 'good enough' residue. Verify with a re-grep of the rendered output.
Why: A customer email with mixed old/new branding looks careless and undoes the rename. Half-measures on a visible rename are worse than not starting.
Failure mode: When rebranding Orgy -> Ollie in the weekly email, I made one edit and framed leaving the rest (alt text, comments, from-name) as acceptable. David: "lets not be lazy fix it there are only 2 Orgy ref on the email."
OTP's enemy statement is "you bought the operating system and the needle didn't move." The pitch is not better meetings; it is: the system was fine, what was missing was the workforce that runs it between the meetings. Frame all homepage/sales copy against needle-not-moving, not against meetings.
Why: This is the buyer's actual lived disappointment (paid for an operating system, company looks the same two years later) and it positions OTP against incumbents on outcomes instead of features.
Failure mode: The letter's hero framed the enemy as "the meeting" / busywork. David corrected the thesis: the real problem is that companies bought operating systems and software (Ninety, Bloom Growth, etc.) that did not move the needle. Years later the company had not grown and was not better, and they needed to change how they did things.
For any UI change, design from the user's mental model, not the data model: "my list shows my work; work I assigned to others shows under Waiting on Others." When a meeting todo is assigned to someone else, stamp the creator as delegator so it routes to the delegation view. Before shipping UI changes, run the UX lens (impeccable / web-design-guidelines skills + src/DESIGN.md), not just a minimal code patch.
Why: A technically-correct patch that ignores the user's mental model just moves the confusion. OTP's own product language already has the right home for these items (Waiting on Others); fixes should land in the model the user already understands.
Failure mode: Fixed the dashboard todo confusion (teammates' meeting todos looked like the viewer's own) by adding an owner label to the rows. David corrected: that's not thinking like a user. Labeled-or-not, other people's todos don't belong in "my to-dos" at all.
When David asks for a jaw-drop brand page, build an EXPERIENCE, not an article: full-viewport cinematic hero, scroll choreography, one idea per screen at massive scale, motifs that live in the page as motion, ruthless copy cuts, no standard nav/footer chrome breaking the spell, no section-grammar scaffolding, no FAQ accordion bolted onto a manifesto.
Why: The gap between "well-executed page" and "omg I love this" is the whole assignment on brand surfaces. Safe editorial structure is invisible at best; for-the-brave positioning demands the page itself be brave.
Failure mode: Built the /ollie manifesto page as a competent editorial layout (repeated mono eyebrow labels on every section, index rows, alternating light/dark sections, FAQ accordion at the bottom) and David rejected it outright: "this really really sucks." The brief was "reader drops on the ground saying omg I fucking love this" and the output was a safe template that reads as AI scaffolding.
When David gives a design reference URL, open it in a browser and STUDY it visually (proportions, type sizes, spacing, alignment) before designing; match its register, not just its layout skeleton. Elegant means restrained: modest type scale, centered calm hierarchy, generous whitespace, thin rules. Never hand-draw SVG artwork to imitate produced brand art; crop/reuse the actual asset or use nothing.
Why: A reference URL is the brief. Reading its HTML structure without seeing it rendered led to importing the skeleton with the wrong soul, twice. Amateur freehand art next to professional motion work destroys credibility instantly.
Failure mode: Second rejection on the /ollie page. David asked for sakana.ai/fugu: elegant, Japanese sense of design (restraint, whitespace, calm, modest type, precision). I delivered giant 9vw headlines, one shouting line per viewport, and hand-drawn SVG chevron "birds" that rendered as crude fat marker scribbles. I treated "jaw-drop" as scale and boldness when the reference was quietness and precision, and I drew freehand SVG art instead of using the actual video's artwork.
For fleet-wide spec maintenance: (1) tarball backup of ~/.claude before any agent touches specs; (2) partition files into DISJOINT clusters, one agent each, with CLAUDE.md owned by exactly one; (3) give every auditor the same stale-fact canon and the rule "verify a launchd plist exists before believing any schedule claim"; (4) auditors apply surgical edits directly for factual fixes but RETURN structural proposals for David instead of applying them; (5) synthesizer closes cross-cluster contradictions the auditors flag at each other.
Why: Agent specs rot faster than anyone audits them: this pass found live specs for a retired agent (jeff.md ending in "Go."), four phantom schedules, Todoist writes in five files, terminated employees still routed DMs, and a Bassim score-inflation bug. Periodic fleet audits with disjoint ownership are cheap insurance against agents acting on dead infrastructure.
Failure mode: SUCCESS: Claude ran a five-cluster parallel level-up of the entire agent army (80 files, ~140 surgical edits) without a single file conflict or lost spec.
Pattern for UX dead-end hunts: (1) fan out parallel read-only explorers per surface (meetings, teams/members, KPIs/todos, onboarding/settings) asking for file:line + user-visible symptom + minimal fix; (2) fix the unsatisfiable states first: any required dropdown that can render zero options must explain where its options come from and link there (owners/attendees come from the org chart, meeting membership from teams); (3) empty states must branch on WHY they are empty (org has no teams vs user not on a team need different CTAs); (4) never report an async side effect as done: invite emails now await sendEmail (which returns null on failure, never throws) and return emailSent so the UI can tell the truth; (5) a guided setup checklist computed server-side from actual data (seats/team/KPI/meeting/members exist?) beats static onboarding because it survives skipped onboarding.
Why: These are the recurring shapes of broken UX in OTP: forms with prerequisites the user cannot see, empty states that misdiagnose their cause, and optimistic success messages over fire-and-forget side effects. Fixing the shape, not just the instance, is what makes the product feel intuitive.
Failure mode: SUCCESS: Claude ran a full UX dead-end audit and fix pass across OTP (4 PRs, #113-#116, all deployed)
The sweep pattern that worked: audit by rule-cluster in parallel (fakery, insight-to-agency, jargon/states, first-meeting goal-walk), then execute severity-first. Key catches to re-check every run: (1) seeded/synthetic data leaking into numbers a reader believes are real (the is_template flag existed but was never enforced; counts now use src/shared/synthetic-orgs.ts); (2) the conversion moment must be ON the default path (end-meeting now lands on Ollie followups, not the list); (3) funnels don't exist until instrumented (insight topic: surfaced/accepted/value_delivered); (4) credentials in seed script comments (one prod DATABASE_URL scrubbed; password rotation still owed). Worklist for run 2 in otp-platform/mission-standard/WORKLIST.md.
Why: The Mission Standard is a repeatable bar, not a one-off audit. Recording the found failure classes makes run 2 start from run 1's ceiling instead of re-discovering it.
Failure mode: SUCCESS: Claude ran Mission Standard sweep run 1 (PRs #121-#123, deployed): 4 parallel rule-audits over OTP, 5 CRITICAL + 13 GAP found, all CRITICAL and 9 GAP closed same-session
When an agent-pushed OTP to-do references a document, include a clickable https link (a Google Doc), not a local file/vault path — David reviews to-dos on mobile. To update an existing to-do's description, use PUT /api/v1/todos/:id (not PATCH). otp-todo.sh has no update verb, so PUT directly with the API key.
Why: A file path in a to-do is dead weight on mobile — the reviewer can see the reference but cannot open it, which reads as "the link is missing/broken." Every agent that pushes doc-linked to-dos (Radar, Pepper, Dan) hits this.
Failure mode: Dan pushed an OTP to-do referencing a document but put a local Obsidian vault path ("2nd Brain/Agent Army/Dan/...") in the description. David opens to-dos on his phone — a vault path is not tappable, so there was "no link to click." Also used PATCH to update the to-do; the OTP todos API update method is PUT /api/v1/todos/:id (PATCH hits the marketing site and returns HTML).
To put an agent-army/IDS issue on an OTP meeting board, POST /api/v1/tickets with the team's teamId (category 'other' for strategic issues, priority low/medium/high/critical, ownerEntityType+ownerExternalId). The MCP submit_ticket tool CANNOT do this — it has no teamId param (it is the generic 'report a bug to OTP' path), which is why nobody ever got issues onto the board. Team IDs: 'AI Army' = 065d1d4b-c7da-4e80-b3ed-d6b101471d2c (the David+Dan agent-army meeting); Leadership Team = c1e1a485-414e-48d5-ae44-e81bd110b554. Update/solve via PUT /api/v1/tickets/:id (idsStatus, priorityRank, resolution).
Why: Agents could push KPIs and todos to OTP but not issues, so every L10 IDS board rendered empty and David kept discovering the hole live. Issues=tickets + teamId scoping is the missing piece; without it a meeting-readiness check would keep mislabeling a working API as absent.
Failure mode: Dan claimed 'no issues API exists in OTP' because there is no src/routes/api/issues.ts. That was wrong. OTP stores IDS issues in the TICKETS table (schema.ts: 'issues live in the tickets table'), with full IDS support (idsStatus, priorityRank, teamId, owner fields). The agent-army IDS board was empty only because our issues lived in a local markdown file and were never pushed as tickets scoped to a team.
OTP has TWO distinct features both called "Ollie Insight": (A) the per-meeting followups wizard that turns a transcript into meetings.ai_summary (src/shared/meeting-followups.ts, transcript-only), and (B) the reusable "address engine" (ollie_insights table, src/services/ollie-insight.ts + shared/ollie-insight.ts + partials/ollie-insight-block.ejs) that gathers org data (KPIs/rocks/todos/meeting-summaries) per SCOPE. When David said "KPIs shouldn't be in the meeting analysis," the fix was in System B's meeting-scope evidence gathering, NOT System A. The block partial (ollie-insight-block.ejs) is fully scope-generic (builds the API URL from data-oib-scope/scopeId client-side), so adding a brand-new 'quarter' scope end-to-end took only: add to INSIGHT_SCOPES + RULES_BY_SCOPE (shared), the scopeQuerySchema enum + a resolveInsightScope branch (api), a gatherEvidence branch + max_tokens (service) -- then just include the existing partial with scope:'quarter'. No new render/generate/receipts UI. Pattern: when a feature is "one engine, many surfaces," new surfaces are a scope + evidence branch, never new UI.
Why: The name collision hides which code to touch; picking the wrong system wastes a whole edit pass. And recognizing the scope-generic block means big-feeling asks ("a quarterly synthesis button") are small, low-risk diffs. Both are recurring shapes in OTP's Ollie work.
Failure mode: SUCCESS: Claude separated OTP's two "Ollie Insight" systems and added a whole new scope by reuse
New endpoint POST /api/v1/meetings/:id/agent-record (PR #154): an agent submits the written meeting record; OTP runs the same redaction ruleset, stores it to meetings.transcript, logs an audit baseline + agent_record event, and the existing /ai/followups generate turns it into to-dos/issues/headlines/insight unchanged. Agent path: ~/.claude/otp-meeting.sh record <meetingId> --file=<record> --source=l10dan. Wired into /l10dan conclude step 6. Verify deploy by probing the endpoint returns JSON not the marketing SPA HTML before pushing records.
Why: Ollie only reads transcripts, so agent-run meetings had no way into it — the empty-insight hole David hit live. This closes it: any agent meeting can now produce Ollie follow-ups. Also a general UI rule captured: buttons reflect capability/state (action taken -> button disappears).
Failure mode: SUCCESS: Dan shipped the Ollie agent-record path so agent-facilitated meetings (David + AI L10, no audio transcript) can feed Ollie Insights.
Before treating a KPI/data source as blocked, re-read the LIVE source, not the note about it. The Havok "non-client %" KPI was marked blocked for ~14 weeks on "the timesheet has no client column" — a 105-day-old memory. The live sheet (1VPlH5ZqTowOe2nJDFwOjuvXrCXUnOzeOb3aO-M936Xo) had since grown per-person tabs with a full Client/Time/Date schema; one read unblocked it. Pattern for wiring a messy sheet into Tally: (1) get_spreadsheet_info to list tabs, (2) read a person tab to learn the real schema, (3) confirm recency by reading the tail (last date), (4) add a focused extract mode to tally.py rather than reshaping the sheet — here `client_attribution_since` (reads ALL valueRanges via a new _values_2d_all, parses h:mm via _parse_hhmm, added dotted DD.MM.YYYY to _parse_date, classifies internal by an `internal_contains` substring), (5) `tally.py --dry-run --kpi "<title>"` to prove the number off live data before pushing. Human-owner + a 1:1-team KPI (owner HUM_BOGDANTABAKA, team = David-Bogdan 1:1) pushes fine via find_or_create_kpi.
Why: Blocked-status notes rot silently while the underlying source improves; a KPI can sit "pending" for a quarter when it was buildable weeks ago. Re-reading the live source first is the cheap unblock. And the tally.py extract-mode pattern makes any timesheet/sheet a live KPI without asking a human to restructure their doc.
Failure mode: SUCCESS: Dan/Tally shipped the Havok client-attribution KPI live in one session after a 14-week "blocked" note turned out stale
Before editing any OTP view to fix an on-screen bug, grep unique visible strings from the screenshot (e.g. "WAITING ON OTHERS", "always only yours") across src/views to confirm WHICH template renders that exact surface. Multiple pages can render similar-looking todo lists (me-todos.ejs vs dashboard-daily.ejs). Verify the rendering route (reply.view target) too.
Why: Two round-trips and two merged PRs produced zero visible change because the edits were on the wrong template, which read as "nothing is fixed" and eroded trust. A 10-second grep on the screenshot text would have pointed to the right file immediately.
Failure mode: Fixing an OTP todos UI bug, I edited src/views/pages/me-todos.ejs twice and shipped two PRs, but the surface David actually uses is the dashboard-daily "Waiting on others" widget (src/views/pages/dashboard-daily.ejs). Nothing he saw changed.
Two reusable patterns: (1) Before building any OTP email/engagement feature, grep src/services for existing infrastructure -- re-engagement.ts, lifecycle-scheduler.ts, and user_engagement_log already carried cadence caps, suppression, logging, and a daily cron, so the todo-aware upgrade was ~350 lines instead of a new subsystem. (2) When resolving "which org does this Clerk user belong to", organizations.clerkOrgId only knows the org CREATOR; invited teammates must be resolved through org_members.clerkUserId + claimedEntityIds. This gap is why per-user personalization (open todos) missed non-creator members like Nate.
Why: One engagement channel with shared caps is what keeps daily utilization pressure from becoming annoying double-mailing, and the creator-vs-member resolution gap will bite any future per-user feature (digests, notifications, billing seats) that starts from organizations.clerkOrgId.
Failure mode: SUCCESS: Claude shipped the smart engagement email engine (PR #186) by upgrading the existing re-engagement service instead of building a parallel system
Frame help/success/onboarding call copy positively: state plainly that the call is there to help and guide them, and describe what will actually happen on it (we'll walk through your setup, get you unstuck, answer your questions). Never say "this is not a sales call" or "no pitch" — describe the help, don't disclaim the sell.
Why: Defensive "not a sales" language triggers the exact suspicion it tries to defuse and undercuts a genuine help offer. David flagged this immediately.
Failure mode: Wrote a "customer success call" Calendly description that leaned on "no pitch, no slides" / not-a-sales-call framing. Protesting that it isn't a sales call makes it sound like one.
Any script that answers "is anything missing / is everything covered?" must fail LOUD, never return an empty set as reassurance. Two rules: (1) assert the expected top-level key exists (`if 'data' not in resp: raise`) before computing a result; a zero/empty answer from a health check is a claim that must be proven, not a default. (2) Cross-check a zero result against one known-positive case before reporting it -- here, one direct call to a single account would have shown $57 of spend and exposed the lie instantly. Also: `~/.claude/meta-ads.sh accounts` exits 0 and prints nothing; do not build on it. Sweep with `/{business_id}/adaccounts?fields=name,account_status&limit=500` then batch `/{act_id}/insights` 50 at a time.
Why: "Nothing is wrong" is the single most dangerous output an audit can produce, because nobody investigates it. A silent empty result on a billing sweep means real revenue is never invoiced and nobody ever finds out. The failure mode is not a crash, it is confident silence.
Failure mode: SUCCESS: Claude caught a silent false-negative in a Meta billing sweep. Querying the Graph API adaccounts edge with nested field expansion (`fields=name,insights.date_preset(this_month){spend}`) returned an error payload with NO `data` key. The sweep script read it as an empty account list and confidently reported "0 accounts, $0.00 unbilled MTD spend" -- a clean bill of health that was entirely fabricated. Direct per-account queries then revealed 7 unbilled accounts spending $5,673 MTD.
When you own a strong long-form asset (a landing page that argues the WHY), the cold email must NOT re-argue it in miniature. Cut the product explanation entirely. The email's only job is to earn one click: proof you read their work, one line naming THEIR problem in THEIR language, the link, out. No feature, no price, no call ask. Match the page to the reader's specific pain rather than sending everyone to the same URL. And assign/log an A/B arm on every send, because an untested signature shipping at 100% for months is not a decision, it is a habit.
Why: A cold email is the worst possible venue for explaining a product: no trust, no attention, no context. A page is the best. Using the email to sell the READ instead of the PRODUCT plays each asset to its strength. The deeper failure was cheaper: months of sends with no arm logged and no reply column filled, which means the volume produced zero learning. Volume without measurement is just noise you paid for.
Failure mode: SUCCESS: Crafter cold-email rework -- the email body was trying to explain the product to a stranger in three sentences, while the website spent full pages arguing the why. The two assets contradicted each other and the email was losing. Outreach log had been dead since Apr 26 with effectively no replies.
When compiling guru influences into an Ollie persona, push voice fidelity hard: the dominant influence's cadence, vocabulary, and signature moves should LEAD the writing, not decorate it. Channel the style ("in the room, guiding in their voice") while keeping the hard never-impersonate line: never claim to BE the person, never claim endorsement. Also: differentiation surfaces (the Lab) must let users edit each guru's forked principles inline and must show WHERE a principle or voice move landed in the output (highlights on the insight, influence tags on todos/issues/headlines) so voices can actually be compared.
Why: The entire value of per-team Ollie voices is that the voice audibly changes what the team hears. If influences only shift content, not voice, the differentiation is inaudible and the Lab proves nothing. Attribution marks are what make the difference legible.
Failure mode: Ollie Lab guru influences rendered as principles-with-a-hint-of-style: the output read like Ollie citing a thinker, not like the thinker's voice guiding the room. David: "it should be as if they were speaking with their voice guiding us in their voice... It should be like they are in the room."
Voice fidelity lives in structure, never slogans. Ban catchphrase quoting explicitly in the persona compile ("never quote their slogans; that is imitation's cheapest form"). Give each library guru a hand-written VOICE DNA block: sentence rhythm and length, how they open, how they build an argument, what they notice first, how they land a point, emotional register, what they never do. For custom gurus, instruct the model to reconstruct the person's published voice from structure and register, not taglines. Method instruction: "before writing, ask how NAME would structure this and what they would notice first; write from inside that mind."
Why: Catchphrases signal imitation and break trust instantly; structural voice makes the reader feel the thinker in the room without a single borrowed phrase. This is the difference between a costume and a mind, and it is the entire premium of per-team Ollie voices.
Failure mode: Ollie's guru voice channeling produced surface mimicry: dropping the thinker's catchphrases ("Start with why") instead of writing from inside their rhetorical DNA. David: "not cheap tricks... it needs to be in the DNA of what Ollie is saying. Dig deep, make it real."
Essay-page hero art is a first-class deliverable, not a wireframe: match the family's craft bar (a real scene that tells the page's story, Ollie present, warm accent treatment, gradients/glow, staged ignition-style animation, reduced-motion complete state). Before drawing, read the sibling page's actual SVG to absorb its techniques, then design a scene, not a diagram.
Why: The hero art IS the argument at first glance on these pages; the candle page's flames are the thesis made visible. A schematic undercuts an inspiration page precisely where it must inspire.
Failure mode: The /the-voice-in-the-room hero art shipped as a lazy schematic (dot, four thin lines, grey boxes with circles for people) while its sibling pages (/one-candle, /mission-to-the-moon) carry hand-crafted narrative scenes with the mascot, gradients, glows, and staged animation. David: "you got lazy there for sure."
When running counterfactual/what-if simulations, make the events endogenous: model causal mechanisms (innovation rate, demographics, conflict outcomes) as functions of the changed variable and run Monte Carlo over branching timelines, rather than overlaying new participation rates on the fixed historical record. Show a distribution of divergent timelines, not one re-skinned version of real history.
Why: Fixed-event counterfactuals smuggle in the answer (the world converges because you forced it to). Branching simulation is what actually answers "would things be different" questions, and it's the difference between a re-labeled chart and genuine out-of-the-box analysis.
Failure mode: Built a counterfactual history simulation that held all real-world events fixed (same Industrial Revolution, same wars, same tech timeline) and only varied participation rates within them. David flagged that this assumes the conclusion: if the labor/care allocation changes, the events themselves change — maybe industrialization comes later (or earlier), wars resolve differently, the whole timeline branches.
In live L10/Delta Meeting facilitation, sequence is: surface ALL signals first (grouped, neutral, no recommendation attached), let David react and pick what to IDS, and only then frame decisions. Decisions come after shared context, not before. Also: open every meeting with Ollie's read of the PRIOR meeting's record (pulled from the OTP followups/insight), as a standing first section — David explicitly values it.
Why: A decision framed before the signal review railroads the meeting toward Dan's framing and skips the part where David's pattern-recognition works on the raw material. The facilitator's job is to lay out the board, not to compress it into a pre-picked fork. And the Ollie prior-meeting read is the continuity loop that makes each meeting compound on the last.
Failure mode: Dan facilitated the live L10 by pushing straight to the rock-set timing decision immediately after the scorecard, without first walking through all the signals (client wins, CC trends, attribution reading, unverified tiles, churn signals) so David could see the whole board. David: "we need to review all the signals before moving so quickly, you are trying to skip ahead too fast."
L10 prep MUST include a live scan of OTP itself: the team's rocks/priorities board, the issues (tickets) board, todos, and KPIs, pulled fresh at prep time. OTP is the source of truth; local files are mirrors that go stale the moment David works in the product directly (which is the whole point of OTP). Additionally, per David's 7/13 ruling: local shared-state files are now formally HISTORICAL unless a live consumer reads them; staleness flags on superseded files are noise, but staleness in the OTP scan is a real miss.
Why: David increasingly works inside OTP directly (weekend rock-setting), so any prep that skips the live product misses his most recent decisions and re-surfaces solved or stale items. The mirror-drift failure has now happened twice (June 15, July 13); the fix is structural, scan the product, not the mirror.
Failure mode: Dan's L10 prep read local mirror files and this morning's Tally/KPI pipeline but never scanned the live OTP boards (rocks/priorities and their attached issues) before the meeting. Result: Q3 rocks David added over the weekend were missing from the prep, and stale issues sitting on the rocks board went unnoticed. David caught it live, again (first time was June 15).
Every KPI on any board must pass the needle test before it earns a tile: one sentence stating the causal chain from this number to the company goal (margin, retention, revenue, rock completion). Dan owns running this test — on every existing tile quarterly and on every proposed tile before creation. Activity metrics (emails drafted, projects counted, pushes made) are health checks at best; they live in the readiness script, not on the scorecard. Wiring a dead tile is worthless if the tile measures the wrong thing.
Why: The strategic co-founder seat exists to hold the big picture David cannot hold while operating. A perfectly-wired scorecard of needle-irrelevant numbers is worse than an empty one, because it manufactures the feeling of accountability without the substance. This is the second-order version of "every seat owns a number": every number must own a reason.
Failure mode: Dan treated the scorecard as a plumbing problem (are tiles wired, do values push) instead of a strategy problem (does each KPI move the needle toward the goal). David: "we/I make all of these changes thinking you are looking at the big picture, that does not seem to be the case. Your job is to make sure that we reach our goal, and the KPIs should support that. How many emails Pepper reads does not move the needle. Each KPI needs to answer how it moves the needle, and how."
L10 prep must WALK THE ACTUAL MEETING before the meeting: open/fetch the exact meeting David will see (via API), verify every section renders with real data (scorecard snapshot has values, rocks board current across ALL teams incl. corporate, issues/todos loaded), run the needle-test on every tile, and FIX or stage fixes for everything found - all before 8am. The brief reports what was already repaired, not what will be discovered. Prep = simulate the meeting end-to-end; the meeting itself is only for decisions the human must make.
Why: A meeting that debugs itself live burns the scarcest resource (David's attention) on work an agent could have done at 7am. "The work happens between the meetings" is the entire operating philosophy of the meeting cadence - prep that only compiles data without verifying the meeting surfaces is half a prep. This is the root cause behind L059/L060/L061; fixing it structurally prevents all three recurring.
Failure mode: The 7/13 L10 scored 4/10. David's reason: prep ran in the morning but the meeting still spent most of its time discovering and fixing things live (weekend rocks missed, empty scorecard render, dead tiles, needle-less KPIs, corporate rocks invisible) - "we are fixing the meeting within the meeting with an absence of information. The work happens BETWEEN the meetings and this is not the case here."
When auditing whether an invariant holds across a codebase, verify it per STATEMENT, not per file. A per-file grep count is an aggregate, and aggregates hide the exact case you are hunting: the mixed file. Then encode the audit as a test that scans every call site, and mutation-test that scanner by reintroducing a known bug to confirm it actually fails. A scanner that only agrees with current code proves nothing. Related pattern seen the same day: when a data model gains a concept (rock levels, agent-owned KPIs), the model and the dashboard get wired up and the OTHER surfaces silently do not. Ask which surfaces read this table, not just which one is broken.
Why: The three holes the per-file count missed included the worst one: the blueprint serializer baked a private rock into a shareable template, which would carry it into another account. The failure mode of an aggregate check is a false clean bill of health, which is more dangerous than no check at all because it stops the search.
Failure mode: SUCCESS: Dan audited a privacy invariant (shadow rocks are owner-only) and found 7 holes, but the FIRST pass counted guards per FILE and missed 3 of them, because a file can contain one guarded query and one unguarded query and still look guarded in aggregate.
Invert it. Give the user ONE copy-paste block (MCP connection + a self-registration prompt) that they drop into Claude Code, Claude Desktop, ChatGPT, Cursor, or any MCP client. The AGENT then connects to OTP and registers ITSELF: it reads its own system prompt, calls a register/enroll MCP tool with its name, role, what it owns, what it does not own, and its KPIs, and OTP creates the seat and KPIs automatically. The user's only job is copy, paste, done. No forms, no parsing, no filling anything in.
Why: The agent already knows what it is -- making a human retype it is redundant work and a competence gate. Any flow that requires the user to know how to describe or configure their agent is not idiot-proof and will lose the non-technical user. The agent is the most reliable source of truth about itself, and it is already sitting on the other end of the MCP connection, so let it do the work. Design rule: when an AI is on the other side of the pipe, push the setup work to the AI, not the human.
Failure mode: Built OTP's "Connect an agent" flow as a human-driven form: the user pastes their CLAUDE.md into a textarea, OTP parses it, and the user hand-fills name / role / owns / does-not-own / KPIs before a seat is created. It made onboarding an existing agent the USER's clerical job and assumed the user knows how to describe their own agent.
Before any L10, dry-run tally.py and for every regex_in_file KPI check the source file's mtime against the KPI's time grain; if older than one grain, re-pull the number from the live system (Accelo, Search Atlas, Sheets) before pushing. New registry entries must use kind/regex/group (never type/pattern) and carry pending:true until their emit line exists. The unknown-kind branch now honors pending.
Why: A KPI pushed from a stale mirror is worse than a missing one: Crystal's tile would have said 32 when reality is 44, and the failure alert noise from mis-schema'd pending entries erodes trust in the one agent whose whole job is keeping the scorecard honest. Live-source-first is the same lesson as L042/L060 applied to Tally's own pipeline.
Failure mode: SUCCESS: Tally pre-L10 KPI sweep found and fixed three silent scorecard rot points: (1) registry entries added at the 7/13 L10 used type/pattern keys but the runner requires kind/regex, so their pending:true flag was never honored and they reported as failures; (2) the runner's unknown-source-kind branch ignored pending entirely; (3) two regex_in_file KPIs (Crystal 32, Beacon 0) were feeding from stale files (Jun 8 and Jun 22) while the live sources (Accelo: 44 projects; Search Atlas: 8 keywords tracked, not 43) had moved.
Any browser feature that accumulates unrecoverable state in page memory (MediaRecorder audio, unsent drafts) must (1) block/defer every programmatic self-reload while active, (2) be flushed by navigation-triggering handlers (End meeting awaits OTPAudioRecord.finish() before navigating), (3) guard beforeunload, (4) never drop data on a failed upload — keep the blob and offer retry. Long-term: stream chunks to the server (timeslice) so the page is never the only copy.
Why: A full Delta Meeting's recording/transcript was permanently lost — the audio never left the browser. Guard rails shipped in audio-record.ejs + l8-leadership.ejs; chunked streaming upload is the queued follow-up (touches billing-metered meeting-audio.ts, needs billing lock).
Failure mode: SUCCESS: Conatus — root-caused OTP meeting recording loss (2026-07-16): the browser recorder holds all audio in tab memory until Stop, while the live meeting page self-reloads on routine actions (reloadKeep, SSE scheduleReload, section-refresh fallbacks) and End Meeting navigates away — any of these silently killed a live MediaRecorder with zero warning.
Durable pattern for any browser feature holding unrecoverable state: (1) layer the fix — guard rails first (cheap, same-day, stops the active bleeding), durable streaming/persistence second; (2) keep the old one-shot path untouched as an automatic fallback so degradation can never be worse than before; (3) money-path parity — finalize calls the exact same precheck/ingest/charge sequence as the one-shot path, charge only after successful transcribe+ingest; (4) hold the billing lock across the whole build, release only after merge; (5) verify with a harness that runs the REAL shipped script (44 assertions) so old behavior is provably byte-identical with the new feature off; (6) cap every attacker-spinnable counter (segment count was a finalize-loop DoS lever).
Why: Completes L067's open loop: the "never lose a recording" guarantee is now structural, not procedural. The layered-fix + fallback-preserving pattern is reusable for every OTP feature that buffers user work in the browser (draft notes, offline edits), and the billing-parity discipline is how streaming touched the money path with zero semantic change. Both PRs merged and confirmed live on orgtp.com 2026-07-16 evening.
Failure mode: SUCCESS: Conatus — full resolution of the 2026-07-16 meeting-recording loss, shipped to prod same day in two layers: PR #205 guard rails (meeting page defers all self-reloads while recording; End Meeting flushes the recorder before navigating; beforeunload guard; pinned REC pill; failed uploads keep the blob with retry) and PR #206 streaming upload (recorder streams ~10s chunks to a server recording session; crash/reload loses ≤10s; resume banner stitches segments into one transcript; phone QR flow covered).
When invoking Steve Jobs as a design standard, treat him as holding BOTH axes to one bar: visual craft (typography, proportion, detail) and end-to-end experience (defaults, subtraction of steps, invisible mechanism). Never frame him as the "UX half" opposite a visual system; frame the visual system as one half of the single Jobs-level standard.
Why: David's design north star for OTP is the full Jobs standard. Splitting it wrongly would let screens pass a visual checklist while the flow, or the craft, gets held to a lower bar. The correct frame keeps one bar over both layers.
Failure mode: When framing the design-standard marriage (Fugu + Jobs), Conatus split it as "Fugu = visual craft, Jobs = journey/UX", understating that Jobs was also a master of visual design (Reed calligraphy class, typography on the original Mac interface).
For /coach-report and any Dash run: do MCP-dependent pulls (Search Atlas, Google Sheets, Calendar) in the MAIN session; delegate only file/CLI/analysis work to subagents. Check a subagent's tool access assumption before waiting on it. Treat CCM STL as unusable until the CloudCRM timezone offset is fixed platform-wide, and never report STL from July data.
Why: Two full delegation rounds were wasted waiting on pullers that could never succeed; the report would have shipped without SEO (repeat of L057) and without CCM if the main session had not redone the pulls. The STL corruption finding upgrades the known Villa-only issue (June) to system-wide, which changes every STL-based alert and coaching metric until fixed.
Failure mode: SUCCESS (with lesson): Dash /coach-report 2026-07-19 — delegated Search Atlas, CCM sheet, and calendar pulls to three subagents; SEO and CCM pullers were fully blocked because spawned subagents do NOT inherit the session's MCP servers (search-atlas and google-workspace tools were absent from their toolsets). Main session had both and pulled everything directly. Also: CCM July speed-to-lead timestamps are corrupt SYSTEM-WIDE (large negative timezone artifacts on every project, not just Villa Sport).
Two mechanics to remember: (1) .gitleaksignore fingerprints are commit-hash-bound, so ANY commit touching a line with secret-shaped placeholder text (like 'Bearer YOUR_API_KEY' in docs) re-mints the fingerprint and re-triggers the scanner — squash merges guarantee this recurs. The durable fix is neutralizing the placeholder so the rule can't match (angle brackets: 'Bearer <your-api-key>'), plus fingerprinting the immutable history. (2) CI checkouts with fetch-depth:0 fetch ALL refs and gitleaks scans all of them — one bad commit on an unmerged branch fails every branch's CI simultaneously; the ignore entry must reach each scanning checkout's .gitleaksignore, which means pushing it to the branch being scanned AND to main.
Why: Symptom (lint-and-type-check job failing everywhere at once) looks like a code regression but is actually the secret scanner; without knowing the two mechanics, the obvious fix (add one fingerprint) only patches one branch and the mole pops up on the next touch of the file.
Failure mode: SUCCESS: Conatus diagnosed a repo-wide CI outage caused by gitleaks fingerprint whack-a-mole — every branch's CI (including main pushes) went red at once from ONE unmerged branch's commit.
Before a swamp send, reconcile the changelog against the week's real PRs (git log origin/main --since since the last issue, filter to feat/ and customer-facing) and write entries for anything unlogged -- do not assume changelog.ts is complete. When the shared repo is contested by a concurrent session, do all changelog authoring in an ISOLATED git worktree (git worktree add off origin/main, symlink node_modules to reuse deps), PR it, merge via gh after CI is green, and run the REAL send from the worktree -- never edit or send from the shared working tree. Date drop-wave entries to the issue's MONDAY anchor (the sender's window upper bound), NOT to the send day: entries dated the Tuesday send day render as future and hide, and getRecentEntries (OS-today) masks this in preflight.
Why: The changelog is the single source of truth for both /whats-new and the email; an unmaintained changelog silently undersells the product to every subscriber. On a machine with concurrent agent sessions, the shared working tree is not safe for a multi-step author+send; a worktree makes the work deterministic and collision-proof. The Monday-anchor date rule is invisible until a dated-Tuesday entry silently disappears from the send.
Failure mode: SUCCESS: Swamp #29 -- two non-obvious operational wins. (1) The digest only reflects changelog.ts, so a big shipping week (~14 features) went out as "2 things" because most PRs never got changelog entries; David caught it. (2) Running the swamp send while another session was actively branch-switching in ~/otp-platform stranded an early commit on a feature branch and made a direct push to protected main a no-op.
For any EJS page whose logic lives in an inline script, add a test that parses every inline script with new Function() and fails the build on a syntax error, then prove the tripwire by running it against the broken version before keeping it. tsc cannot see inside a template and EJS renders a broken string happily.
Why: It was caught only by driving the actual rendered page in a headless browser, not by tsc, lint, or 1204 passing tests. First-run screens are where paying customers land, and a dead one is invisible from the server side.
Failure mode: SUCCESS: Conatus found OTP onboarding Door 4 (Give Ollie everything, PR #232, shipped 2026-07-19) had been completely dead in production since Sunday. onboarding-import.ejs line 214 had an apostrophe inside a single-quoted JS string, so the browser discarded the whole inline script and the file drop, analyze and commit buttons did nothing, silently, for every new customer who picked that door.
David: A2P pages do not allow form fills. Remove the lead form from A2P review landing pages entirely. Keep the SMS disclosure, the Privacy Policy and Terms links, and the registered business name and address on the page, because those are what carriers actually read. Route the quote path to phone plus the GHL chat widget. Reword any disclosure copy referencing "check the box above" so the page does not describe a mechanism that no longer exists.
Why: The form was the architecture of these pages, so this is not cosmetic: it orphans the consent audit trail and the POST /:slug/lead route, and it changes what the A2P campaign registration can declare as its opt-in method. Getting it wrong in either direction risks 10DLC rejection, which blocks SMS for the client entirely. The pattern repeats for every future A2P location page.
Failure mode: Built the Dryer Vent Squad A2P landing pages (Katy, DFW) around a web lead form with an optional SMS consent checkbox, treating that form as the opt-in proof mechanism for A2P/10DLC review, plus a consent audit trail behind it (consent.js and sms_consent_text/timestamp/IP/user-agent captured on POST /:slug/lead).
The unbilled-spend sweep (billing-report step 3b) reads ~/.claude/billing/sweep-exclusions.json and drops those account IDs from the Review tab entirely. Use it for accounts that spend on our Google/Meta but are NOT billed on % of ad spend (white-label / flat-fee). Added 2026-07-23 per David: Phillip Jeffries, M.V. Parker Law, Champy's Chicken (+Nashville), Emily Shalant, Jet City Blinds, J&K Engines, Meyer Law, True Path, GettaMeeting, Lazzara Law, Studstill Firm. Only add IDs here on David's explicit instruction. Separate from the Clients-tab dont_bill mode (which still shows a DO NOT BILL row on the Billing output).
Why: Without a persistent list these 12 accounts (~$24.4K/mo) resurface as "confirm arrangement" in every monthly sweep, wasting David's review time. The file makes the exclusion durable across sessions.
Failure mode: SUCCESS: Billing sweep now has a persistent white-label exclusion list
Three patterns for the OTP recorder. (1) "No way to record a second time" is TWO bugs: the missing UI affordance AND a server ingest that overwrites (ingestTranscriptOneShot did set({transcript})). Fix both or the feature silently destroys data; recording ingests now pass mode:'append', idempotent on the tail so worker retries cannot double-append, while paste/import stays 'replace'. (2) "Mic recorded silence after a long pause" on a phone is a dead MediaStreamTrack (readyState 'ended' or muted): MediaRecorder.resume() succeeds and records nothing. Recover by swapping in a fresh getUserMedia stream, and ALWAYS flush the retired recorder's final chunk BEFORE bumping the server segment number, or the old container's tail bytes land in the new segment and corrupt it. Auto-recover only when the track is provably dead; when it is alive but silent, warn and offer a button, since a genuinely quiet room looks identical. (3) Check git log before rebuilding from a support ticket: background transcription had already shipped that morning, so ticket 3 needed a sequential-to-concurrent R2 upload fix plus a crash fix, not a rebuild.
Why: The recorder holds the customer's only copy of a meeting until upload completes, so every bug in it costs unrecoverable audio, and two shipped in one day. The widget is inline JS in an EJS partial with no import path, which is why it went untested; it can now be tested by rendering the partial, extracting the script, and running it in a vm against a fake MediaRecorder (src/views/partials/meeting/audio-record.test.ts). Use that harness for future recorder changes.
Failure mode: SUCCESS: Claude (OTP dev) fixed three meeting-recording tickets (PR #315) at root cause, and caught that PR #309 from earlier the same day left savedStatus() calling itself on the non-pending branch (a stack overflow that swallowed the save confirmation on the R2-off path) because the recorder widget's inline JS had no tests.
When editing OTP trust/security claims, edit src/config/trust.ts (the file the /trust route imports and ships in the dist image). trust.yaml at the repo root is only the audit copy carrying `# source:` code citations for legal; nothing reads it at runtime. The two had already drifted (legalEntity OTP,LLC vs OrgTP,LLC; a stale lastUpdate; and — critically — trust.ts shipped a prohibited EOS mark "L10" that trust.yaml did not). Always mirror any claim change into BOTH files, and treat trust.ts as authoritative for what the public actually sees.
Why: A trademark-compliance violation (EOS "L10" mark) was live on the lawyer-facing trust page for weeks because the de-EOS pass only fixed trust.yaml, which never ships. Editing the audit copy feels like fixing the page but changes nothing a visitor sees. This is a recurring drift trap worth a permanent CI equivalence check between the two files.
Failure mode: SUCCESS: the public /trust page renders src/config/trust.ts, NOT trust.yaml — the "source of truth" file never loads at runtime
Never position OTP (or Ollie) as an employee. A founder, and even an employee, does not want another employee; they want to know the organization is held. "Employee" also imports employee mentality: waits to be told, owns a lane not the whole. OTP guides the organization. Frame it as the guide/what holds the company, positioned between employee (too small), guru (too big), and co-founder (not really).
Why: Positioning error at the core-promise level: hiring an employee increases a founder's load (managing, explaining, checking); OTP's promise must reduce the holding. Wrong noun poisons every downstream page, price, and demo.
Failure mode: Framed OTP's destination as "the first employee a company hires whose job is to remember" in the product discovery brief.
When a vendor was chosen specifically to absorb an operational burden, do not propose solutions that hand that burden back, even when the vendor documents them. Default order for the consent-screen concern: (1) own the story in product copy (tell users they will see Composio, frame it as the security vault), (2) at most white-label the two or three marquee providers if branding ever matters commercially, (3) never the whole catalog.
Why: Vendor docs happily describe features that shift work back onto the customer. The right frame is the original build-vs-buy decision: Composio IS the OAuth team. Recommendations that quietly re-hire that job in-house waste the subscription and David's time.
Failure mode: Asked how to get OTP branding on the Composio OAuth consent screen, Claude recommended white-labeling via custom auth configs, which means creating and maintaining an OTP-owned OAuth app per provider (Slack app review, Google verification, secret rotation, scope upkeep). David rejected it: the reason OTP uses Composio at all is ONE managed OAuth surface across hundreds of products; a per-provider OAuth app pipeline recreates the exact burden Composio was chosen to eliminate.
The company boundary applies to every artifact that feeds Ollie, not just what is said in the room. When writing a meeting record for the Sneeze It L10, describe mechanisms in company-neutral operational terms (the meeting pipeline, the prep gate, the record path) and keep OTP product identifiers (PR numbers, endpoints, feature ship dates) out of the record entirely. Before pushing any record via agent-record, scan it with the same company-mismatch check the preflight applies to the board.
Why: Ollie's insight renders inside the Sneeze It meeting, so a record contaminated with OTP content produces a contaminated insight automatically, one week later, with no human in the loop. The boundary check must move upstream to where the source is authored or the violation recurs on autopilot.
Failure mode: Dan wrote the 7/20 agent-record for the Sneeze It L10 full of OTP product identifiers (PR #154, the agent-record endpoint, ship dates), so the Ollie Insight generated from it reads as OTP product narrative inside the Sneeze It meeting. David caught it live on 7/27: "Ollie still thinks OTP and Sneeze It are one." The context-bleed boundary was enforced in live speech but not at record-writing time, and the record is the insight's source.
Prep does not end at the Slack brief. Any signal the brief nominates for IDS gets pushed to the OTP board as a ticket (otp-issue.sh, with teamId, AI Army = 065d1d4b) BEFORE the meeting, so David opens the meeting with the issues already loaded and workable in-product. The Slack brief is the narrative; the board is the working surface.
Why: The meeting runs inside OTP, so an issue that exists only in Slack is invisible at the moment of solving. This is the same source-of-truth lesson as L060 applied to prep outputs, not just prep inputs.
Failure mode: Dan surfaced pre-meeting signals (Pulse dark, pipeline shape, HiTone, overdue queue items) only in the Slack prep brief. David at the 7/27 L10: "you should have wrote those signals in OTP, too late now." Signals that deserve IDS never landed as tickets on the AI Army board, so the meeting could not work them in-product.
When David closes an alert as handled, retire the RULE that generates it, not just the instance. For HiTone: edit the CLAUDE.md trigger line and coach-report spec to read "billing confirmed active since Jul 2026, do not flag on spend." In general: trace any recurring alert to the config line that emits it and fix that line, or the alert regenerates forever.
Why: A closed instance with a live trigger is an alert factory. It spends David's attention on the same resolved question weekly and erodes trust in real billing flags.
Failure mode: HiTone billing keeps re-surfacing to David even though he confirmed billing is correct and active (Jun 29, and again 7/27: "HiTone is being billed, you ask about that a lot"). Root cause: the BILLING TRIGGER rule still lives in CLAUDE.md's active-clients list and in the coach-report spec, so any agent that reads the config re-fires the flag whenever HiTone spend appears. The Jun 29 closure was recorded in the rocks file, but the upstream trigger rule was never retired.
Prep must scan for WON/signed revenue explicitly, not just open pipeline: GHL won opportunities across ALL pipelines since the last meeting, plus Proposify signed events, categorized new/expansion/reactivation. Signed expansion revenue is the Q3 headline metric; it leads the scorecard section, and its dollar values get pushed to the manual Expansion tile the same morning. A tile with no automated source still gets its value entered at prep time from the won-deal scan; manual source does not mean no value.
Why: The board exists to catch exactly this: revenue proving or disproving the quarterly thesis. A prep that inventories dead tiles but misses live signed money reports the plumbing and skips the water. David finding revenue wins that his facilitator missed inverts the entire point of the seat.
Failure mode: Prep missed two signed expansion deals (Glo30 and WOA franchise additional revenue) that David saw on the board himself and had to point out at the 7/27 meeting: "you skip that a lot, did you not see them?" The prep scan read open opportunities in one GHL pipeline and the Expansion KPI tile (manual, no values), so signed/won expansion revenue had no path into the brief. The single most strategy-relevant signal of the quarter, expansion revenue from existing accounts, was invisible to prep while five dead tiles got named in detail.
Same-session capture rule for ALL agents: when you witness David build or ship something real (a dashboard, an integration, a signed deal, a process), record it THAT session: a headline line in the daily note, and a todo or KPI value in OTP if money or a rock is touched. Do not wait for the weekly meeting; the builder remembering to report is not a capture mechanism. Dan additionally runs a Shipped This Week sweep every Monday prep as the backstop.
Why: Work done outside meetings is systematically invisible to a meeting-based OS, and the founder's most valuable hours happen outside meetings. Three instances surfaced in one meeting (dashboard, prep signals, signed revenue). Invisible work costs real money: David unknowingly built part of Bogdan's open reporting-cost rock.
Failure mode: David built a WOA dashboard for iCart (central database, less Zapier, better reporting) and the only witness was the AI in that session; no headline, todo, or record reached the operating system. David at the 7/27 meeting: "the only one that knows is you, this is a true failure of the operating system that needs to be reconciled." Same morning, two signed expansion deals (Glo30, WOA franchise) were also absent from every system of record.
Facilitation = operating the product live. The moment a section starts, its artifact moves: an issue under discussion is verified rendering on the board before discussing it; the moment David decides, the ticket is solved with its resolution, the todo is created, the KPI is pushed, in that minute, not at conclude. After every state change, verify the rendered surface. Conclude should be a read-back of changes already made, never a batch of pending writes.
Why: A meeting inside OTP is only real if the product state changes while the humans watch. Deferred writes recreate the mirror-drift problem inside a single meeting, and David cannot trust a board that lags the conversation.
Failure mode: Dan facilitated IDS discussion in chat but did not move the meeting's product surfaces in real time: the Pulse issue being discussed was not visible in the meeting's IDS section, and the already-solved invisible-work issue still sat open on the board. David 7/27: "you should be moving the meeting along like a human, doing things as we work."
The needle test has a step zero: before wiring, fixing, or reporting any KPI, confirm the program/offer it measures still exists in the business, by asking David or checking recent revenue/activity, not config files. A dead program's tile is retired, not wired, and its language gets swept from all agent config at source (L099 pattern) so no agent rebuilds it. Guarantee/T20 program: DEAD as of 2026-07-27.
Why: Config outlives strategy. Wiring effort spent on a dead program's metric is worse than a dead tile, it would have shipped a number that misrepresents the business as still running an offer it killed, and every agent reading the tile would have inherited the fiction.
Failure mode: Dan was one step from wiring the guarantee-clients-retained KPI (emit line written, Tally about to fire) when David said the guarantee program is dead and no longer offered. The wiring work treated the tile's PROGRAM as alive because the config said so; nobody had asked whether the business still runs the thing the tile measures.
When a media element silently stalls (networkState LOADING, error null, no console output), suspect CSP: a 302 redirect from a same-origin playback route to a cross-origin storage URL violates media-src (falling back to default-src 'self') and Chrome blocks it with zero surfaced errors. Diagnose live by attaching a securitypolicyviolation listener and probing a known cross-origin media URL. Fix: add the presigned bucket origin to media-src, computed from storage config at boot, never hardcoded.
Why: This failure mode is invisible by design (no element error, no console noise) and the natural debugging paths (storage probes, presigned URL tests, range requests) all pass, sending you everywhere except the CSP header. Cost about an hour across two sessions; the probe technique turns it into a 2-minute check.
Failure mode: SUCCESS: Claude/Conatus diagnosed why OTP meeting recordings never played in the browser (player stuck at 0:00, click did nothing) while server-side storage checks all passed.
Never invent a person's first name from an email address or initial. If the name is not stated, refer to them by the email address or ask David, and only record a name once confirmed.
Why: Guessed names propagate into client-facing artifacts (emails, dashboards, docs) and getting a client's name wrong damages trust; an email initial is not evidence of a name.
Failure mode: Claude inferred the Drybar Ballston client's first name as "Jodi" from the email address jsterling@sterlingcapitalllc.com and used it in the summary, credentials file, and memory. The client is Julie Sterling.
`ghl.sh update-opp <oppId> <stage>` with the optional value argument omitted hits an inline Python syntax error (`data['monetaryValue'] = ` with nothing after it), so the stage payload is never built: the opp's updatedAt bumps but the stage does NOT change, and nothing errors loudly. Workaround until ghl.sh is fixed: always pass the value explicitly, e.g. `update-opp <id> sql 0` (confirm the opp's current monetaryValue first so you do not overwrite a real value). Always verify stage changes by re-reading the opportunity after the write.
Why: A stage move that silently no-ops corrupts the pipeline of record without any error signal. Both /otp-sales and /sneeze-sales route stage moves through this command; without the verify-after-write habit the bug would have shipped 2 phantom SQL bumps today.
Failure mode: SUCCESS: Sneeze-Sales found and worked around a silent ghl.sh write failure
Before bumping a stage or queueing a task on any reply, check the PERSON and company against the active client list (CLAUDE.md), pepper-clients.md, and known client people, not just the cold-load exclusion at contact creation. Jordan Anderson is a client (Workout Anytime / Proof Fitness): never prospect-touch him on any domain. Replies inside threads a team member already owns (for example Zeynep scheduling) get no David task; the owner handles it. A reply landing in the prospect book is a signal to verify WHO it is, not proof they are a prospect.
Why: The prospects-only law fails at the edges: client people reply from domains that sit in the prospect book because franchisee outreach and client domains overlap (WOA). A wrong SQL bump plus a David task double-touches a client relationship and burns David's queue on non-sales work.
Failure mode: Sneeze-Sales treated Jordan Anderson (Workout Anytime) as a prospect: bumped his opp MQL to SQL on an inbound reply and queued David a respond-within-24h task. David corrected: Jordan Anderson is a client. Also queued David a confirm task on Lindsey (Fitness Factory) when Zeynep already owns that thread.
When an OAuth integration fails silently, verify each layer with direct probes instead of reasoning from app behavior: (1) print env var names with cat -v, since a pasted quote becomes part of the variable NAME and the app reads undefined (Railway kv showed "GOOGLE_CALENDAR_OAUTH_CLIENT_ID with a literal leading quote); (2) test client credentials against the provider token endpoint with a bogus auth code, since the error distinguishes exactly: invalid_client means bad id/secret, invalid_grant Malformed auth code means credentials are VALID; (3) treat the database as ground truth for whether a flow completed, since users saying connected can mean a different surface (Composio integrations page vs the calendar card).
Why: Three probe layers turned what could have been hours of guessing into minutes: no callback in logs proved the flow died at Google, cat -v exposed the quote character, and the bogus-code token probe verified the replacement secret BEFORE the user retried, avoiding another failed round trip. Reusable for every OAuth integration OTP adds (Microsoft, Zoom, future providers).
Failure mode: SUCCESS: Claude shipped Recall calendar auto-join and debugged two invisible OAuth config failures the same afternoon (quoted env var name, invalid client secret)
When graduating a feature out of Labs, grep the whole repo for the feature key and isFeatureEnabledForOrg calls before merging; every gate (page, API, scheduler, MCP tools) must come off in the same PR. A fail-closed flag check on a deleted key silently disables the feature for everyone, which reads as random 404s, not as a flag problem.
Why: Fail-closed gating is correct security posture, but it means catalog removal IS a kill switch. The bug shipped invisible because the page worked while the API did not, and the generic 404 body hid the cause; the honest error message shipped hours earlier (names the workspace and feature) is what made the real diagnosis possible.
Failure mode: Projects went GA (PR #347 removed it from the Labs catalog and the page gate) but the API routes kept gating on the removed key; isFeatureEnabledForOrg fails closed on unknown keys, so all project CRUD 404'd platform-wide for two days until Dawson and David hit it
Any otp-platform endpoint that scopes data or authz by getAuth(request).userId is broken under impersonation. The rule: gate and scope by the EFFECTIVE viewer (request.impersonation.as when active, else auth.userId), return/audit with the RAW session id, and for context-pinning actions (org switch) re-issue the impersonation cookie via startImpersonation rather than moving the admin's own cookies. When reviewing or writing any new route, grep for getAuth(request).userId used in a WHERE clause — each one is a latent impersonation bug.
Why: Impersonation is how David supports customers (view-as Tom, Kristen, etc.). Every raw-session usage silently shows the admin's data under the customer's banner or 403s the customer's own surfaces — a privacy leak in one direction and a support dead-end in the other. Fixed instances: dashboard (2026-06-02), PRs #381, #382, #384 (2026-07-28).
Failure mode: SUCCESS: Dan identified a recurring defect class in otp-platform — four separate surfaces broke under super-admin impersonation in one day (portfolio pages listing the admin's portfolios, portfolio API 403ing "Could not load team", the sidebar org label showing the admin's org, and the org switcher 403ing "Could not switch organization"), all with the same root cause.
Do not treat a missing or small wallet balance as a signal of anything. An org with no wallet is simply not an active OTP user, so its balance says nothing about product readiness. New orgs are seeded with $25 of credit to incentivise starting, so a funded wallet is the default going forward rather than a hurdle. When assessing whether a metered feature is usable, filter to orgs with real activity and check the NON-wallet prerequisites, since those are the ones that actually gate anyone.
Why: Reporting wallet balances as blockers manufactures work out of the ordinary shape of the user base: most rows in that table are dormant signups, not stuck customers. It also buries the prerequisite that does bite, because a real blocker listed next to four fake ones reads as one item in a list instead of the single thing to fix.
Failure mode: Flagged orgs with low or missing wallet balances as a readiness problem for OTP scheduling, treating wallet funding as a live blocker worth David's attention.
Before retiring a scarcity, cohort or badge claim, count each cohort in the database and check whether one phrase names several cohorts. Shared wording is not shared meaning. If counts contradict the instruction's premise, return to the human with the numbers instead of executing literally.
Why: Broad approval rests on an assumed premise. When the premise is partly false, literal execution silently destroys value: a deleted live offer throws no error, it just yields fewer signups. One query per cohort is cheaper than a loss nobody detects.
Failure mode: SUCCESS: Claude verified cohort counts before executing an approved codebase-wide sweep of OTP's "first 50" claim. It was closed for signups (50/50) but live for Founding Publishers (45/50) and Founding Partners (6/50). Literal execution would have deleted two accurate live offers.
For multi-agent feature builds: (1) research agents return structured briefs before any code, and briefs override the spec when they conflict (two detectors were impossible as specced: meetings have no booked-duration column, decisions are not rows). (2) Put all shared-file wiring (server.ts) in ONE sequential final task so parallel agents never collide. (3) Always run an independent fresh-context diff review before committing: it caught two honesty blockers the per-task verifications missed (ratified moves netting costs away; org-wide gains summed over subtree-scoped costs). (4) Discovery worth acting on: subscriptions.plan_rate is never written by any code path in otp-platform, so any revenue/cost feature reading it ships dark until billing populates it.
Why: The per-task agents were all green individually; only the cross-seam review found the invariant violations. Repo guard tests (private-issue leak scan, blueprint coverage) also fired exactly as designed, proving lint-style guard tests catch what unit tests cannot.
Failure mode: SUCCESS: Claude shipped OTP Impact Phase 1 (PR #427) via 11-agent build: parallel research briefs, wave execution with pure-function cores, then adversarial diff review before commit
When adding a NEW page to an existing app, open a sibling page that already ships (for OTP admin surfaces, /admin/support) and copy its outer container, top padding and control classes verbatim before writing any markup. Do not hand-roll spacing from DESIGN.md tokens alone -- the tokens do not tell you the page-level offsets that keep content clear of the fixed header. For any default that is a money amount, confirm the number rather than inferring it from the option list order.
Why: A new page laid out from first principles looks subtly wrong in ways the author cannot see without loading it: header collisions and control scale only show up in a browser, not in typecheck, lint or design-lint, all of which passed. Copying a shipped sibling inherits every page-level decision already made and reviewed.
Failure mode: Built /admin/join-link with hand-rolled Tailwind layout (max-w-3xl, custom padding, custom input classes) instead of copying the container and control classes from an existing admin page. Result: the page header collided with the fixed top nav so the title was unreadable, and the form controls were oversized versus OTP's 32px control scale. Also picked $50 as the default starting credit without asking; David wants $25.
When promoting any OTP Labs feature from beta to live, do THREE things, not one: (1) flip `status` in src/shared/lab-features.ts; (2) grep src/views for the feature's `surfaceUrl` -- if the ONLY link is the Labs-injected rail item, add a permanent entry to layouts/main.ejs in BOTH the `_sbItems` array and the mobile settings menu; (3) grep the page for stale "this is a Labs feature" banner copy pointing at a /settings/labs toggle that graduation just removed. Also verify any second, independent gate (e.g. an env check like recall-calendar.ts calendarIntegrationEnabled) and make the registry copy match what is actually configured in production -- check `railway variables --kv` rather than trusting the existing description.
Why: Graduation looks like a one-line status change and is not. The rail item, the page's own Labs banner, and any env-based second gate all key off the old state, so a naive flip can make a feature LESS reachable than it was in beta while appearing to ship it. Confirmed live: PR #437 shipped the flag plus the nav entry together, and the promoted page rendered correctly with the Calendar section visible and no Labs opt-in.
Failure mode: SUCCESS: Claude caught that graduating an OTP Labs feature from `beta` to `live` silently DELETES its left-rail nav item, which would have shipped calendar auto-join into being unreachable. `getOrgLabNavItems` (src/services/lab-features.ts) filters on `f.status === 'beta'`, so only beta features get a rail item injected. /settings/meeting-presence had no other link anywhere in src/views, so flipping the flag alone would have removed the only way to navigate to it.
Four reusable rules for /coach-report and any Dash run. (1) When `meta-ads.sh token-check` returns Valid:False, do not report the portfolio as quiet: the CCM Stats sheet "Ad Spend" column carries the Meta-side spend for every call-centre project, so it is a working fallback for spend and leads. Label the source on the card. (2) `mcp__google-workspace__get_events` silently caps at max_results and truncates the NEWEST events, not the oldest. A 90-day pull capped at 250 returned nothing after Jun 30 and would have reported "no client meetings in July." Always check the max date in the response against time_max and re-pull in narrower windows. (3) Never average Google cost-per-conversion together with CCM cost-per-booked-appointment in a franchise network benchmark. Group the cost metric by source, keep only the largest comparable group, suppress the table when fewer than 2 comparable peers, and name the metric explicitly on the card. (4) When CCM shows appointments booked but zero shows, that is unconfirmed data, not a zero show rate. Do not report show rate; ask for confirmation instead.
Why: Each of these silently produces a confident wrong number in a client-facing artifact. The calendar cap fabricates a churn signal, the mixed-metric average makes a healthy location look 60x worse than a peer, and a missing Meta token makes a $136K/month portfolio look dead. Cards go to clients, so a wrong number costs trust directly.
Failure mode: SUCCESS: Dash /coach-report 2026-08-03 — shipped 50 cards with Meta Ads fully down, and caught two silent data traps that would have produced wrong client-facing numbers.
Refines the CCM-fallback rule captured in L125. The CCM "Ad Spend" column matches Meta actuals closely (Rockstars Frisco $453.23 vs $452.54, Okeechobee $150.07 vs $151.09, China Grove $384.79 vs $388.21) but ONLY when the sheet has a complete row for every day in the window. WOA Winder read $133.98 against Meta's $303.08 because the project stopped appearing in the sheet after Jul 30. So: before using CCM as a Meta proxy, count the daily rows per project across the window and flag any project with fewer rows than days. A project that silently drops out of the sheet reads as a spend and lead collapse when it is a recording gap. Second lesson: do not attribute a lead decline to an ad-platform outage without checking delivery. Meta REPORTING was dark to us from Jul 27, but Meta DELIVERY was fine (portfolio leads only -9% week over week, spend flat), so the much larger per-project CCM declines were a recording or routing artefact, not an ad problem.
Why: The fallback is genuinely good enough to save a report, but only with the completeness check. Without it, a project that falls out of the sheet produces a fabricated churn signal that a coach would take to a client. And blaming a visible outage for an invisible decline is the easy wrong answer that stops the real investigation.
Failure mode: SUCCESS: Dash — Meta token restored 2026-08-03, and the CCM fallback used during the outage was measured against real Meta data once it came back. The fallback was accurate to within 1% on 3 of 4 spot-checked projects but understated WOA Winder by 56%.
Upload the coach report to Drive exactly ONCE per run, at the very end, after the validation sweep passes. Never upload an intermediate build, even when the intent is to re-upload a better one later, because the shared folder is read by coaches the moment a file lands. If a rebuild is genuinely needed after an upload, trash the superseded file in the same action rather than leaving both. Use update_drive_file with trashed=true (recoverable); neither Drive MCP exposes a hard delete. Verify by file size against what was generated locally before trashing anything, and get David's approval first since the folder is shared.
Why: A shared folder is a publishing channel, not a working directory. Every extra file is an opportunity for a coach to open the wrong numbers and take them to a client. The same duplicate pattern already exists on 2026-07-11, 2026-05-25, 2026-04-06 and 2026-03-10, so this is a recurring habit rather than a one-off.
Failure mode: Dash uploaded three files to the shared Coach Reports Drive folder in one morning, as the data improved from Meta-blind to full to corrected. Zeynep has reader access, so two of the three were coaches' paths to stale client numbers until David approved trashing them.
Never read a green tile as evidence that the agent named on it is alive. Decoupling a KPI from its agent protects the number but removes the number's ability to report the agent's death, so agent liveness needs its own signal: check the shared-state file mtime alongside the tile before making any seat decision. When a seat review comes up, present tile value and agent liveness as two separate lines, never one.
Why: A seat can look healthy and be vacant. Here the 32.6% measured Erica and Amanda's human performance, not agent output, so the tile would have argued against repurposing a seat that had already stopped running. Seat decisions made on decoupled tiles are made on the wrong evidence, and the error is invisible because the metric is genuinely accurate about the thing it actually measures.
Failure mode: Dan argued in the 8/3 meeting that Arin's 32.6% appointment rate showed the Arin seat was working, so repurposing it was not urgent. That inference was wrong. The Arin KPI was deliberately decoupled from Arin-the-agent in June (source kind composio_action, reading the CCM sheet directly) so it would survive a repurpose. Meanwhile arin-latest.md had been stale for 287 hours, about 12 days. The agent was dark and its tile was green the entire time.
Before escalating any data problem to a vendor, query the vendor's live API for the object in question and reconcile it against our own source of truth. Here the API showed the project had held exactly 8 keywords since creation (first_keyword_added_at), so nothing was ever lost, and our own cluster map defined 28 pillars rather than the 43 the KPI asserted. The genuine vendor issue turned out to be a different and much larger one: 19 of 27 projects silently blocked on NO_QUOTA, including six client projects. Split the ticket: send the vendor only what the vendor actually owns.
Why: A false claim to a vendor costs credibility and burns the support cycle you need for the real issue. It also hides the internal defect behind a vendor excuse, so it never gets fixed. In this case verifying first both protected the vendor relationship and surfaced a client-delivery problem nobody had noticed.
Failure mode: An item routed to a vendor (Search Atlas) as "the vendor dropped our 43 tracked keywords to 8" was accepted at face value from an L10 without checking live vendor state first. Sending it would have been a false data-loss claim against the vendor.
When a measurement pipeline is broken, check whether the metric itself is the defect before repairing the plumbing. Beacon's KPI was "pillar keywords in top 10 (of 43)" on a domain whose pillar pages had launched six weeks earlier. That number reads 0 for quarters regardless of whether the work is excellent or abandoned, so repairing the vendor feed would have restored a tile that still said nothing. Test any KPI with two questions: can this number move within one review cycle, and if it moved would I trust what it means? Beacon failed both (its actual top-10 rankings were unrelated junk queries). The fix was to change the instrument to Google Search Console, which is free and already authorized, and the metric to pillar-cluster impressions, which moves weekly. Rename the existing tile in place via PATCH /api/v1/kpis/{id} rather than creating a new one, so the seat keeps its position and no dead tile lingers on the chart.
Why: A KPI that cannot move is not accountability, it is decoration, and it quietly consumes the attention a real metric would earn. This one also hid a genuine finding for six weeks: the pages were being surfaced 3,479 times and earning one click, which is a titles and intent problem no top-10 counter would ever reveal. For OTP specifically this is a Constitution matter, since the axiom is that an org reconciles what it says with what it does; running a fake KPI on OTP's own chart dogfoods the disease the product exists to cure.
Failure mode: SUCCESS: Beacon rewired from a vendor-dependent KPI that had never once produced a real number to a free first-party one that works.
After editing any .ejs or CSS in otp-platform, run `node scripts/design-lint.mjs` as well as tsc/tests/smoke:render. It has no npm script, so the standard local verify passes while CI's lint-and-type-check job fails. It counts DESIGN.md violations per file against scripts/design-lint-baseline.json and fails on any increase. When it fires, fold the new selector into the existing rule instead of running --update on the baseline.
Why: A view change can look fully verified locally and still red CI, costing a full round trip. On PR #459 a focus ring (not a resting shadow) tripped css-resting-shadow 6 -> 7; folding the picker into the existing input and focus rules fixed it and was what DESIGN.md wanted anyway, so the gate caught real design debt rather than noise.
Failure mode: SUCCESS: Claude found the otp-platform verify recipe is incomplete for any view change
A bounce is never proof an account is fake. Before deleting any account, check three things: (1) does an auth-provider user exist, (2) when did they last sign in, (3) what data and pending invites hang off them. Then read the bounce SHAPE: a hard bounce in ~3 seconds means the domain or mailbox does not resolve, which usually points to a typo in an otherwise real address; a bounce 12-14 hours after send is a soft bounce from retry exhaustion (full mailbox, suspended account, reputation deferral) and the person is real. Report the split and get explicit confirmation before deleting anything that is not provably fake.
Why: Bounces cluster on real customers with mistyped addresses, not on fabricated signups. A bad email is a recoverable lead: One Jump's owner mistyped a subdomain and became unreachable for seven weeks, while the colleague he invited was reachable at the correct corporate domain the whole time. Treating "bounced" as "fake" deletes paying-customer-shaped signups and destroys the only evidence needed to win them back. Deletion is irreversible; suppression achieves the actual goal (stop the bouncing) at zero cost.
Failure mode: SUCCESS: Claude caught that a "delete these fake bounced accounts" request included a live customer org. Of three addresses flagged as fake, only one was: test@abctest.com had no Clerk user at all. business@mail.onejumpinc.com was Jessup Jong, owner of the live "One Jump" org, last signed in 3 weeks earlier, with a chart, team, meeting, and a pending invite to a colleague expiring in 5 days. Deleting as asked would have destroyed a real signup mid-onboarding, irreversibly.
Never write a file that other processes read with open(path,'w') plus write. Write to a temp file in the SAME directory via tempfile.mkstemp, then os.replace(), which is atomic on POSIX so readers see the old file or the new one and never a half-written one. Wrap it so the temp file is unlinked if the write throws. When you find one writer of a shared file doing this, grep for every other writer of the same file and fix them all; fixing one leaves the race intact. Verify with a concurrency test that reproduces the failure on the old code and shows zero on the new, rather than assuming. Note that mkstemp yields 0600, which tightens a credential file from the usual 0644 and is an improvement.
Why: A truncate-then-write race on a shared credential file is invisible in normal use and only appears under concurrency, so it presents as a flaky, unreproducible alarm on a load-bearing data source. That trains the operator to dismiss the monitoring. Worse, the failure mode is indistinguishable from a genuinely revoked token, so the real alarm and the false one look identical. Atomic replace removes the entire class of failure at the source rather than papering over it with retries.
Failure mode: SUCCESS: Radar traced an intermittent phantom "Google Ads token invalid" alarm to a non-atomic credential write, not to the token or the network. Three separate processes write ~/.claude/mcp-google-ads/google_ads_token.json (google-ads.sh, google_ads_server.py, auth_setup.py) and all three used open(path,'w') followed by write. That truncates the file first, so any concurrent reader json.load()s a partial document and throws. Because get_access_token ends in 2>/dev/null, the exception surfaced as an empty token and got reported as a dead credential. A hammer test measured 178 partial reads out of 400 writes, a 44% failure rate under contention.
Before repainting any utility class in a view, grep it in src/styles for an !important clamp selector and move that hook in the same commit. Verify repaints with a pixel diff against a before-screenshot, never on a green lint run alone.
Why: A clamp keyed on a class name is an invisible coupling no test or linter can see, and cleaning up that class name is exactly the change that breaks it.
Failure mode: SUCCESS: Claude found OTP dashboard-daily hairline-row grammar was produced by an !important clamp in input.css keyed on the row wash class name. Repainting that class onto tokens silently detached every row from the clamp, restoring borders and radius and shifting padding. Tests and design-lint both stayed green.
For the Sneeze It AGENCY lane, outreach is account-based, not volume-based. The outreach that produced actual revenue (Beem, now a multi-location client) was individually researched and drafted into David's Gmail, one account at a time. Build the big list for SELECTION, not for sending: a master Google Sheet where every row carries enough context to be decidable in ten seconds, David marks an X on the rows he approves, and each approved account then gets real research and a bespoke Gmail draft. GHL becomes the record of the resulting deal, not the send engine. Do not templatize, do not sequence, do not meter.
Why: Sneeze It accounts are worth $44k to $136k a year each (HiTone is 8 locations at $5,472/mo; WOA is 13 club accounts at ~$136K/yr expansion). At that account value, thirty minutes of research per email is trivially correct economics, and 500 templated sends optimize the wrong variable entirely. The data agrees: 221 contacts sit parked at MQL against 11 that ever advanced, and the templated 19-contact batch sent 7/27 produced nothing measurable in nine days, while bespoke research-led outreach produced a real client. IMPORTANT SCOPE NOTE: this INVERTS the existing learning that says default to a multi-touch sequence rather than hand-personalized copy. That rule was derived from OTP coach outreach at list scale and remains correct there. It does NOT transfer to Sneeze It agency prospecting, where the ICP is small (roughly 800 franchisors) and each account is large. Check which lane you are in before choosing the artifact shape.
Failure mode: I diagnosed the Sneeze It cold-outreach problem as insufficient volume and proposed waves of 500 templated emails metered out through GHL at 25/day. I was optimizing for throughput when the evidence in front of me said throughput is exactly what has never worked.
Do not treat an open GHL opportunity in the Sneeze It Sales Funnel as evidence of a live relationship. David confirmed 2026-08-05 it is a graveyard: 221 of 234 open opportunities are parked at MQL and have never moved a stage. Only a WON opportunity (they became a client) is a hard block. Open, lost and abandoned opportunities are just a dated touch and should fall through to the COOLING or REACTIVATION tier based on how long ago that touch was. Separately, out-of-ICP companies (equipment manufacturers like The Abs Company, food, dental) belong in a dedicated ICP-exclusion file, not in do-not-blast.md, which exists for compliance and unsubscribe obligations.
Why: Blocking on open opportunities inverted the meaning of the pipeline. A stage that nothing ever leaves is a record of who we once imported, not who we are talking to, so using it as a suppression signal removes the exact brands most worth re-approaching. Reading a dead pipeline as a live one is the same class of error as reading a capped API response as a complete one: the data is technically present but means something different from what its name suggests. Check whether records in a stage actually move before you let that stage gate behaviour.
Failure mode: The suppression builder treated any OPEN opportunity in the Sneeze It Sales Funnel as a hard block, on the assumption that an open opportunity means a live deal that must not be cold-emailed.
David explicitly authorized DIRECT SENDING for the master-prospect-sheet lane on 2026-08-05, after I raised the risk and he reaffirmed. The human gate moved rather than disappeared: it is now the X he types in column A of the Sneeze It Master Prospect List, which is a per-company approval made before any email exists. Hard limits he set: max 30 sends per day, from dsteel@sneeze.it, one scheduled run per day, with a manual override command to run on demand. Every send stamps the sheet with sent status and date so a no-reply follow-up can fire ~30 days later. L096 is NOT repealed anywhere else: the `/sneeze-sales` GHL harvest still never applies a sequence tag and still never sends. Check which lane you are in.
Why: The reason behind L096 was never "an agent must not send" as a principle; it was that no cold email should reach a client or someone David is already talking to. A per-row human X satisfies that intent more directly than sequence enrollment did. The risk that remains is different and worth naming to a future operator: the X approves the COMPANY, not the SENTENCE, so nobody reads the email before it goes. That makes research accuracy and suppression freshness load-bearing in a way they were not when David hand-sent drafts. Rebuild the suppression list before any run, and never send to a BLOCKED domain, a bounced address, or an unsubscribe, regardless of what the sheet says.
Failure mode: Standing rules said the Sneeze It outreach agent never sends email: L096 required David to personally vet each contact and enroll them in the sequence himself, and the morning-pass floor says "prepare drafts, never fire them." Under the new sheet-driven process those rules would block the whole loop.
In-home care and senior care FRANCHISORS are IN ICP for Sneeze It as of 2026-08-05 (David: "low target and worth a try"). The qualifying trait is a multi-location franchisor with a real lead-generation budget, not a membership billing model. Do not hold or flag them. Because David framed it as a try rather than a conviction, tag these sends so their reply rate is measurable separately from fitness instead of being blended into one number: an experiment you cannot read the result of is not an experiment. The residential end (nursing homes, assisted living facilities) is still untested and was not part of this ruling.
Why: The ICP was written as a description of existing clients, who happen to be membership gyms, and I applied it as a boundary on who could ever be a client. Those are different things. What Sneeze It actually sells is paid lead generation plus a call center that works the leads fast, and any franchisor buying leads for local operators has that problem regardless of how they bill their customers. Watch for this shape generally: an ICP inferred from the current book will keep reproducing the current book, and the operator is usually the one who can see past it.
Failure mode: I flagged seven in-home care franchisors (Home Instead, Assisting Hands, Always Best Care, Comfort Keepers, Senior Helpers, Synergy HomeCare, FirstLight) as out of ICP and held them from sending, reading the Sneeze It ICP as membership-model fitness and wellness only and treating the inherited Nick-era "senior living" exclusion as covering them.
When a resource has an access rule on its primary page, extract that rule into ONE shared helper (canReadMeeting in services/meeting-read-access.ts) and apply it at every read surface, list AND single-row, instead of letting each endpoint hand-roll a partial check. Also: Ollie chat executes tools via app.inject with the user's own session (makeSessionOtpFetch), so fixing the HTTP endpoints automatically scopes Ollie's chat answers — no separate AI-context fix needed. Org-API-key callers carry no member row and stay org-wide by design.
Why: Access-control drift is a leak class, not a one-off: every new read surface (followups, exports, recordings, captures panel) shipped without the team gate because the rule lived inline in one page handler. A single source of truth makes the next surface safe by default, and knowing Ollie chat rides the user's session means endpoint-level authz is the one place to fix AI data exposure too.
Failure mode: SUCCESS: Conatus — root-caused critical R3V ticket 45f74cff (users could open any meeting): /l8/meeting/:id had a team/attendee/creator gate for months, but the followups page, all per-meeting read APIs (transcript, exports, recordings, agenda, headlines, share, SSE), and the meeting-captures panel only checked the 'restricted' flag — the gate existed but was never shared.
Adding a scheduled command to run-claude.sh takes THREE edits, not one: (1) APPROVED_COMMANDS, (2) a BUDGET case entry, (3) a RESOLVED_PROMPT case entry that expands the slash command into a literal instruction. The file contains SEVERAL separate `case "$PROMPT" in` blocks, so never anchor an insert on a command name alone; anchor on something unique to the target block (for the resolver, the line `RESOLVED_PROMPT="$PROMPT"` immediately above its `case`). Verify by extracting the resolver block and running it with the prompt as input, rather than by reading the diff: after the edit, `/outreach` must echo a real instruction string, and at least two pre-existing commands must still resolve to prove nothing was clobbered. `bash -n` passing proves only syntax, not that the edit landed in the right block.
Why: Both errors share a shape: a change that looks complete because the part you touched is correct, while the part you did not know about is untouched. The whitelist edit read as done because the command appeared in the file. The anchored insert read as done because the diff showed the right text. Neither was checked against behaviour. The saving grace was that run-claude.sh fails LOUD on an unresolved command, writing FATAL to the run log and an entry to alerts.log, so the job did nothing and said so rather than reporting success. That is the pattern worth copying into every scheduled job: an unconfigured job must be distinguishable from an idle one.
Failure mode: I added /outreach to APPROVED_COMMANDS in run-claude.sh and declared the scheduled job ready. It fired at 10:12 on 2026-08-06 and did nothing: approving a command is only half the wiring, and headless bare mode cannot resolve a slash command from the commands directory. Then, fixing it, I anchored the insert on the string ' "/otp-sales")' and landed the resolver inside the BUDGET case statement instead of the RESOLVED_PROMPT one, which would have left the original bug in place while also giving the job no budget entry.
A blocker on one item never blocks the others. Work every item that can be worked, exhaust every avenue on the ones that cannot, and only then report. Park the genuinely undecidable ones and keep going. A question for David is fine and should be raised, but it is raised ALONGSIDE completed work, never instead of it. Concretely for any queue-processing run: partition the queue into workable and blocked at the start, finish the workable set completely, and treat "I found a bug in my own tooling" as work to do inside the same run rather than a finding to report at the end of it.
Why: David's scarcest resource is his attention, and a report that says "here is why nothing happened" spends it without buying anything. The blockers were real but they applied to 2 of 12 rows; I let them set the pace for all 12. The deeper error is that stopping felt like diligence: I had genuinely found real problems, so surfacing them felt like the responsible act. It was not, because surfacing a problem is only half the job when the other ten items were sitting there workable the whole time. Exhaust the possible before you escalate the impossible.
Failure mode: The outreach run hit two items needing David's ruling (Powerhouse Gym, Fastest Labs) and a bug of mine blocking three more, and I reported all of it and stopped. Nine rows sat unsent while I wrote a summary. Worse, I had Clay-verified addresses for two of those companies (Annie Long at Senior Helpers, Jennifer Chasteen at Synergy HomeCare) already in hand from the previous session and did not use them. I presented blockers as a reason the batch could not proceed, when they were only reasons those specific items could not proceed.
When David asks for a report he can distribute, assume the audience is internal leadership, not the people being measured. Per-agent performance comparisons belong in that document at full detail with no warning attached. Do not volunteer to redact, soften, or produce a second sanitized version of performance data unless David names an external or team-wide audience. If the audience would genuinely change the content, ask which audience up front before building, never as a caveat appended at the end.
Why: Manager-level reporting exists to name who is converting and who is not. Attaching a sensitivity warning to that treats normal management reporting as a risk, and hands David an extra decision he did not ask for while he is trying to walk into a meeting. It also spends the final impression of the deliverable on a hypothetical instead of the findings. General shape: resolve audience before writing, not after.
Failure mode: Arin built the 14-day call center review Google Doc for David, then closed by flagging the Amanda vs Erica per-caller comparison as sensitive and offering to cut a sanitized second version before distribution. David corrected: the doc is internal and was never going to the call team. Both the hedge and the offer of a redacted variant were wasted.
A franchisee client does NOT block the franchisor, and the two are different companies on different domains (powerhousegymbridgeport.com vs powerhousegym.com). David 2026-08-06: "we do powerhouse bridgeport one location not corporate so if this is the corporate location good to go." Reverse the instinct entirely: an existing franchisee relationship is the strongest possible ASSET in a franchisor email, because it is proof delivered rather than claimed. Lead with it. The direction that matters is one-way: never cold-email a franchisee of a brand whose CORPORATE relationship we are mid-conversation with, and never email the specific client location, but corporate remains open and warmer than cold. Check which entity a domain actually belongs to before assuming brand-family contamination.
Why: I generalised one true fact (do not email a client) into a rule that would quietly delete the addressable market. Franchising is precisely the structure where one brand contains many independent buyers, so brand-level blocking is the wrong shape for this ICP. The asymmetry is worth remembering: blocking a franchisor to protect a franchisee costs a six-figure account and protects nothing, because the franchisee is not the one receiving the email. It also throws away the single best proof point available, which is that the brand already works with us somewhere.
Failure mode: I held Powerhouse Gym corporate (powerhousegym.com) as a policy risk because Powerhouse Gym Bridgeport/Stratford is an active Accelo client on that brand, and I proposed hard-blocking any brand domain where an active client sits anywhere in the system. That rule would have blocked every franchisor whose franchisee we already serve, which is most of the best targets we have.
Separate first-party link tracking from email-provider tracking. A tokenized link pointing at our own domain (orgtp.com/join/sales/<token>) is fully click-trackable no matter which mailbox sent it -- the prospect's browser hits our server and we stamp it, no provider cooperation needed. Manual personal-mailbox sending costs only the steps BEFORE the click: email opens (no pixel) and bounce/delivery visibility. When someone says "it came from a personal email so we can't track it", check whether the link destination is ours before agreeing.
Why: The two are routinely conflated, and the conflation kills features that would have worked. Here it nearly closed a ticket whose core ask (who clicked, who signed up, success rate) was fully buildable. The residual limitation is real but narrow, and worth stating precisely rather than as a blanket "not trackable": a dead address renders identically to a live address that ignored you, and the mint timestamp is not the send timestamp.
Failure mode: SUCCESS: Claude -- a join-link analytics ticket was nearly dropped on the false premise that sending from a personal mailbox makes clicks untrackable.
Before treating an OTTO queue as work to approve, check three things: (1) autopilot_ai_settings limits, where 0 means nothing is ever generated regardless of autopilot_is_active; (2) whether pending rows actually carry a non-empty recommended_value; (3) is_active as the approval flag, since the API's status field describes the current value's condition (compliant / invalid_length) and the status query param is accepted then silently ignored. When changing autopilot settings, always read-modify-write the complete settings object, because a partial PATCH risks dropping the other issue types.
Why: Counting pending tasks as available SEO wins overstates the work by orders of magnitude and invites a blanket approve. On orgtp only 171 of 779 pending titles were genuine defects; approving all of them would have overwritten 600+ pieces of good human copy with generic AI copy, damaging the pillar pages the Beacon impressions KPI depends on. The zero-limit config also means every client site, including Workout Anytime at 2,233 pages with issues, has an OTTO that has never run.
Failure mode: SUCCESS: Beacon found that OTTO's "pending task" count is not a backlog of approvable SEO fixes. On orgtp.com ~7,100 pending tasks all had an empty recommended_value, because autopilot_ai_settings carried limit 0 for every issue type on all 19 Search Atlas projects. autopilot_is_active reported true the whole time, so the config looked healthy while generating nothing.
When calculating any client credit or make-good on misspent budget, net out the value of what was actually delivered before proposing an amount. Formula: credit = spend under review minus (conversions delivered x a defensible cost per conversion). State which benchmark rate is used and why. Never default to crediting 100% of spend when the spend produced results.
Why: Crediting gross spend overpays the client and understates the work that did land. It also sets a precedent that any misallocation equals a full refund regardless of outcome. Netting delivered value is both fairer to Sneeze It and more defensible to the client, because it shows the math instead of a round apology number.
Failure mode: Drafted a client make-good credit at 100% of the misdirected spend ($2,501.53), treating the entire amount as a total loss. The spend was not a total loss: it delivered 6 real conversions to the client, and the draft gave that value away for free.
When changing the arity or shape of a function that other code passes a hand-rolled structural stub into, grep every caller for stubs BEFORE trusting typecheck. In otp-platform, registerOtpTools is called with `as unknown as Parameters<typeof registerOtpTools>[0]` in three places (the HTTP route, Ollie's collector in src/services/ollie-tool-registry.ts, and the parity test). That cast makes any missing method invisible to tsc: migrating otp-tools.ts from server.tool() to server.registerTool() typechecked clean while Ollie's collector, which only implemented .tool(), would have collected zero tools and removed every OTP tool from the chat box at runtime. Rule: a structural stub behind an `as unknown as` cast is an untypechecked interface. Treat it like a second implementation and update it in the same PR. Also: when two hand-maintained lists answer the same question (Ollie's WRITE_TOOLS/READ_TOOLS vs the MCP readOnlyHint/destructiveHint annotations), tie them together with a test rather than trusting them to be edited in step; ours had already drifted twice (discover_intelligence POSTs and INSERTs but was classified read, so it ran with no confirm card; sync_rules_to_file only rendered text but was classified write). Finally, verify a guard by breaking the invariant and watching it fail, not just by watching it pass.
Why: The failure mode is invisible to every automated check: tsc passes, all 3232 tests passed before the shim was fixed because the shim's own test used the same stale stub. It only surfaces as "Ollie can suddenly do nothing" in production. The generalizable form is that casts convert compile-time contracts into runtime hopes, and a codebase with structural stubs has as many implementations of an interface as it has stubs.
Failure mode: SUCCESS: Claude migrated all 56 OTP MCP tools to registerTool with directory annotations, and caught a shim that would have silently emptied Ollie's tool registry in prod.
Before diagnosing an OTP support ticket, identify the exact surface the reporter was on, then verify the reported cause can even occur in that state. From the 8/7 R3V batch: (1) "reassign meeting to another team" read as a feature request but was a creation bug — the UI offers "No team (personal)" and POST /meetings silently substitutes the Leadership Team; reassignment already worked via PUT /meetings/:id. (2) "Ask Ollie can't file a ticket" was mistaken identity — OTP has TWO assistants: Ask AI (corpus-only, no tools, so its refusal was truthful) and /ollie-chat (full MCP registry, has submit_ticket). (3) A plausible cause for a typing-freeze was ruled out by a precondition check: the transcribing banner only renders when the meeting has NO transcript, and the reporter had already generated Ollie insights. Say so honestly rather than shipping a fix under a false claim.
Why: Users describe symptoms, not causes. A plausible cause that survives no precondition check produces a fix that fixes nothing while closing the ticket. Checking which surface and which state rules candidates in and out cheaply, and turns "feature request" into "bug" often enough to change what actually gets built.
Failure mode: SUCCESS: Claude — three of six R3V support tickets had root causes different from what their titles said, and only reading the actual surface found them
Before designing any adoption, activation or gamification feature, query production for the funnel FIRST and count organizations rather than events. Then look for the outcome the product already produces and is only labelling wrong. Three rules that fell out and should be reused: (1) make progress steps OBSERVED FACTS re-checked on every render, never stored completion flags, so a step goes back down when its fact stops being true (a revoked key must not leave a badge behind); (2) start the ladder with a rung the user has ALREADY cleared (endowed progress) rather than at zero; (3) report the biggest ABSOLUTE drop, not the smallest number, because they are different steps -- here the largest loss was 49 orgs at signup-to-first-meeting, not the agent gap being investigated.
Why: A metric that requires a hand-written query gets checked once and then never again, which is exactly how a zero on the company's core thesis survived for months next to a dashboard that looked healthy. Counting events instead of organizations is the specific trap. And honesty is load-bearing on any adoption surface: the moment a number flatters, the whole surface is worth less than showing nothing, so no points, no badges, no streaks, and zero must render as zero.
Failure mode: SUCCESS: Claude found OTP's north-star metric was zero and nobody knew, by querying production before designing anything. 61 orgs, 12 ran a meeting, 1 ever created an agent seat, 0 agents ever called OTP. The reason it hid: /admin/usage counts ACTIONS (looks healthy, a few orgs run many meetings) while the thesis needs per-ORGANIZATION counting. The fix reused data we already had: Ollie's real work was already recorded per org in wallet_ledger.metadata->>'feature' and had only ever been rendered as billing. Read as a timesheet, the same rows prove an agent already works there.
The blocker for chart-drawn agent seats was structural: register_agent always mints a new seat, so pre-drawn YAML seats had no claim path. The fix was a claim_seat MCP tool (PR #564) plus file-based badge-in tooling (otp-badge-in.mjs, otp-agent-work.sh with keys in ~/.claude/otp-agent-keys/). Loop per agent: mint claim-mode enrollment token from /dashboard/agents/connect?agent=AGT_X, claim_seat with the exact id, get_my_seat, log_work. New agents use register_agent through the same page without the agent param. Also: the OOS publish gate counts BODY words only (frontmatter free); the L-rule text blocks in the body are a stale inert copy (claims table is canonical, carryForwardLearnings preserves it), so slimming deletes them safely with a claims-table diff as the verification oracle.
Why: Ten agents connected in one pass (Radar, Dan, Dash, Pepper, Crystal, Pulse, Neil, Arin, Tally claimed; Outreach registered fresh). Agent adoption went from zero to one org same day the leak was diagnosed. The claim-vs-register distinction and the body-only word gate are non-obvious and will recur for every org with template-drawn seats and every future OOS slim.
Failure mode: SUCCESS: Conatus ran the first agent connect pass; Sneeze It became the first org on OTP (of 61) with live agents on the chart
When Clay and LeadMagic's email-finder both return nothing for a confirmed named human, do not mark the row "needs research". Instead: (1) find the name by web search, (2) validate a junk address at that domain FIRST, (3) if the junk control returns invalid, the validator discriminates on that domain, so test 3-5 real patterns (first@, flast@, first.last@) and any that returns valid is a real mailbox. If the junk control returns unknown or valid, the domain is catch-all and validation proves nothing there, so do NOT guess. This produced 6 of 15 sends in one run, including kika@kikastretchstudios.com and dfink@iflexfranchise.com, which no enrichment tool returned. Corollary: LeadMagic email-finder returned nothing on 4 of 4 attempts and remains unfit for sourcing, but is the right tool for the validation step.
Why: The held-row problem was misdiagnosed as "no named human", which framed it as a Clay/enrichment gap. It is actually an address-verification gap: most held rows HAD a confirmed decision maker. Naming the bottleneck correctly is what moved throughput. The control test is what makes pattern guessing safe rather than reckless, which matters because the 7-day hard bounce rate was already 9% against a 10% stop line, and unverified guessing on catch-all domains is precisely what pushes a sending domain over that line.
Failure mode: SUCCESS: Outreach tripled per-run sends (5 to 15) by treating address VERIFICATION, not name-finding, as the real bottleneck, and by testing address patterns against a per-domain junk control.
When extracting an EJS partial, pass EVERY value it reads explicitly; a template-scope var/function can never be inherited by an include no matter how the include is written. And build render-test fixtures from the route's actual reply.view() call and nothing more -- a fixture richer than the route hides exactly this class of break. Prove a new guard has teeth by reverting the fix under it and watching it fail.
Why: A fixture more generous than production turns a render test into theatre: it renders a page that cannot exist, so a page-down bug ships green through CI. The verification I did after merging (health commitSha) confirmed the DEPLOY, not the PAGE, and an anonymous GET only returns a sign-in redirect -- so nothing I checked would ever have caught a crash in the authenticated render.
Failure mode: Extracted an EJS row into a partial on /l8 and passed only { m: m }. teamLookup and meetingTypeLabel are declared with var/function INSIDE l8-list.ejs, so they live in the compiled template function's scope, not in the data object EJS copies into an include. Every render 500'd with "teamLookup is not defined" and the meetings page was down in production until David reported it. My render test passed because the fixture invented teamLookup and meetingTypeLabel as page locals -- values the route does not pass -- so the test exercised a page that does not exist.
Sales-invite prospects (/admin/join-link) are added to the Swamp subscriber list at mint, so they already receive the weekly. Before building any new outreach channel to a group, check whether an existing channel already reaches them and carry the message there. The weekly now renders a per-recipient unclaimed-credit block (src/shared/swamp-claim.ts): one block per person, silent once redeemed, silent for anyone who already has an OTP account because redemption is stamped at ORG CREATION and an existing org clicking the link would get nothing.
Why: 18 prospects held unredeemed credit, 0 clicks and 0 redemptions, and had been reading the weekly for weeks with no mention of the money set aside for them. The reach existed; only the message was missing. The account-holder rule is the part that is easy to get wrong: it would send a money promise the product cannot honour.
Failure mode: SUCCESS: Swamp -- the weekly email already reached the prospect list that was never told about its credit
Read the no-EOS-recipients rule at its actual scope. It governs cold blasts, list sends, and the weekly Swamp digest. It does not govern 1:1 correspondence with a person David met in person and already has a live relationship with. Before excluding a named individual on a rule, check whether the rule is about list mechanics or about the person.
Why: Over-applying a permanent rule silently drops real relationships out of David's pipeline. Shemtov was met in person on Jul 13, was described as genuinely taken with OTP, and had already been sent a workspace link on Aug 3 from that same address. Excluding him would have quietly killed a warm follow-up on a technicality that did not apply.
Failure mode: Applied the permanent no-EOS-recipients rule as a blanket block and excluded Rabbi Mendel Shemtov from a rabbi outreach list because his address is mendel.shemtov@eosworldwide.com. David corrected this: Shemtov is one of the three.
When a row company name carries a city or territory suffix but the domain is the brand corporate domain, do not enrich the corporate domain hoping for the local owner. Find the unit own site or Google Business listing and enrich that domain instead, or mark needs-research noting the row needs a unit-level domain. Never email the corporate address on a franchisee row because it reaches a different company than the one approved. Also mark CBD and cannabis retail rows out-of-icp on sight, since paid ads for them are restricted on both Meta and Google.
Why: This pattern accounted for most held rows in the 2026-08-12 run and is the concrete shape of the no-named-human problem that got the outreach schedule killed on 2026-08-10. It converts a vague sourcing complaint into a fixable data problem, namely that the sheet needs unit-level domains on franchisee rows.
Failure mode: SUCCESS: Outreach found why prospect rows keep landing in needs-research instead of sending. Rows named for a single franchise unit (BodyBrite South County, Buff City Soap Birmingham, Cardio Plein Air Haute-Yamaska) carry the franchisor corporate domain in the domain column, not the local operator domain. Clay resolves it to corporate HQ and returns no usable local contact.
~/.claude/google-ads.sh pinned API_VERSION=v21, which Google began blocking with "Version v21 is deprecated" - and it failed INTERMITTENTLY, so roughly half of a 12-query batch succeeded and half returned INVALID_ARGUMENT. Bumped the pin to v23 (v22/v23/v24 all work; v24 costs ~2.5x the query resource units). When any Google Ads pull returns partial or inconsistent errors across identical queries, check the pinned API version first before assuming rate limiting or token trouble.
Why: A partial-failure sunset is far more dangerous than a hard failure: an agent that does not inspect every error body will report averages computed from half the periods and present them as complete. Every agent reading Google Ads (Dash, coach report, billing report) shares this one pinned constant.
Failure mode: SUCCESS: Dash caught a silent Google Ads API version sunset mid-pull
When an /outreach send is refused with "Blocked by classifier": do not retry verbatim, and do not reach for an alternate send path such as the Gmail MCP directly, since that bypasses the four-place recording the sheet depends on. Instead (1) confirm the block is scoped to `send` by running a harmless subcommand like `status`, (2) get David's explicit authorisation, (3) add a SPECIFIC allow rule "Bash(python3 ~/.claude/scripts/outreach-queue.py *)" plus the absolute-path twin to ~/.claude/settings.json, (4) re-run the send WITHOUT a "cd ~ &&" prefix, using the absolute script path. Never try to add an autoMode.allow entry: the classifier blocks edits to its own config by design and that boundary should be respected, not routed around. Separately, a "Stage 2 classifier error - usually transient" refusal is a different thing and one retry is legitimate there.
Why: A broad allow rule does not clear the auto-mode classifier for outbound-email actions, but a narrowly scoped rule naming the exact script does. Without this an /outreach run looks completely broken and produces zero sends despite a healthy queue, healthy suppression and a working script. The fix is permanent and one-time, so recording it stops the next run losing an hour to the same dead end. The negative half matters as much as the positive: self-granting classifier permissions and side-channel sending would both technically work and are both wrong.
Failure mode: SUCCESS: Outreach — /outreach sends were blocked mid-run by the Claude Code auto-mode permission classifier, not by a missing tool or bad data. Non-obvious because Bash(python3 *) was ALREADY in the settings allow list, so it looked like a tool failure rather than a permissions one.
Treat a Make alert email as a timestamp of an event, never as current state. Before surfacing any Make scenario as down, read live state via the Make MCP: organizations_list for the team id, then scenarios_list, and judge on isActive, isPaused, dlqCount and the errors-to-executions ratio. Report the ratio, not the alert. A scenario whose errors equal its executions has never worked and outranks anything that merely stopped once.
Why: Make scenarios stop, alert and auto-restart constantly, so alert emails generate false urgency while burying real failure. Reading live state on 2026-08-14 turned five urgent-looking alerts into one genuine finding: Beem Atlanta Glenwood Fix Any Names (5623886) had 44 errors across 44 executions, a 100 percent failure rate since creation five weeks earlier, which no alert email distinguished from routine noise. The same read caught an unfinished cutover, where the scenario named "(Stop Using) CCM - Speed To Lead - New Lead Tracker" still carried 2,149 executions while its replacement had 66.
Failure mode: Radar reported five CCM Make scenarios as currently stopped and needing escalation, based only on the "scenario has been stopped" alert emails in the inbox. Querying Make directly showed all five had already restarted and were active with zero DLQ. The briefing escalated a resolved condition, and the emails also hid the item that actually mattered.
Treat every club/location number as an opaque string used exactly as issued. Never pad, trim, or parse one to a number, and never infer it from a display name. When a location read 404s, test the padded and unpadded forms before concluding the data is missing — the request is more often malformed than the club absent.
Why: A 404 from a padding mismatch is indistinguishable from a missing club in logs, so the wrong diagnosis (client not connected, data not synced) is the natural one and can cost days. It also cuts the other way: 04462 parsed as a number becomes 4462, which is a different club or none.
Failure mode: SUCCESS: iJoin — verified ABC production sandbox access and found that club-number padding is part of the identifier, not a format to normalise. GET /rest/9003/clubs returns 200 ("09003 PROD Test Club"); GET /rest/09003/clubs returns 404. The club's own display name carries a leading zero the API identifier does not.
Before writing any code for an OTP support ticket, search for prior work by the ticket's 8-character id PREFIX, not the full UUID: `gh pr list --state all --search "<keywords>"`, `git branch -a`, and `grep -rn "<prefix>" src/`. A hand-typed full UUID in a PR body or commit message is often mistyped, so a full-UUID search returns nothing even when the work exists. Then verify the fix is actually live rather than trusting the PR narrative: compare /health commitSha to the merge commit, re-run typecheck plus the ticket's tests, and confirm any customer-facing claim (e.g. the retention sentence on /trust) renders in production. Fix the PR body id before closing the board row.
Why: The ticket instructions said "I think we started it" and that was right: 698 lines across 12 files had already shipped as PR #616. Writing the feature again would have duplicated a day of work and risked a conflicting second file store. The mistyped UUID (fab4877c-61c1-46bc-... instead of fab4877c-67c1-4b26-...) would have left the merged PR pointing at a ticket that does not exist, so the board row could never be traced back to the code that resolved it. The 8-char prefix was correct everywhere in the code comments, which is why prefix search finds what full-UUID search misses.
Failure mode: SUCCESS: Claude (OTP dev) - an OTP support ticket assigned as "fix this" was already built, merged and deployed; the real remaining work was a wrong ticket UUID in the PR body that would have broken traceability when closing the board row.
Before building any OTP feature that overlaps an existing domain, do two reads first. (1) Grep for a subsystem that already does part of the job and check whether anything CALLS it: `src/services/coaching/` was 250KB of working code that no route or job had ever invoked, so the new feature became its first caller instead of a duplicate. (2) When a house rule says a term is banned (e.g. de-EOS), grep the whole tree before assuming compliance: OTP deliberately ships EOS marks on /templates/level-10-meeting and in the site footer under a nominative-use disclaimer. Encode that as an ALLOWLIST in the guard test with a staleness assertion, never as a blanket ban that would fail the build on shipped pages. Also: when a numeric helper clamps (Math.max(0, ...)), never route a signed difference through it. delta and ratingGap both silently became 0 for every falling score.
Why: Building the scorer beside the dormant Coach would have produced two coaching systems with different confidence levels and no way for a customer to tell which was which. Writing a blanket "no EOS marks" guard would have failed CI on a live SEO page that earns traffic and is legally covered. And the clamp bug would have shipped silently: it only manifests when a team's meetings get WORSE, which is exactly the case the feature exists to surface, so no happy-path test would ever have caught it. All three were found by reading before writing.
Failure mode: SUCCESS: Claude (OTP dev) - shipped Ollie meeting scoring end to end, and the two things that mattered most were both discovered by reading the existing codebase rather than by building: a large never-called subsystem, and a deliberate trademark carve-out that looked like a compliance miss.
When building a feature that JUDGES something (scores, grades, health ratings, risk levels), the unit tests only prove the arithmetic. Before showing it to anyone, run it against a real org's full history and inspect the DISTRIBUTION, not samples. Specifically: (1) count how much of the total measured weight lands on exactly zero, since a dimension that is 45-for-45 zeros is a bug, not a finding; (2) check whether any dimension is NEVER unmeasured, which means an absence is being scored as a failure; (3) compare two orgs, because a defect that only fires on one customer's data shape is invisible with one; (4) never route a signed difference through a clamp written for a bounded score. Also: a judgement built on thin evidence must be withheld with a reason naming what to capture, not published as a low number, and the list must show the items that scored NOTHING because those are the ones worth acting on.
Why: Concretely: R3V had 17 meeting to-dos, all with named owners and none with due dates, so an all-or-nothing rule scored all 45 of their meetings 0 and made 75% of their measured weight zero. Fifteen meetings published 0.0 on 20% coverage. Falling scores rendered as "level with your average" because a signed delta went through Math.max(0, ...). And the list hid 27 of 36 meetings, which were exactly the ones a coach needed. Sneeze It's data shape (161 of 197 commitments fully formed) hid every one of these. A green suite proves the code does what you specified; only real data tells you the specification was wrong.
Failure mode: SUCCESS: Claude (OTP dev) - shipped a scoring feature that passed 4,000+ green tests and was still wrong three separate times. Every defect was found by running it against a real customer's data, not by testing.
(1) The coach report is now generated by ~/.claude/gen-coach-report.py <scratchpad>: it parses meta-ads.sh/google-ads.sh 'active 7/30' text plus ccm.json/otto.json/rank.json and emits the HTML with rule-based wins/recs and a Watch flag per churn tripwire; edit the client config block, do not hand-write 50 cards. (2) Search Atlas keyword rankings via MCP return ~1MB per project; call https://keyword.searchatlas.com/api/v1/rank-tracker/{id}/keywords-details/ directly with python requests (urllib gets 403), passing searchatlas_api_key plus period1/period2 start/end, and keep only keyword/pos/vol/hist. (3) The 'active' CLIs report leads only for action_type=lead; VENT purchases (18/wk) and Champy's site visits need a direct /insights actions call, otherwise those cards read as zero-lead failures.
Why: Hand-assembling ~50 cards each Monday is where the errors and the 3-upload duplicates came from; a data-driven generator makes the weekly run a data pull plus a config edit, and the two API quirks (1MB MCP payloads, lead-only action filter) would otherwise re-cost an hour every week.
Failure mode: SUCCESS: Dash /coach-report 2026-08-16 shipped 53 cards with all sources live and made the run repeatable
In l10dan prep: (1) Headlines sweep must include David's own work, not just agent files: check git log across ~/ijoin-platform, ~/otp-platform and any repo touched in the last 7 days, the memory files dated this week, and the calendar, then lead with the biggest thing David did. (2) IDS: for the top signal, run the decomposition BEFORE the meeting (numerator vs denominator, one client vs all, one week vs trend) from live data, so the brief carries a data-backed candidate issue with a suggested owner and board, not just a signal. (3) When an issue belongs to another team, write it in plain English from David with no agent names, file it on that team's board with a human owner, and close it on the AI Army board with the cross-reference.
Why: David's standard is no discovery left for him in the room. Headlines that miss his own week and issues that get decomposed live both put the discovery back on him. Cross-team issues that mention agents by name are unreadable to the humans on the receiving board.
Failure mode: Dan's Delta Meeting prep (8/17, rated 8.5) still under-delivered on two sections. Headlines: the brief led with agent state files and missed David's own biggest work of the week (ijoin.ai mostly built over the weekend, two products named for the Product Engine rock) until David named it in the room, even though the repo, memory files, and git log were all local and readable. IDS: the brief arrived with grouped signals but no data-backed candidate ready; the numerator-vs-denominator check on the Arin drop was run live from the CCM sheet during the meeting, and it changed the issue entirely (dial collapse = auto-dialer; the real issue was WOA converting at ~11% vs ~30% for a month).
Treat Workout Anytime Yadkinville as excluded from call-center coaching and portfolio analysis (same class as China Grove). Never flag its leads-with-zero-dials as a caller miss. Add it to the standing exclusion list in the Arin/Dash rules and check that list before naming any project in a recap or coaching note.
Why: A project that appears in the sheet but is out of contract scope produces false "uncalled leads" coaching that erodes the callers' trust in the recap and could push them to dial leads we are not paid to work.
Failure mode: Arin's CC recap draft (2026-08-20) told the callers to clear Workout Anytime Yadkinville's uncalled leads. David corrected: we do not call for Yadkinville anymore. Yadkinville rows still appear in the CCM sheet with leads and zero dials, which reads as an uncalled pile when it is actually out of scope.
Clay is the validation and enrichment layer for the Outreach Engine. Contacts arrive at the intake API already carrying validation_status from Clay's validation waterfall; old/CSV lists get exported to Clay for validation and pushed back. LeadMagic stays only as a dormant pluggable provider behind an explicit VALIDATION_PROVIDER env, never the default or the recommendation.
Why: LeadMagic was the Nick cold-prospecting stack, retired 2026-07-03. The current sales stack is Clay to GHL and Clay to Outreach Engine; recommending a retired tool's key adds a vendor, a cost, and a contradiction with the process of record.
Failure mode: The Outreach Engine build wired LeadMagic as the default email validation provider (leftover assumption from the retired Nick-era tooling) and told David to set LEADMAGIC_API_KEY on Railway. David corrected: validation runs through Clay, not LeadMagic.
Two things. (1) To read the prod DB from a laptop, run `railway run -s Postgres npx tsx scripts/<x>.ts` — capital P, and the script must build its own pg.Pool preferring DATABASE_PUBLIC_URL. Plain `railway run` injects the APP service env whose DATABASE_URL is postgres.railway.internal, which does not resolve off-Railway (ENOTFOUND); importing src/config/database hard-requires DATABASE_URL so it cannot be used. (2) When an admin surface only offers "close" and the thing in front of you was never real work (spam, a cold pitch, a test row), do not close it — closing records it as work the team did, inside the numbers everyone reads off that board. Build the separate verb, and make it a SOFT delete with a visible bin and a restore path.
Why: The DB-read trick cost real time twice now and the older note in memory documents a command that no longer works. The close-vs-spam distinction is the more valuable half: reaching for the nearest available verb quietly corrupts the metric the surface exists to report, and a delete with no visible other side is indistinguishable from a permanent one to the person clicking it.
Failure mode: SUCCESS: Claude — reading the OTP prod DB locally, and why "close" was the wrong verb for spam
Use the full Kennedy machinery, not one device: (1) open with a visceral, do-it-yourself demonstration of the pain (a test the reader can run on their own team), (2) agitate with specificity before naming the product, (3) present the mechanism as the escape, not a feature list, (4) single low-friction reply-word CTA that earns a concrete deliverable, (5) P.S. with disqualification or risk reversal. If copy reads as an explanation, it is not Kennedy yet.
Why: Cold email to founders lives or dies on the first two lines making the pain concrete. Feature explanation is what every SaaS email does; the demonstration-test opener is what gets forwarded to the leadership team.
Failure mode: The OTP intro email draft (cp_f64eae3f263836654812) wore a Kennedy veneer (three-negation opener) but then feature-dumped: paragraphs explaining what OTP is. David: "you can do better dan kennedy."
Check-in is BOTH people. Dan gives his own personal item and business item, unprompted, right after David gives his, and then waits. Never move to the next section until both sides have checked in and David signals ready. Slow down: one section per message means one section per EXCHANGE, not one section per message dumped in sequence.
Why: The partnership frame is the point. An agent that collects the human's update and then reports numbers is a servant running a form. A co-founder shows up to the check-in too. David caught it in the room, which is the exact failure mode the Clean Agent Meeting rock was supposed to close: holes found live instead of in prep.
Failure mode: In the Delta Meeting check-in, Dan relayed David's personal and business update back to him and immediately moved to the scorecard, never offering Dan's own check-in. Two-person meeting, one-sided check-in, and it read as rushing David through his own section.
When an agent needs access to a system David owns, default to a service-level credential issued from the account that owns the resource, not to adding an existing personal account as a member. The invite route makes the agent's access a dependency on one human's personal membership, which breaks the moment that person is removed or that account changes. Offer the invite only as the fast temporary unblock, and say so.
Why: Dan optimised for "fewest steps right now" and missed the architecture point. David is building a system of record for client delivery; the agent reading it should hold its own credential. This is the same principle as the per-seat OTP agent keys already in use.
Failure mode: Dan told David "do NOT create a new API key" for the new Trello, and pushed the invite-davidsteel12 route instead, twice, including in a filed to-do. David overrode it: he wants the connection made via API.
Before porting an agent to a new system, inspect the live data shape first. Trello looked like a PM system but held zero due dates, zero comments, and checklists that were contact rosters rather than deliverables — so the inherited KPIs (overdue projects, milestones, delivery risk) were uncomputable. The seat was re-scoped to what the data supports: request flow, roster completeness, and board adoption. Also: define workflow "churn" as BACKWARD list moves only. Counting total moves flags every healthy card, since a clean card moves forward twice by design. Always pair a small-n metric with an explicit fallback and an (n=X, directional) label.
Why: Porting a spec verbatim onto a new tool produces confident fiction — numbers that look like the old report but measure nothing. Inspecting the live schema first turned an unbuildable JD into five KPIs that compute today, and surfaced the real finding: 2 of 11 people use the boards, which governs every other number.
Failure mode: SUCCESS: Crystal rewritten from Accelo to Trello — measure what the tool actually holds, not what the old spec measured
On day one of a rollout, report capability, not compliance. Frame the doc as what Claude can now answer about the system, with example questions the person can ask, and present current numbers as a neutral first snapshot explicitly labeled as too early to mean anything. No targets, no goal columns, no assigned homework until the tool has real usage history. Name Claude openly when the reporting capability itself is the subject. Invest in visual formatting (brand colors, shaded table headers, hierarchy) rather than shipping plain text.
Why: A scorecard on day one reads as judgment of the two people who actually adopted early, and it buries the thing that would drive adoption: showing people what they now get for free. Capability framing pulls people toward the tool; compliance framing pushes them away.
Failure mode: Wrote a team-facing doc for the first day of a new tool as a scorecard: baseline-vs-target table, a two-week task plan, and an adoption number (2 of 11) framed as a gap. Also buried the actual point, which was what reporting is newly available.
Before designing any integration that CREATES records in a partner system, check what the partner system does to its own attribution fields on API-created records, and check the account-count ceiling on the app tier you plan to ship. For Jobber specifically: (1) clientCreate permanently overwrites Jobber's Lead Source with the connecting app's name and Jobber users cannot edit it, so real attribution must live in our own DB plus an app-configured custom field, and we must never promise a franchisee that Jobber's native Lead Source will show the true source; (2) a Draft-state Jobber app is cut off past 5 paying Jobber accounts, so a franchise rollout needs either Jobber approval for a Custom Integration or App Marketplace publication, decided before onboarding resumes, not after.
Why: Both failures are invisible in testing and only appear at scale: the Lead Source overwrite looks fine until the client asks why every lead says the app name, and the 5-account cap looks fine until franchisee number six connects and API access is blocked mid-rollout. Reading the platform docs for what the vendor does TO your data, not just what you can do WITH it, is the cheap step that catches both.
Failure mode: SUCCESS: Claude found two Jobber platform landmines that would have silently broken the HBFG Jobber<->GHL connector after launch, by reading the API docs before designing rather than after.
After changing a brand default booking link in the outreach engine, verify at the campaign level before the next send: (1) body_html must contain {{calendar_url}}, not a literal URL; (2) campaigns.calendar_url must be NULL, not blank — renderTemplate uses `campaign.calendarUrl ?? brandDefault`, so an empty string wins over the default. Query: select name, calendar_url, body_html like '%{{calendar_url}}%', body_html ~* 'calendly' from campaigns where status not in ('done','archived','sent').
Why: A brand default that no campaign reads is a setting that looks changed and is not. The link is also what makes the visit pixel work: ijoin.ai is an ownHost so the merged URL gets a signed ?sv= claim, while a Calendly URL gets wrapped in a /c/ redirect and the prospect is invisible on arrival.
Failure mode: SUCCESS: Outreach — changing a brand's default booking link does NOT change what a queued campaign sends. Found the iJoin Touch 1 campaign still pointing at Calendly after the ijoin.ai/intro switch, for two independent reasons at once: the CTA href was a hardcoded URL rather than {{calendar_url}}, and the campaign row carried its own calendar_url override that beats the brand default.
When a major feature lands inside a crowded issue as a mechanism card, treat the case FOR it as still unsent and worth its own issue. Write it as "turn this on and here is why", never as a fresh reveal, because attentive readers already met the machinery. Verify every behavioural claim in code before writing it: this issue's claim that the score reaches the facilitator and coach but not the whole team was checked against resolveRecipients() and the meeting_score_settings defaults (notifyFacilitator true, notifyCoach true, notifyAttendees false) before it went in the copy.
Why: A quiet shipping week is not automatically a skipped issue. Announcement volume and persuasion are different jobs, and a feature can be fully announced yet entirely unsold. Checking recipient behaviour in code keeps the difference between a marketing claim and a true one, which matters most in the entry asking someone to switch something on.
Failure mode: SUCCESS: Swamp turned a one-item week into a coherent issue by re-framing an already-announced feature. Meeting scoring shipped 8/15-8/16 and went out as issue #33's pick, but written as a mechanism card (rubric, evidence thresholds, exclusions) buried as one of twelve. David named the real story at the gate: better meetings mean better outcomes, it is about steering the ship in the days between meetings, not volume. That argument had never been sent.
When any threshold-plus-minimum-sample rule is written, check the granularity: divide 100 by the minimum sample to get the smallest non-zero rate the rule can observe. If that value exceeds the threshold, the rule is incapable of returning "acceptable" and will fire on the first ordinary event. Set the sample so the threshold sits several increments above the granularity floor (6% over 100 gives 1% steps). Also gate cold-email sending on COMPLAINTS rather than unsubscribes: complaints are what burn a sending domain, unsubscribes are the polite exit and are expected. Resuming after such a pause needs the override path, not a plain resume, because a plain resume re-reads the same tranche and re-pauses within seconds.
Why: A breaker that cannot pass looks identical to a real deliverability problem from the outside, and the pause reason it writes ("the copy or the audience is wrong") actively misdirects the operator into rewriting healthy copy. The campaign sat stalled for a day and the stated cause was false.
Failure mode: SUCCESS: Outreach found why the EO blast silently stalled at 75 sends. The unsubscribe circuit breaker was configured at unsubscribePausePct 2.0 with unsubscribeMinSample 25. At n=25 a single opt-out is 4%, double the threshold, so the rule could only ever return "pause" or "perfect" and a tranche passed only at exactly zero unsubscribes. It was a stop sign wearing a guardrail's clothes, and it stopped a campaign showing 25.3% clicks, 0 bounces and 0 spam complaints.
When a feature reports "the data is there but users never see it", check the WRITE ORDER before checking the render logic. An optional section that reads a column written by a later, user-triggered step is dead code in practice. Fix by moving the send to the step that produces the data (here: the accept/approval step), and lock it once-per-entity with a conditional UPDATE on a timestamp column claimed BEFORE the slow part, so concurrent triggers cannot both send.
Why: Kristen reported the symptom as a missing feature ("can we email Ollie insights?"), but the feature already existed in the template. Building it as new would have produced a second, duplicate sender. The real defect was a timing gap, and timing gaps in "best-effort, non-fatal, backgrounded" code are invisible to tests and to logs: nothing fails, the section is just absent.
Failure mode: SUCCESS: OTP build -- "Ollie's read" section in the post-meeting recap email was structurally always empty, and nobody noticed because the code looked correct. The recap sends at endMeetingCore; the read is written later by the follow-ups wizard into meetings.ai_summary. The template said "render it when it exists" and it never existed yet.
iJoin positioning: never describe it as a page, a connector, or an AI integration. The category claim is that the join flow stops being frozen infrastructure and becomes testable creative. Today a club has ONE join page, handed to them by their member-system vendor, unchangeable without a ticket, impossible to A/B. iJoin makes join flows unlimited and disposable: one per campaign, per location, per offer, all live simultaneously, all writing into the same system of record, tested against each other like ad creative. For a media buyer the line that lands is that they optimise every surface except the one where the money actually changes hands. The live write into ABC, Glofox or Zenoti is the ENABLER, not the pitch. Also: when writing Touch 2 or later, verify the follow-up carries the thesis at least as far as Touch 1 did, never less far.
Why: Selling the enabler instead of the category makes iJoin sound like a plugin competing on features against member-system vendors, which is a fight on their ground and a small idea. The reframe is what makes it worth a conversation with a multi-location operator, and it is the part no competitor is claiming. A follow-up that narrows the thesis also de-sells: a reader interested by Touch 1 gets a smaller reason to act on Touch 2.
Failure mode: Drafting iJoin Touch 2 outreach, the P.S. thesis read "the page was never the hard part, writing a real membership into the system you already run is the hard part, and it is the only thing iJoin does." That shrinks iJoin to a connector. David corrected it: iJoin is a whole other way of thinking about joining online, and the point is unlimited join pages you can test against each other. Touch 1 already carried the bigger idea, and the follow-up narrowed the thesis instead of advancing it.
iJoin has THREE pillars and copy must carry all three. (1) Unlimited join flows, one per campaign, per location, per offer, all live at once, all writing to the same system of record, tested against each other like ad creative. (2) The Beacon/Pixel, which follows the individual journey: who came back, what they looked at, how far they got. (3) Workflows that fire on what somebody DID, not on what day it is. Reached the plans and stopped, opened the agreement and stopped, each gets its own next move. The join page is the DEMO, not the product, and should always be introduced as "the easiest part to show you." Strongest available proof: our own outreach engine already works this way, so a follow-up triggered by a prospect reaching the booking page and not booking is itself a live demonstration of pillar 3. Say so in the copy.
Why: Pitching the page alone reduces a category to a feature and invites a comparison against whatever join page the member-system vendor already ships. The pixel plus action-triggered workflow is what turns it into a system nobody else is selling, and it is the half that a multi-location operator with a media budget actually values. Demonstrating pillar 3 inside the email that sells it is proof rather than a claim, which no competitor can copy without having built it.
Failure mode: L238 captured only one of iJoin's three pillars. David expanded: iJoin is also the Beacon/Pixel that tracks a person's journey, and customised workflows that fire on ACTION rather than timing. His words: "so much more than just a page, but the page was easiest to show."
The ICP bar for someone who ALREADY ENGAGED is lower than the bar for building a cold list. On a cold list a vendor or a domain mismatch is noise you paid to add. Someone who opened the email, clicked through and reached the booking page has selected themselves, and an adjacent-industry vendor doing that is a partner or referral conversation worth having, not a mistake to filter. So: apply vendor and domain-coherence screening at LIST BUILD time, and do not re-apply it as a gate on engagement-triggered follow-up. Surface the oddity for a human to see, do not withhold the send. Reserve holds for real conflicts (active client, litigation, direct competitor) rather than "this looks off".
Why: Filtering the engaged list on the same rules as the cold list throws away the highest-intent contacts in the system, which are exactly the ones the whole tracking apparatus exists to find. It also makes an automated trigger quietly narrower than the human it replaced, so the automation looks like it is underperforming when it is actually being over-restricted.
Failure mode: I held four vendor-ish contacts out of the hand-built iJoin Touch 2 list (SportsArt, Paramount Acceptance, Peterson Partners, a mismatched domain) and then flagged two more the action trigger picked up automatically (iKizmet, a fitness BI vendor, and a name/domain mismatch at HCOA), recommending we hold them. David said let them go.
Do not tell David to wait for data before building a surface he will have to look at anyway. The list view, the intake, the schema and the page are needed whatever the match rate turns out to be; only the thresholds and the automation depend on the numbers. Build the surface immediately so data lands somewhere legible from day one, and leave the tuning knobs configurable. "Earn complexity / data before design" applies to AUTOMATED DECISIONS, not to giving a human somewhere to look. Two weeks of data piling up in a vendor's UI is two weeks of nothing to read.
Why: David builds the machine first and tunes it against reality, which is how the outreach engine, the scanner classifier and the action triggers all got good inside a day. Advising a pause defers the moment data becomes visible and useful, and the visible version is what generates the judgment the tuning needs. It also reads as caution about his strategy rather than about the code.
Failure mode: After shipping the Clay pixel I recommended waiting two weeks for match-rate data before building the Clay-to-outreach connection and the Companies view. David overruled it: "what no build it now, we are going to make us a sales machine." The recommendation confused two different things: tuning thresholds needs data, but building the surface does not.
Two reusable patterns. (1) The auto-mode classifier blocks COMPOUND shell commands (for-loops, && chains) around gh pr merge, git worktree remove --force, and polling loops, while the identical single-action command passes: when a batch is blocked, unbatch into one plain command per action instead of retrying or routing around. (2) The sweep recipe held end to end: read board via railway run -s Postgres npx tsx scripts/read-support-board.ts, check prior work by 8-char ticket prefix + gh pr list keywords BEFORE building, one Explore agent per subsystem then one worktree builder agent per ticket in parallel, squash-merge in order, poll /health commitSha until it equals the last merge commit, then otp-support-close.sh close --status=resolved with customer-facing resolutions (no em dashes, never blame the user).
Why: The classifier single-vs-compound distinction turns a hard stop into a 10-second adjustment next sweep. The recipe confirmation means the next board sweep can run the whole loop without rediscovery: this one went from 3 open tickets to 0, all fixes live, in about an hour wall clock.
Failure mode: SUCCESS: Claude (OTP dev) - 2026-08-26 support sweep: 3 R3V tickets fixed via parallel worktree agents (PRs #656/#657/#658), merged, deploy verified on /health, board closed to zero, in one pass
When wiring a Clay HTTP API enrichment column to a POST-only endpoint, expand the collapsed "Method" field under SETUP INPUTS and explicitly select POST, even though Clay marks it Optional. A 404 (rather than a 400) from a URL you know exists is the tell that the HTTP method is wrong, not the payload. Also: do not wait for organic traffic to validate a webhook hop - type a fake domain into the source column, confirm Status Code 200, then delete the test row and the record it created.
Why: All four brand signals (ijoin, OTP, Orger, sneeze it) were configured identically and would have failed silently forever, reporting a healthy "waiting for visitors" state while never delivering a single company. An Optional label on a field the receiving endpoint cares about is a trap, and forcing an end-to-end test is the only way to find it before real traffic arrives.
Failure mode: SUCCESS: Outreach - Clay HTTP API columns 404'd because Method is labelled "Optional" and silently defaults to GET
When automating a step where a machine picks who gets contacted, do not remove the human approval, MOVE it: have a person write and switch on a rule once, in advance, that states the whole condition in one English sentence (this site, this page, this company size, this campaign, this many people). Then make "no rule matched" the default outcome, so anything nobody wrote a rule for is never acted on. Render the rule back as that sentence in the UI, and create every rule switched off. Separately: in any request/response pipeline where a ledger row guarantees "do this once", that same row will exclude the subject forever if the request fails or is never answered - so roll back on a failed send AND time out unanswered requests, because both failure modes look identical to healthy.
Why: The cost asymmetry is total: a bug in the matcher wastes a credit, a bug in a refusal cold-mails a client's staff or an email-security vendor and cannot be undone with a patch. Making the rule the unit of approval keeps the speed the business wants while keeping a human accountable for the decision, and making no-rule mean silence guarantees the blast radius can only grow deliberately. The ledger-row trap is the same shape as a Clay HTTP column defaulting to GET and 404ing for a day: a system reporting success while doing nothing.
Failure mode: SUCCESS: Outreach - auto-enrolling strangers into cold email is only defensible if the judgement moves to a rule approved in advance, and the default is silence
Clay's HTTP API JSON body editor auto-pairs every quote and brace, so authoring more than one field in it reliably produces corrupt JSON - four attempts, four mangled bodies. Its query-parameter rows are plain text inputs that accept a column token cleanly. Rather than keep fighting the editor, make the receiving endpoint read fields from the query string as well as the body (body wins where it has a value, so a per-row answer always beats a static parameter). Two Clay UI facts worth keeping: `+ Add column` in the grid is a DIFFERENT menu from the `Tools` panel and only the grid one actually selects an enrichment; and to clear that body editor use cmd+A then DELETE - BackSpace only removes the current line no matter how many times you press it.
Why: Hours can disappear into a hostile third-party editor. The receiver is code we own and can test; the vendor's editor is not. Moving the awkwardness to the side you control turns an unreliable manual configuration into something a test suite covers. The guard that matters is unchanged either way - the verdict is still normalised server-side, so this bought convenience and not trust.
Failure mode: SUCCESS: Outreach - when a vendor UI cannot express the shape you need, change the receiver to accept the shape the vendor CAN express
When David has already given the direction and the safety posture (here: build it, autoSend off so he can review), finish the whole job and report once. Create the records, populate them with real starter content, switch on what is safe to switch on, and flag the judgement calls in the final summary rather than pausing for each one. Reserve a mid-task check-in for a decision that is genuinely unsafe or useless to guess at - not for choices with an obvious default like which campaign copy to seed or which sending domain to attach.
Why: Every pause costs David a round trip and reads as stalling, especially on a build he has already scoped and said he wants finished today. The approval he actually wanted was on the OUTPUT (the drafted emails, the enrolled names), not on each configuration step along the way. Deferring cheap decisions upward is not caution, it is offloading work back onto the person who asked for it to be done.
Failure mode: David: "you are stalling each step of the way" - after each build step I stopped and handed a decision back instead of finishing, turning one task into many round trips
When shipping something fast that a person is not yet confident in, give it its own page rather than blending its numbers into dashboards that have earned credibility, and build that page around named failure conditions instead of totals. Lead with plain-sentence warnings, then show a funnel with the DROP named between each stage (five bare numbers hide the only interesting fact), then the refusal reasons. Write each warning to distinguish BROKEN from merely QUIET: "a rule matched nothing while traffic was arriving" is a real alarm, "a rule matched nothing because nothing arrived" is a new system on a slow day. Test every warning as a PAIR - the condition fires, and the innocent version of the same shape does not.
Why: Every failure mode in a multi-hop pipeline across two systems is silent and renders as a healthy zero: a misconfigured HTTP method 404s forever while the vendor UI still says "waiting for data". Totals cannot show that; only a stated expectation can. And a warning that fires on innocent conditions is worse than no warning, because it trains the reader to ignore the box that will eventually hold the real one.
Failure mode: SUCCESS: Outreach - a dashboard for a system nobody trusts yet should be built around what is BROKEN, not what is happening
When a dashboard reports on an automated system that makes DECISIONS about people or accounts, list the subjects by name with the verdict on the same row, not just totals. "12 identified, 2 enrolled" cannot be checked by anyone; "hitone.com - passed over - client or never-contact domain" can be agreed or disagreed with. Also keep states distinct that a count would merge: a record with no decision yet ("not looked at yet") means the queue has not drained, while "passed over" means a decision was already made - identical in a total, opposite in meaning. And never let "enrolled" read as "mailed": show the send status explicitly.
Why: The reason to watch a newly built automated system is to check whether its judgement matches yours. A total cannot answer that question no matter how many totals you add. Naming the subject and the verdict together is what turns a monitoring page into something that can actually catch a wrong decision before it becomes an email to the wrong person.
Failure mode: Built a monitoring dashboard showing only counts. David: "I cant tell what companies or who was added, that is not a full system."
1:1 prep briefs must lead with strategic material the seat owner can't see alone: live financial data (AR aging, P&L trend, margin), patterns across items (churn clusters, cost creep), exposure planning, and the person's capacity/ownership growth. Routine open threads go in a short "housekeeping" footer at most.
Why: A 1:1 is David's scarcest leadership time. Rehashing what the team member already tracks wastes the meeting; the value is the cross-cutting picture and forward decisions only the CEO-level view can bring.
Failure mode: Radar's 1:1 prep brief for Janine listed only transactional low-lying fruit (open email threads, overdue todos, billing housekeeping) — items the team member already knows about and handles daily.
GHL marketplace app name, app type, and Target User are ALL immutable after creation. Get all three right at creation time; a wrong name is permanent short of recreating the app and re-installing every location. When docs don't confirm editability, say "unverified" first, not "yes, editable."
Why: Relay now permanently ships as "Replay" on the sub-account app across 60 HBFG installs, and Jobber will stamp that name into Lead Source on app-created clients. Naming at app creation is a one-shot decision.
Failure mode: Claimed the GHL marketplace app display name was editable after creation ("the display name is an editable field") — David tried and the field is locked, same as app type and Target User.
Clay's JSON body editor is reliably automatable for exactly ONE field: clear with cmd+A then Delete (BackSpace only removes a line), then type `{"email` -> Right -> `: "` -> `/` -> pick the column. Any attempt to add a second or third key by repositioning the cursor corrupts it. Clay's Query parameter ROWS are never transmitted, and column tokens do NOT render in the Endpoint URL either - only static text in the URL works. So: static values go in the endpoint query string, exactly one per-row value goes in the body, and anything more must be typed by a human. Stop after two failed attempts at a hostile third-party editor and hand over the exact text to type.
Why: A vendor UI that fights automation will consume unlimited turns and can leave a working integration in a worse state than it started. The user's time is better spent on twenty seconds of typing than on watching repeated failed attempts, and leaving a broken column silently posting nothing is a worse outcome than a partially-featured working one.
Failure mode: Spent far too many attempts trying to author multi-field JSON in Clay's HTTP API body editor via browser automation; it auto-pairs quotes and braces and left the column in an invalid state that posts nothing
Three portable discoveries: (1) tldraw sync clients treat ONLY WebSocket close code 4099 as fatal — any other code (4401/4404) makes the client silently reconnect-loop instead of surfacing the error, so auth denials on sync sockets must close 4099. (2) The EOS-marks guard regex matches the word "ids" (IDS mark, case-insensitive) inside tool descriptions — write "a board's id", never "board ids", in any Ollie tool copy. (3) A debounce that re-arms on every change never fires under continuous editing — always pair a debounce with a max-wait ceiling and a shutdown flush when it guards persistence.
Why: The 4099 and debounce-starvation defects both passed typecheck, build, and 4363 green tests — only an adversarial review pass against the running behavior caught them. They are silent-data-loss and silent-error classes that will recur in any future realtime OTP feature.
Failure mode: SUCCESS: Claude (OTP dev) shipped the full OTP Whiteboard (board.orgtp.com) in one session — tldraw multiplayer canvas, meetings attach, three Ollie surfaces — via 16 parallel fresh-context agents in dependency waves, with an adversarial review pass between build and ship.
Before shipping any SDK with a commercial license tier, read its license enforcement CODE (grep the installed package for license/gate/watermark), not the marketing page, and smoke-test on a production hostname, not localhost. For tldraw specifically: production requires a license key passed as the licenseKey prop (TLDRAW_LICENSE_KEY env -> page -> Tldraw prop, now plumbed); a 100 day free trial key is self-serve at tldraw.dev.
Why: Dev-mode license exemptions make this failure class invisible to every local check: typecheck, tests, boot smoke, and localhost browser testing all pass while production blanks for real users 5 seconds after load. The delayed teardown also masquerades as a rendering bug, which cost a misdiagnosis round through CSP first.
Failure mode: Shipped the OTP whiteboard believing unlicensed tldraw in production only shows a watermark. Wrong: tldraw 5.x LicenseProvider HIDES THE ENTIRE EDITOR 5 seconds after load on any production host without a license key (state unlicensed-production, LicenseGate display:none div). David hit a blank board twice. It never reproduced in dev because localhost/http counts as development and gets ALL_FEATURES.
When Clay data must reach another system in bulk: (1) Clay's workbook UI has NO CSV export — don't hunt for it. (2) Use Clay's HTTP API action column instead, with all values in the Endpoint/Query-parameter fields (its JSON body editor mangles multi-field bodies and arrays); teach the receiving endpoint to accept one record per request via query string, mirroring /intake/company-contacts. (3) Feed companies in bulk by POSTing directly to the existing prospecting webhook — but expect Clay to 429 its own webhook under concurrent action-column bursts; repeated 'Run N empty or out-of-date rows' passes converge (~100/pass). (4) Action columns do NOT auto-run on newly arrived rows — run the column manually after each wave or new data silently sits unpushed. (5) Clay Find People mis-resolves small business names to famous corporations (~25% wrong-company rate) — always classify post-import (icp=in/wrong-company/review), never trust the resolution.
Why: This turned a one-off 5,000-prospect ask into permanent infrastructure: any person Clay finds now auto-flows into Outreach validated and gated. The five gotchas each produced silent failure modes (invisible 429s, unpushed rows, wrong-company contamination) that would have burned credits and polluted audiences on every future run.
Failure mode: SUCCESS: Outreach — imported 892 validated gym decision-makers from Clay with no CSV export and no manual re-entry, as a standing pipeline
When a rank tracker project reports refresh_blocked or last_successful_refresh older than the requested window, re-query with period2 ending on or just after last_successful_refresh (here period2 2026-08-05 to 2026-08-13 returned the full keyword set). Always read last_successful_refresh and refresh_blocked from the response and stamp the refresh date on the card. Sign convention confirmed: avg_position_delta = previous minus current, so positive means the keyword improved. Also: the CCM Project Stats tab reports a nominal size of 5,504 rows but data ended near row 4,450; read the header first, then probe from the end backwards rather than trusting the sheet size.
Why: A zero result from a quota-paused tracker is not "no rankings"; shipping it that way would strip the SEO section from six client cards and tell coaches the keywords vanished. The generator (~/.claude/gen-coach-report.py) is now the stable path; each run only edits the config block, VENT purchases, last-meeting dates and the notice text.
Failure mode: SUCCESS: Dash /coach-report 2026-08-31 shipped 55 cards with all sources live, and caught a Search Atlas rank-tracker trap: with the plan quota exhausted (refresh_blocked=true since Aug 12), keywords-details returns count=0 and an empty results list when period2 is set to the current week, which reads as "no keywords tracked".
Generating the prior Delta Meeting's Ollie Insight is part of Dan's prep, not a gap to report. In prep: trigger insight generation for the last completed meeting (find or build the API path behind the Generate button; if truly UI-only, ask David to click it at WEDNESDAY prep, never on meeting morning), then read the aiSummary into section 1. The meeting never opens with section 1 empty.
Why: Section 1 of the contract exists so the meeting starts from Ollie's read of last week, not Dan's paraphrase. Reporting the gap instead of closing it puts prep work on David at the table, which is exactly what prep exists to prevent.
Failure mode: L10 prep reported the prior meeting's Ollie Insight as a gap ("Generate button never clicked, no CLI read path") and opened the meeting without it. David had to generate it himself mid-meeting.
Never present a queue as approval-blocked without checking send velocity first. queued high + sent ~0 over multiple days = blocked; queued high + sent at steady weekly volume = pacing by design (domain warm-up caps). Also: a Monday-7am "this week: 0" is the week-reset clock, not a stall - autoSend fires at 9am.
Why: Framing a healthy, maximized pipeline as a David-bottleneck puts a false decision on his plate in the meeting and misreads the one system that is working exactly as built.
Failure mode: L10 signal A presented 1,285 queued Sneeze It emails as "idle behind David's approval click." The live stats show 1,383 sent LAST week against that queue - the blasts are approved and draining at the deliverability pace. The "awaiting approval" footnote in outreach-engine-latest.md described the 8/20 state and went stale.
To reach the OTP production database from a laptop, use: `railway run -p 6e2dde6f-7fcd-406a-a98a-b8cc6aa9c53c -e production -s Postgres npx tsx scripts/<x>.ts`. The project ID is `fortunate-commitment` (Postgres lives there, not in otp-platform) and `-e production` is mandatory whenever `-p` is passed. Scripts that need this should read `process.env.DATABASE_PUBLIC_URL || process.env.DATABASE_URL` and build their own pg.Pool rather than importing src/config/database.ts, which hard-requires DATABASE_URL.
Why: 146 scripts carry a documented invocation that does not work off-Railway. During an incident with a 3-hour customer-notification clock, a responder burns minutes on two failing commands and a code-archaeology detour before they can query anything. Now documented in docs/incident-response-runbook.md step 2.
Failure mode: SUCCESS: SOC 2 G7 incident-response tabletop found that the production database is unreachable from a laptop via the command every otp-platform script documents. `railway run npx tsx scripts/<x>.ts` fails with ENOTFOUND postgres.railway.internal (the app service's DATABASE_URL is a private-network hostname), and the fallback documented in scripts/read-support-board.ts (`railway run -s postgres`) fails with "Service not found" because Postgres lives in a separate Railway project.
Before announcing any feature to customers, verify three things separately, never inferring one from another: (1) the code is MERGED, (2) the deployed commit sha actually carries it (check /health), and (3) the feature RENDERS for a real user, including any third-party license key, CSP allowance or env var its vendor requires in production. Check every URL in the announcement resolves; a hostname that appears in a commit message is not proof it is live. When a surface is not reachable yet, route the copy to the one that is.
Why: A merged PR, a green CI run and a deployed sha all say the feature exists. None of them say a customer can use it. Vendor-side gates like a license key fail silently and only in production, which is exactly the condition a broadcast walks 98 people into at once. The gap between "shipped" and "usable" is where an announcement turns into a support wave and a credibility hit.
Failure mode: SUCCESS: Swamp caught that announcing a shipped feature is a different claim from announcing a WORKING one. Issue #35 led on the new multiplayer Whiteboard. The code was merged and deployed (prod /health on the exact HEAD), but unlicensed tldraw 5.x HIDES THE EDITOR five seconds after load on production hosts. Had TLDRAW_LICENSE_KEY not been set, the email would have sent 98 customers to a canvas that goes blank. Separately, board.orgtp.com was in the feature's own commit message but returns 000 (DNS not live), so it was kept out of the copy entirely.
Build the env file in two steps rather than one: dump the app service vars while EXCLUDING DATABASE_URL, then append DATABASE_PUBLIC_URL taken from the Postgres service (not from otp-platform, which only carries the internal host) renamed to DATABASE_URL. Verify before running that the resulting DATABASE_URL does not contain railway.internal. Delete the env file immediately after the send; it holds live production credentials.
Why: Issue #34 lost time to this and ended with the operator running the command manually. The public URL lives on a different Railway service than the one the app vars come from, which is the non-obvious part: querying the app service alone will never surface it. This turns a blocking failure into two lines of setup for every future weekly send.
Failure mode: SUCCESS: Swamp resolved the recurring prod-DB blocker that stopped the issue #34 run and forced David to execute the dry run by hand. The prod DATABASE_URL on the otp-platform service resolves to postgres.railway.internal, which is unreachable from a laptop, so the documented env dump alone fails at gatherSubscribers.
(1) tldraw core ships only the commenting license hook and toolbar item; pins, popovers, composer and the CommentTool live in @tldraw/commenting, which must be installed and wired (tools=commentTools, overrides=[commentToolOverrides], InFrontOfTheCanvas=CanvasComments, import commenting.css). Records syncing is not proof the feature works; check what draws them. (2) Any standalone signed-in page rendered with ejs.renderFile must load clerk-js as a session keeper, because the app layouts are what refresh the 60-second Clerk session cookie; without it the first fetch works and every later one 401s. Add a 401-refresh-retry around the page's API calls too.
Why: Both failures looked like license or auth problems and were neither. Reading the vendor's own package split and the layout's clerk loader found the real causes in one pass; the previous PR had shipped comments believing the license flag was the only gate.
Failure mode: SUCCESS: Dan diagnosed two whiteboard bugs from one screenshot: comments saved but invisible, and AUTH_REQUIRED on a signed-in board a minute after opening
When a human confirms a system is working (Bogdan 8/31, Kristen 9/1, David 9/2 'good to go'), close the flag that morning and remove the check from source-health, rather than carrying a 'softened' version. A missing log is only a flag until someone with eyes on the system says it is fine; after that it is noise.
Why: Every flag on the morning board costs David attention. A flag that survives human confirmation trains him to skim the Watch line, which is the one line that must stay credible for real dark sources (Meta token, CCM feed).
Failure mode: Dan kept carrying 'job:outreach STALE / no run log' in the Watch line for six mornings after David and Bogdan had already confirmed the outreach sending was live and healthy. The flag was measuring missing observability, not a real problem, and David had to tell Dan to stop.
When a design fix is CSS-only and the local otp-platform dev server cannot start (needs DATABASE_URL; port 3000 is taken by another app), open the live page in gstack browse, inject the edited rules as a style element via $B js, and screenshot before/after. Page-level inline style blocks in the body beat head injections, so append the override to the body or add !important for the verification pass only. Then run scripts/design-lint.mjs and the view lint tests before opening the PR.
Why: It gives real before/after evidence on real content in minutes, with no database, and catches collisions the source diff hides (a blue primary on the blue compare band). The Sep 2 sweep shipped four CSS fixes to production on this loop with CI green first try.
Failure mode: SUCCESS: Conatus verified CSS-only design fixes against the live orgtp.com pages without a running dev server
David: "no black buttons, no reason for it." OTP has one filled button colour, the blue --primary. Black ink fills are not a secondary style; a quieter action is a ghost/outline or a text link. On coloured bands (blue or orange) the filled button inverts to surface-on-colour so it stays visible.
Why: Black filled buttons add a second primary weight that carries no meaning, which breaks the "colour equals meaning" and "one primary per screen" rules in src/DESIGN.md. David reads the lime-plus-black combo as something he liked but accepts losing for a single coherent system.
Failure mode: In the Sep 2 orgtp.com design sweep I kept the black .btn-ink buttons (pricing Get Started / Upgrade / Send Inquiry, home Keep me posted, the orange band CTAs) as a second filled button colour beside the blue primary.
When asked to clean up a site's design: (1) crawl the sitemap, every single route plus two samples per templated family, running a computed-style scan in the page (button fill colours, radii over the token scale excluding true circles, 999px pills, resting shadows, off-token large backgrounds, em dashes, sub-12px text, horizontal overflow, heading fonts) and a viewport screenshot per page; (2) aggregate into one table so the patterns, not the pages, are visible; (3) fix each pattern once at its source: alias stray tokens to the real ones in the shared stylesheet, clamp utility classes globally, and run an idempotent transform over page-local style blocks; (4) verify by rendering every page locally with the layout and no database (ejs.renderFile with stub locals), serving the static output, and re-running the same scan to get before/after numbers. Tooling kept in otp-platform/scripts/design-crawl.
Why: A sampled sweep of 8 pages missed pages carrying their own colour systems (David: "the scan design skills don't go far enough through the site"). Measuring all 90 pages found 235 off-scale radii, 124 resting shadows and 30 non-primary button fills; one transform brought them to 2, 0 and 10 in a single PR with CI green, where per-page edits would have taken days and drifted again.
Failure mode: SUCCESS: Conatus converged the orgtp.com marketing site on one design system by crawling every page and fixing patterns at their source instead of page by page
When a pending_transcriptions row is 'fatal' with last_error 'transcription failed', do not trust the label: download the R2 segments with the app's own storage client (railway run -s otp-platform with DATABASE_URL overridden to the Postgres service's DATABASE_PUBLIC_URL) and try Deepgram on each segment to get the provider's real error. A WebM segment recorded after a pause can lack the EBML header (bytes 1a45dfa3) and is rejected as 'corrupt or unsupported data' while concatenating it onto the header-carrying segment transcribes cleanly. Fix the worker (coalesce header-less continuations, surface the provider error), deploy, then requeue the row (status pending, attempts 0) so the production worker attaches the transcript through the real pipeline; verify transcript length on the meeting before closing the ticket.
Why: The audio was never lost, only mislabelled; the 'transcription failed' string hid a specific, fixable cause and a customer waited a week. Diagnosing against production storage with the real provider took minutes and turned a manual recovery into a permanent fix (PR #683).
Failure mode: SUCCESS: Conatus recovered a customer's 97-minute meeting recording that the transcription worker had marked fatal, and fixed the cause so it cannot recur
When the system can already enumerate the valid choices, render a dropdown of them, exclude ones already claimed elsewhere, and preselect the likely match from the name. Every bind surface gets a matching unbind/disconnect on the same listing, so a claimed choice can be released without hunting for the other side.
Why: David: "binding the call tracking is very difficult ... instead of CallRail Company ID why not a dropdown with all the locations in CallRail that can be connected? and if connected to one it is not available to another unless you disconnect, in fact all the bind buttons on the account should have a disconnect as well." A free-text id field pushes lookup work onto the operator and invites one-character mistakes; the list was one API call away.
Failure mode: Relay's CallRail binding step asked the operator to paste a CallRail company id (COM...) into a text box, with a hint to go find it on another page or in CallRail's URL bar. Relay already had the full company list from the API.
Before building a "trial" or "offer" email, grep the wallet/billing code for what every new account already receives (OTP: seedSignupCredit gives every org $25 at signup). Sell the existing thing with attribution (utm_campaign, utm_content per rung) instead of inventing a coupon table. Reuse the existing series renderer (LifecycleEmail shape + renderLifecycleEmail) so unsubscribe scope, footer and brand come free; add only the rung data, a unique (subscriber, rung) send log, and a weekday tick. When two automated programs can reach the same inbox, give the newer one an explicit window (PRESIGNUP_SEQUENCE_WINDOW_DAYS) and make the older one wait for it.
Why: The form had been posting to partner_signups for months with zero rows arriving and nothing sent back; the partial's comment claimed the opposite. Checking what the form actually hit, and what the product already gives away, turned a "build a credit system" request into a four-file change shipped the same evening with a live proof row.
Failure mode: SUCCESS: Claude turned the home "Keep me posted" form from a dead lead endpoint into a pre-signup sequence built on credit that already existed
Two reusable patterns. (1) When a public loginless page must call a paid model, key every per-IP limiter on the RIGHTMOST x-forwarded-for entry (the one the trusted edge appends), not request.ip: under Fastify trustProxy:true request.ip is the client-supplied leftmost entry and any limiter on it is defeated by one header. orgtp.com is served directly by Railway's edge (no Cloudflare), so there is no CF-Connecting-IP to lean on. Add a global in-memory daily ceiling as the real spend cap. (2) tldraw's editor.toImage() inlines page fonts by fetch()ing the Google Fonts stylesheet and loading faces as data: URLs; a CSP without fonts.googleapis.com/fonts.gstatic.com on connect-src and data: on font-src makes every export silently lose the handwriting font. Also: when another session holds the main checkout on its own branch with uncommitted work, build in a dedicated git worktree with node_modules symlinked rather than sharing the tree.
Why: The door is an unauthenticated endpoint in front of a paid model, so a spoofable rate limit is a direct cost exposure, and the same request.ip weakness exists in every other limiter in the app. The font CSP gap had been degrading the signed-in board's PDF/PNG export since it shipped without anyone noticing. The worktree pattern prevented two live sessions from clobbering each other's uncommitted files.
Failure mode: SUCCESS: Claude (Conatus) shipped orgtp.com/draw, the loginless whiteboard door for OTP outreach, in one evening: three parallel builders in fresh contexts, one adoption wave, one independent review that surfaced 13 real defects (11 fixed before merge), PR #688.
When adding native or WASM media dependencies (sharp, heic-convert, pdfjs-dist, @napi-rs/canvas) to OTP, build the Dockerfile's deps stage locally (docker build --target deps) and run the new unit tests inside that container with the repo src mounted, before opening the PR. Alpine (musl) and the image's Node version differ from the Mac; the lockfile must carry the linuxmusl binaries and the runtime must have the builtins the library assumes (pdfjs 6 needs Promise.withResolvers and ArrayBuffer transfer helpers, so Node 22). Bump Dockerfile and CI Node together.
Why: All 16 conversion tests passed on the Mac (Node 25) and would have passed CI, yet on the node:20-alpine image the PDF path threw once and, with a polyfill, rendered pages with errors swallowed as warnings. Only the in-container run exposed it. Ten minutes of Docker saved a broken feature in production.
Failure mode: SUCCESS: Claude shipped photo/voice to whiteboard import in one evening by running the conversion tests inside the production Docker image before merging, which caught that pdfjs-dist 6 silently drops PDF content on Node 20
Never put a real client, a realistic account number or a person's name in a form placeholder. A placeholder reads as prefilled data to the person filling it in, and another customer's name on a form is a confidentiality smell. Leave the field blank and put format guidance in the hint text under the field. Grep views and route templates for client names before shipping any customer-facing form.
Why: Prospects judge the product by its forms. Example values that look like data make the form feel used, and leaking a paying client's name onto a stranger's screen is a trust problem, not a styling one.
Failure mode: iJoin's public sign-up form and the club authorization card used a real client's name (Proof Fitness), realistic ABC club numbers (04462, 04431) and a person's name (Alex Newman) as input placeholders. David: "remove the prefilled legal name Proof Fitness and prefilled club numbers and prefilled who will sign please, this is all bad form to have that in there."
On anything addressed to ABC Fitness, the vendor is "Sneeze It" and the signer is "David Steel". The legal-name rule (David Sieradzky, The Steel Method LLC) applies to the LLC's own legal matters, not to vendor forms with partners who know us as Sneeze It. Defaults live in src/platform/abc-form.ts (VENDOR) with ABC_VENDOR_* env overrides.
Why: ABC's records and the Data Transfer Agreement are under the Sneeze It name; a different entity string on the release form makes ABC reconcile two names and can stall the release.
Failure mode: On ABC's Club Data Release form I filled the vendor as "The Steel Method, LLC d/b/a Sneeze It" signed by "David Sieradzky", reasoning from the legal-entity rule. David corrected: vendor is "Sneeze It", page two vendor block is "Sneeze It, by David Steel", with the date.
When David points at a page he likes as the reference, match its actual visual system (light ground, soft colour, rounded white panels, generous air), not just its energy or copy. For a logo, do a real pass: a mark that carries one clear idea (for iJoin, two things joining, or walking in through a door), built from a small number of geometric primitives, shown as proper lockups (horizontal, stacked, favicon, on white and on navy), with the wordmark set in the brand face with tight tracking. Never present placeholder shapes as options.
Why: A reference page is a spec. Reading it loosely wastes a round trip on the one thing David could see immediately. And a logo is the brand's face; three throwaway marks read as not caring, which costs more trust than presenting one careful option.
Failure mode: Homepage v2 for ijoin.ai used a dark navy "sky" hero when David had praised the LIGHT styling of /pixel (off-white ground, soft aurora, white rounded panels). And the three logo directions were weak: generic geometric shapes and a smiley ball, thrown together without a real design pass. David: "the logo well that just is awful you can do so much better."
The base David quotes for the call center is the COMBINED figure for both callers, not per head. Split a combined base across the roster by days worked; never multiply it by headcount. Corrected margin is +4.1%. Also, when the user gives both a pay rate ($5/hr + $5/appt) and observed pay ($700/$450 a week), reconcile them before building - they may not agree.
Why: Treating a team-level cost as per-head inverted the conclusion of a live pricing decision for a 4-location client program, turning a thin-but-positive model into an apparently fatal one.
Failure mode: Modelled the $3,000 caller base as PER CALLER and pro-rated it by days worked, which doubled the fixed cost and made the base-heavy comp model show a -41.5% margin.
A report answers three questions in order: is the instrument reporting at all, is this better or worse than the window before, and what do I do in the next ten minutes. Lead with blind spots (a 0 from a disconnected gauge must say so, not display as a result), then blockers, then unworked engagement, then copy evidence. Every recommendation carries the measurement that triggered it and a link to the exact list of people, never an adjective alone. Check three traps before ranking anything: attribution (a visit before the send, or from a de-anonymising pixel, is not that campaign's result), selection (a retargeting campaign mails people because they already visited, so it scores 96% against a cold list's 18% and must be ranked only against its own kind), and measurability (one campaign linked only to Loom, scored 0% of 810, and would have been labelled weak copy when nothing in it was observable).
Why: One confidently wrong recommendation costs the page every other recommendation it makes: David stops reading it and the true findings are wasted. Rebuilt on live data the same page surfaced reply capture silent across 10,613 delivered, both brands under one day of unmailed audience left, two blasts held down since Aug 31, and 1,716 hot contacts against one booking in 30 days.
Failure mode: The Outreach Engine's /dash/reports page was a table of campaigns with sent, delivered %, bounce %, on-site, replied and booked. Every number was individually correct and David's verdict was "this tells me nothing". Replied and booked read 0 on every row because reply capture is disconnected, so the two columns billed as the ones that matter were dead everywhere and the page could not change anybody's behaviour.
When a test fails with no code change behind it, diff the failure against the CALENDAR before calling it a flake: hardcoded dates plus a relative rule is date rot, and it is a real defect in the guard even when the production logic is correct. The fix is not to bump the date, which only resets the timer on the same failure. Inject the clock into the rule (default it to new Date() so every production caller is unchanged), thread the caller's own now where one already exists, and pin the regression at day 0, day (window - 1) and day (window + 1) so the answer depends only on the gap between the event and the question. Use the ORM's typed comparison (drizzle gt(col, date)), never a JS Date inside a raw sql template.
Why: A guard that rots stops guarding silently, and the failure looks like noise rather than a defect, so it gets ignored on exactly the sends where it matters. Here the guard is what stops a second campaign re-mailing the whole list days after the first one, on a programme currently sending 3,250 a week. "Pre-existing, not from this change" is a diagnosis, not an excuse to skip the diagnosis.
Failure mode: A guardrail test protecting the 14-day contact cooldown ("does not re-mail the people the first campaign already reached") started failing with no code change. It stamped its sends at a hardcoded calendar date while the rule itself was evaluated against SQL now(). A relative rule measured against an absolute date passes for exactly the length of the window and then fails every day afterwards. The first instinct, and the one I reported before checking, was to call it a pre-existing flake and leave it.
Always set CENSUS_API_KEY when building on api.census.gov (free, instant, https://api.census.gov/data/key_signup.html). A keyless call does NOT return 4xx: it returns a 302 to an HTML page at /data/missing_key.html, so a client that follows redirects and calls .json() dies with "Unexpected token" several frames from the real cause. Fetch Census with redirect manual, check the Location header for missing_key, and sniff the body for a leading angle bracket before parsing. The Census geocoder at geocoding.geo.census.gov is a separate service and stays keyless. Also filter the ACS suppressed-cell sentinels (-666666666 and similar) to null or they poison every median downstream.
Why: It burns time twice: discovering the key is mandatory when every reference says otherwise, then debugging a JSON syntax error that points nowhere near the cause. Any Sneeze It work touching census demographics hits this on the first request. Wider lesson: a 302-to-HTML failure disguises an auth problem as a parse bug, so any integration failing with an unexpected-token error should be checked for an unfollowed redirect before blaming the parser.
Failure mode: SUCCESS: Audience module build. The US Census data API now requires an API key on every request. Widely repeated guidance, including library docs, says it is optional below 500 requests per day. That allowance is gone as of 2026. Verified across ACS years 2021-2024 at both single-ZCTA and national wildcard scope: all bounced.
Any Google Doc containing tabular data must use real Docs tables. The MCP helper create_table_with_data is BROKEN — it creates the table then fails to populate it ("ERROR: Could not find table after creation"), leaving an empty table. Working method: (1) batch_update_doc with insert_table (end_of_segment) to create the empty table; (2) note the table start index S from inspect_doc_structure, or compute it as doc_length_before_insert; (3) populate cells with batch_update_doc insert_text ops applied in REVERSE order (last cell first) so earlier indices do not shift. Cell insert index = S + 1 + (row-1)*(1 + 2*cols) + 1 + (col-1)*2 + 1, with rows and cols 1-indexed. Empty table length = 1 + rows*(1 + 2*cols) + 1. Apply heading styles afterwards with update_paragraph_style using named_style_type; style ops do not shift indices so order does not matter.
Why: Long financial documents built as monospace text are unreadable in Google Docs, which uses a proportional font — columns collapse and the reader cannot scan across rows. The reconciliation is a review document for a bankruptcy with roughly $1.2M at stake; if it cannot be read line by line it has no value. And the standing rule already existed, so this was a preventable repeat.
Failure mode: Alan built the master reconciliation Google Doc using ASCII text-aligned columns instead of real Google Docs tables. David could not read it and asked for it to be reformatted. This also violated an existing standing rule in memory (feedback_google_doc_tables: real tables, never text-aligned columns).
When a deploy appears to have come from nowhere, check `git log --oneline HEAD --not --remotes` and `git status` BEFORE assuming a dirty tree, and verify the running code with a build-time source fingerprint rather than any commit field. `railway up` ships a directory rather than a git ref, so a commit sha on such a deploy can only have been set by hand and will go stale silently. Never set a deploy-identity variable by hand: derive it from the build, or report 'unknown'.
Why: A hand-set identity field is worse than an absent one. An absent field sends you to a real check; a stale field answers confidently and wrong, and here it cost an evening spent hunting a working directory that had never gone missing. The same trap exists on any service deployed by CLI upload rather than by git trigger.
Failure mode: SUCCESS: a deploy that looked like it came from an unpushed working directory was actually a clean, committed, merely unpushed tree. The evidence that misled us was /health reporting commit 1a4a08e (Aug 15) beside a build twenty minutes old. The stale value came from GIT_COMMIT_SHA, a Railway variable set once by hand with `railway variables --set` and never cleared, so it reported August through every deploy after it.
When a competitive scan cannot reach a vendor's changelog, roadmap or release notes, NEVER report that as "they shipped nothing" or "low threat." A 301 to a dead page, a 404 or a 403 is a CRAWL FAILURE, and it looks identical from the outside to genuine absence. Re-run with a real browser (Playwright or claude-in-chrome) and navigate the site's own UI rather than guessing URLs, before drawing any conclusion about a competitor's activity. Vendors commonly host changelogs on a subdomain (changelog.vendor.com) linked only from a help-center header, so the obvious path (/product-roadmap/, /changelog) can be dead while the real one is actively maintained. Also treat G2's 403 and search engines drowning a query in a same-name distractor (EOS the cryptocurrency vs EOS the business framework) as access gaps to be stated explicitly, never as evidence of absence.
Why: On 2026-09-06 the first OTP competitive scan concluded Bloom Growth was "not a credible AI competitor at all" because WebFetch hit a dead 301 at /product-roadmap/ and a 404 on the help page. A live browser pass found changelog.bloomgrowth.com actively maintained with entries every 1-4 days, including a proactive "Meeting Prep" feature shipped 2026-08-28 that surfaces off-target KPIs and slipping priorities before a meeting and refreshes itself ahead of it. That is the single closest competitor feature to OTP's own core differentiation claim, and it was 20 minutes away from being recorded as a competitor with no AI story. A wrong competitive conclusion is worse than a missing one because it gets built into positioning and nobody re-checks it.
Failure mode: SUCCESS: CI_SCOUT — a failed WebFetch was nearly reported as a competitor shipping nothing
Before reporting any KPI as simply above or below goal, decompose it by its own components at least once per quarter and ask what behaviour the tile REWARDS, not merely what it measures. A tile whose goal points the same direction as a legal, contractual or strategic constraint we are trying to shrink is worse than a missing tile: it pays us to make the problem bigger. Re-verify goal DIRECTION whenever strategy changes, not just goal value. Separately: never write an identifier (project key, board id, account id) into a spec from an assumption. Query it once and record the response, because a wrong identifier returns zero rather than an error, and zero reads as good news.
Why: Both findings share one shape: a number that looks fine, or looks simply low, while the mechanism underneath points the wrong way. Staleness assertions catch a source that stopped updating; nothing catches a source updating correctly toward the wrong outcome. The Beacon tile was wired 2026-08-03, five weeks AFTER the 2026-06-29 EOS cease notice, and nobody re-asked whether ranking harder on those terms was still wanted. Decomposition is cheap and it is the only thing that surfaces this class.
Failure mode: SUCCESS: Dan found that 92.5% of the Beacon KPI (979 of 1,058 pillar impressions) came from EOS-trademarked clusters, the exact marks under a cease notice whose compliance deadline had already passed. For five weeks the tile was read as a simple miss (1,058 vs 3,500 goal = "SEO is behind") and never decomposed. The same prep found Crystal's spec asserted Jira project key "CP" from an unverified assumption; that project does not exist, so the query would have returned a clean zero forever.
When a churn tripwire fires on many accounts at once in the same week, do NOT report it as N client-level alerts or as one uniform market/holiday effect. Pull `time_increment=1` daily insights for 3-5 representative accounts plus `account_status`, `spend_cap`, `amount_spent` and campaign `effective_status`. On 2026-09-07, 18 of ~45 accounts tripped the -20% wire; the daily pull showed three unrelated causes: WOA Hickory had a 15x spend spike on Sep 1-2 ($456/$462 vs a $30/day baseline) that burned its entire $996 account spend cap, after which Meta paused both campaigns and the account went dark from Sep 3 (invisible in the 7d rollup, which just looked "down"); WOA Lakewood Park held spend at $66/day with 2 leads in 7 days and campaigns ACTIVE, a tracking or form break; WOA Kettering simply had its daily budget cut to $50 from ~$78 on Sep 1, so its drop was explained and needed no action. Also verify the window: Meta `date_preset=last_7d` and `last_30d` both EXCLUDE today, so a broad decline is never a partial-day artifact and that hypothesis can be dismissed by reading the script rather than guessing.
Why: A rollup metric hides the mechanism. Reporting 18 near-identical "down 20%" lines buries a dark account that is losing a client real leads every day, and reporting it as one holiday dip explains away a live tracking break. The daily-plus-account-status pull is about 60 seconds of work and is what converts an unactionable trend line into a named cause with a named fix. An account that hits its spend cap looks identical to soft demand in every weekly view.
Failure mode: SUCCESS: Dash — a portfolio-wide "leads down >20%" pattern was 3 distinct root causes, not one trend, and only a daily-level pull separated them
David's design, 2026-09-07: run Dan as a background agent on every /good-morning, and define prep as MOVING THE WORK FORWARD rather than assembling the report. The pass reads the prior Ollie insight, advances open rock milestones and to-dos, and does NOT spend its time gathering headlines/signals — reporting material can be collected at compile time. Attach the recurring obligation to a habit the human already has, never to a day the agent must remember on its own.
Why: Every Dan commitment backed by a mechanism shipped on time (freshness assertion, run log, Tally refusal). Every commitment left to Dan's own judgement failed. The fix is therefore never a better promise, it is a trigger owned by something else. And prep-as-assembly optimises the wrong half: the brief was never the bottleneck, the undone to-dos were, and a brief compiled from work that did not happen is just a well-formatted miss.
Failure mode: For four consecutive weeks Dan missed "Wednesday prep" and treated it as a scheduling/discipline problem, proposing a scheduled job that writes the brief earlier. That diagnosis was wrong. Dan had defined prep as DOCUMENT ASSEMBLY (gather headlines, read signals, compile the brief), which is genuinely doable Monday morning — today's Monday-built brief found the Beacon trademark issue and the Jira gap. What cannot be done Monday morning is the actual open work: to-do T4 sat undone for a week and was closed in four minutes once attention was forced onto it.
Report cost per acquisition on measurable media only, meaning the channels that can be traced to a conversion, and disclose the all-in figure alongside it in the measurement notes. Here: $42,000 Meta plus Google divided by 836 equals $50.24, versus $53.83 all-in. Remove that channel's clicks from the funnel denominator too, for consistency. Then give the untrackable channel its own reach economics rather than a blank: CPM, cost per completed view, cost per click, cost per thousand reached, cost per viewing hour. If a modelled contribution helps, label it modelled, keep it out of the verified count, and show the working.
Why: David: "I want the number 50" and "use a known stat instead of reporting nothing." Both are the same rule: measure each channel on the metric it can actually be held to. The fix was not a massage, it resolved a real contradiction, and it had a bonus - the following month had that channel paused, so the measurable-media basis made the two months comparable for the first time. A blank in a client report reads as laziness or concealment.
Failure mode: Charged a channel that cannot be tracked to a conversion against the cost-per-acquisition figure, while the same report said that channel should not be judged on the acquisition target. OTT's $3,000 was inflating CPA to $53.83. Separately, the report described OTT as contributing nothing, when it had abundant hard delivery numbers sitting unused.
Test the assumption against historical cohorts in the same dataset before writing it. Method: take month M's non-converting records, match on a stable key such as email, and look for them converting in later months; then compare cohorts of different ages. Result here: abandoned checkouts returned at 0.85% after three months, 0.03% after five weeks, 0.00% after one week, and there was no within-month gradient either, because the newest cohort converted highest. The figure was already final. Then look for where the real uncertainty actually lives: off-platform. Someone who abandons online and joins in person is recorded in the club system, never in the cart, so report the number as a ceiling with a sensitivity table rather than a point estimate.
Why: The instinct that a number is incomplete is usually right, but the mechanism is often wrong. Measuring cost nothing and produced a stronger deliverable than the requested asterisk: instead of asking the client to trust that the number will improve, the report showed that closing 30.5% of the abandoned leads at club level, about one member per club, puts the campaign at target, and named the exact match that would settle it. A sensitivity table beats a fabricated maturation curve.
Failure mode: Was asked to add a footnote stating that a month-end conversion figure would improve over the following weeks as late closes landed, and to estimate the late-close rate. Writing that as an assumption would have put a forward-looking claim in a client document that the client could disprove using their own data.
When parsing a month out of a document title or filename, match full month words explicitly (january|jan, february|feb, ...) with word boundaries. Never use a three-letter prefix plus a wildcard: mar/may/jun/aug/sep all appear inside ordinary English words (Marketing, Mayfair, August, September). Also require a four-digit year in the same string, and refuse to guess when zero or two months match rather than falling back to the current date.
Why: A date parsed off the wrong word is worse than no date: it produces a confident vintage that silently ages or freshens every figure in the file. Here it would have kept budget pacing dark for all 54 Coach clients while appearing to have fixed it, which is exactly the failure the vintage was introduced to prevent.
Failure mode: SUCCESS: Sneeze Coach budget mapper dated the "Active Marketing Clients - Sept 2026" sheet to MARCH, because a three-letter month-prefix regex matched "mar" inside the word "Marketing". Every budget in the file would have been stored with a six-month-old vintage, which switches pacing off silently rather than loudly.
When a data source can return a partial answer, model "we were not told" and "the answer is zero" as two separate fields, never one. In Audience this is GoogleDemand.available (the call worked) vs GoogleDemand.metricsAvailable (the rows carry numbers); the verdict stays 'unknown' when metrics are missing and the panel shows the real keyword list with no numbers instead of zeroes. Probe the live API before trusting a parser: both generateKeywordIdeas and generateKeywordHistoricalMetrics were tested against the MCC and against a client account spending ~$3,000/mo, and both returned text with keywordIdeaMetrics absent entirely, so it is the developer token's access level (Basic; full Keyword Planner metrics need Standard access, an application in the Google API Center) and not account eligibility. Also: do not run drizzle-kit generate in sneeze-audience - migrations there are hand-written incremental .sql files applied in filename order by a ledger table, and the generated baseline sorts before them and would CREATE TABLE over a live database on every boot.
Why: A zero that means "unknown" is indistinguishable on a page from a zero that means "no demand", and everything around it on the report is real, so it gets believed. The two readings lead to opposite actions: apply for a token, or move a client's ad budget off Google. The same shape recurs anywhere an API degrades partially rather than failing loudly.
Failure mode: SUCCESS: Audience - Google Keyword Planner returns keyword ideas but ZERO metrics on a Basic-access developer token, and parsed naively that renders as "thin search demand, move the budget elsewhere"
Two rules. (1) LOG THE WEEK AS ITS OWN STEP, with an explicit "deliberately NOT logged" list written into the changelog file. PR #716 did this on 9/6 and it turned issue #36's reconstruction from ~55 PRs into a bounded 9/4-9/7 gap, because the prior wave's editorial decisions were recorded rather than re-litigated. Six straight issues had undercounted before this. (2) NEVER ACCEPT A CLEAN COMPLIANCE SCAN FROM ONE CASING. Run the trademark/vocabulary scan case-insensitively AND on the parsed entries. On #36 the capitalized scan returned clean while the case-insensitive pass found "headline" twice in an unsent entry: EOS agenda vocabulary and the wrong product noun (it is a Signal).
Why: The undercount cost five previous issues either a re-send risk or a scramble to reconstruct dozens of PRs under time pressure. Logging the week when it happens makes the weekly email a read rather than an archaeology dig. And a single-casing compliance scan is a false negative that ships a trademarked term to 98 customers while reporting itself clean, which is worse than no scan at all because it manufactures confidence.
Failure mode: SUCCESS: Swamp issue #36 -- the weekly changelog undercount pattern finally broke, and a compliance mark slipped past a capitalized-only scan for the second time (the issue #30 lesson).
To restyle an existing Google Doc whose URL is already registered somewhere, do NOT use import_to_google_doc (it mints a new URL). Use update_drive_file with file_path + source_format:"html" — it replaces a native Doc's content in place, keeping the file ID, URL, sharing, comments, and version history. Stage the HTML in ~/.workspace-mcp/attachments/ and pass file_path, never inline content. Two conversion gotchas: paragraph background-color is dropped by Drive's HTML converter, so put callouts, pull-quotes and change-log blocks in a single-cell table with the shading on the td; and use real h1/h2/h3 tags (inline CSS still applies) because they map to Docs' named heading styles and populate the document outline pane. On a long document verify content preservation mechanically — diff the source word sequence against the tag-stripped HTML with difflib before publishing — and publish to a throwaway doc first, look at it in a browser, then apply in place.
Why: The doc was 94K characters of hard-wrapped plain paragraphs with ASCII rules and no headings, and it is a financial record where an accidental edit to a figure would be worse than bad formatting. The word-level diff proved nothing was lost, the trial import caught two real defects (a swallowed "Total assets" line and a column header wrapping), and the in-place update preserved the URL that CLAUDE.md and Alan's spec both point at.
Failure mode: SUCCESS: Alan — reformatted the 46-section Bankruptcy Reconstruction Google Doc in place without changing a word or a number
Match a duplicate-person check on name by edit distance as well as on email, gating a near-surname match behind a matching forename so families do not trigger it. Warn and ask, never block, since a deliberate second link is legitimate. Any form POST that creates a row must end in a redirect, or refresh becomes a duplicate-creation button. When a duplicate is reported, check older rows for the same shape before calling it a one-off.
Why: A person is not their email address. Keying identity to the address is the obvious implementation and it fails exactly on the case that motivates the feature: someone who signed up personally and is being sold to at work. A duplicate row is not harmless either, since every minted invite sits in the sent denominator the page reports its own click and signup rates against.
Failure mode: SUCCESS: OTP join-link dedupe. An exact-email duplicate check would have missed the case it was built for: Mendy Shishler signed up as one address and was invited at another, under a misspelled surname. Separately the mint page rendered its result from the POST, so a browser refresh re-submitted the form and minted a second token. That had already happened to another prospect a month earlier and nobody noticed it was a pattern.
Before acting on any "supply running out" alert, check what the existing supply CONVERTS at, and check that the conversion number is measurable at all. Specifically: (1) compare the alert's assumed burn rate against ACTUAL sends per day, since a cap is a ceiling not a rate and OTP had sent 0 for four days while the alert projected 2,500/day; (2) confirm the outcome pipeline is intact end to end before believing a zero, because "no positive replies" and "we destroyed the replies" look identical in a dashboard; (3) if outcomes cannot be attributed to a campaign, fix attribution FIRST, because every list bought before that is unfalsifiable. Also: do not lower a sending domain's daily cap to stretch supply. That number is an earned warm-up ramp, not a throttle, and lowering it discards reputation the domain has to re-earn; pace by staging fewer per day instead.
Why: A supply alert measures the tank, not the engine. Sourcing more leads into a funnel with a measured zero conversion multiplies data spend and burns the sending domain's reputation, which is the one asset that cannot be rebought. And the deeper trap is that the zero itself was untrustworthy: replies bounced entirely before 9/4 and were unattributable after, so the org was about to make a five-thousand-lead purchasing decision on a number that did not exist.
Failure mode: The outreach lead-inventory alert reported "OTP has 2.0 days of supply left, source more before the queue empties" and framed sourcing as the fix. Acting on it would have bought ~5,000 more addresses. Investigation showed supply was not the binding constraint: 19,884 sends in 30 days produced 0 human replies, reply capture had been entirely broken until the 9/4 MX fix, and replies.send_id was never populated so no reply could be attributed to a campaign, list or brand.
When turning a hand-built client report into a repeatable product, separate the two halves and build for the second one. The mechanical half (API pulls, arithmetic, tables) is easy and worth almost nothing on its own: WOA's own script divides spend by every iCart join and prints $12.92, the wrong number that had already shipped. The valuable half is four things that live in no API - the client-specific counting rule (plan name contains 'national promo' and not 'kiosk'), the target the client stated on a call, the phrases that must never appear in their documents, and the corrections a human made over eleven drafting passes. So: (1) give every client a stored, editable BRIEF holding exactly those, and refuse to generate without the counting method; (2) make the forbidden-phrase list a hard export gate checked on the finished document including what the human typed, not a hint in a prompt; (3) compute every figure in code into a fact sheet and check the generated prose back against it, allowing spoken rounding but rejecting any numeral no fact supports, so a failed section becomes a visible gap rather than an invisible invention; (4) keep a human as the last editor before anything leaves.
Why: A generator fed only ad-platform data reproduces the wrong report faster. The measurement definition, the agreed target and the commercial red lines are per-client knowledge that only a human holds, and if the product has nowhere to put them it will be confidently wrong rather than merely incomplete. A rule that exists only inside a prompt fails silently on the one draft nobody reread; the same rule as a gate cannot.
Failure mode: SUCCESS: Coach monthly report - the analysis scripts were only ~20% of what made the WOA/Pete report right, so productising it meant productising the human context, not the arithmetic
In a narrow sidebar, bound a calendar to one week and give it explicit next/previous controls. Never answer "what is on Thursday" with a scroll region: show all seven days at once, mark empty days, and disable the arrows at the edges of what the data source can actually answer for.
Why: A scroll area inside a page that already scrolls hides the thing it contains. The week is the unit people think in, so the panel should show exactly one and let them step. It also forces an honest data window: showing a whole week means the feed has to cover the days already past, not just the future.
Failure mode: Built the home.sneeze.it calendar rail as an open-ended agenda list inside a 300px sidebar with its own internal scrollbar. David: "you are showing too many days ahead only show the week and have the ability to move the date forward scrolling down does not help."
When evaluating any CTV/OTT platform (Vibe, MNTN, tvScientific, Madhive, Roku): (1) Screen the BUY before the PLATFORM. Two independent floors: delivery (can the geo spend at a sane frequency, usually yes, ~$400/mo for a 60K-person trade area) and measurability (a lift test needs ~250 conversions per arm, so one location at 50 leads/mo takes 5 months to resolve while 4 pooled locations at 200/mo resolve in 1.3). A single location can afford CTV and cannot measure it, so CTV is a corporate/multi-location buy, never a per-location upsell. (2) Never quote an in-platform CTV attribution number: a documented audit found it more than 4x the third-party incrementality result on the same campaign, IP-to-email household matching is ~16% accurate, and 7-30 day view-through windows credit conversions that were going to happen anyway. Design the matched-market holdout BEFORE the buy or the data to answer the question is never created. (3) Verify vendor corporate status at the primary source. Vibe.co is mid-acquisition by Walmart (announced 2026-06-23, into Walmart Connect, expected close end of Walmart FY2027) with NO financial terms disclosed - the "$1.4 billion" figure in aggregator write-ups is uncorroborated and must not be repeated.
Why: Two failure modes were live here. First, the obvious way to answer "is Vibe the right platform" is a feature comparison, which produces a confident recommendation for a buy that can never be proven - the agency then bills for a channel it cannot defend and loses the client at renewal. Screening measurability first inverts that and, for a multi-location book like ours, points at the larger and more honest line item. Second, a search-summary layer asserted a "$1.4B" purchase price that the Walmart and Businesswire releases do not contain; repeating it to a client would have been a fabricated financial fact. Check the primary source before repeating any acquisition figure. Our multi-location client base makes a matched-market holdout constructible, which almost no SMB advertiser can do, and that is the real CTV asset rather than any platform.
Failure mode: SUCCESS: ads-ctv skill built - CTV/OTT evaluation must screen on measurability before platform features, and two vendor facts nearly got repeated wrong
Bookings are not a function of outbound pickups. Callbacks, texts and inbound also book, and the CCM sheet carries its own CallBackReqBooked column. Name and check for a benign mechanism before calling any number an anomaly. Keep questions about one named person out of team channels, and route a low connect rate to ops as a phone-health issue using the sheet's Phone Number Health Check tab.
Why: Publicly doubting someone's work on their best booking day does real damage and was entirely avoidable. It is also narrating a proxy as truth: reporting a gap in my own understanding as a defect in someone else's work, while the disconfirming evidence sat in a column I had already read.
Failure mode: Called Amanda Zuze's CCM numbers (105 dials, 4 pickups, 9 booked on 2026-09-08) an arithmetic contradiction, and drafted that doubt into the team Slack channel where her manager reads it.
Before blaming a per-tenant switch, check whether a sibling system already succeeds at the same task, and if so establish whether it uses the SAME code path. iCart bills those exact clubs today, but through ABC's raw-card branch (todayBillingInfo/draftBillingInfo), never calling paypagecreateagreement. So it proves credentials/plans/club numbers are good and proves nothing about the Pay Page. The real finding: paypagecreateagreement has never returned a URL on ANY club including ABC's own sandbox 9003, so there is no positive control and "per-club switch off" is the weakest of three explanations (vs. app_id not entitled, or we are calling it wrong).
Why: A failure that reproduces everywhere, including the vendor's own sandbox, is not a per-tenant configuration problem. Diagnosing it as one sends a human to make a phone call that cannot fix it, and the vendor's answer ("it is enabled") then reads as a contradiction rather than as evidence for a different cause. Always find one place the call succeeds before blaming the places it fails.
Failure mode: Accepted "ABC has not enabled the Pay Page for these clubs" as the diagnosis for iJoin joins failing at Fitness Quest 6716/8321/40074, and recommended a phone call to ABC Support naming those three club numbers.
Before linking tickets by symptom ("end meeting does not work"), pull the underlying record's discriminating field (meetings.meeting_type) from the DB; a record-only meeting and a leadership meeting run different code paths. After writing a guard-style fix, remove each guard, confirm a test fails, restore it; the first test pass was vacuous for one guard and hid three more unguarded sites in the same file.
Why: Title-matching wrote a false "same root cause" claim into a PR that had to be corrected, and it would have written a false resolution onto a closed customer ticket. The mutation pass turned a fix that covered one call site into one that covers all five.
Failure mode: SUCCESS: Claude support-board audit found two look-alike tickets were different bugs, and mutation-testing found three guards the first fix missed
When David asks for a checklist, deliver only checkable lines: one box per item, each a short noun phrase or verifiable statement, grouped under 3 to 4 headings, no NEVER/rules section, no explanations, no examples in parentheses, no closing sermon. Put any rationale in the cover note to David, not in the artifact.
Why: A checklist is a tool the recipient ticks; a do/don't list is a reprimand the recipient reads once. David wants the vendor to use it every week, so it has to be scannable and neutral.
Failure mode: Pepper drafted a vendor "checklist" for Sonya that read as a do/do-not lecture: NEVER section, parenthetical examples, rationale sentences, and instruction prose mixed into the items. David: "this sounds like a do and do not I asked for a checklist."
A hostile reply suppresses the PERSON, never the domain. Block a domain only when the sender provably speaks for the company (corporate role, legal, the franchisor's marketing lead) or when the domain itself has told us no at scale. On a franchise brand where locations share the corporate domain, every address is a separate business; treat the domain as a list, not a company. Unsubscribes on that domain are individual too.
Why: Franchise brands are Sneeze It's core ICP and their franchisees are the buyers. One franchisee's temper is not the brand's decision, and a domain block silently removes the next 81 prospects from every future campaign with no record on any of their rows.
Failure mode: Dirk blocked the whole anytimefitness.com domain (82 contacts, 18 queued) because one franchisee, mike.steele@, replied "stop spamming us". Anytime Fitness franchisees use the corporate domain (jacksonville@, brightonmi@, and named people), so a domain block killed 81 independent owners for one person's reply. David: "Mike Steele is a franchisee not a franchisor, so don't kill the brand just for one asshole."
Any plan step of the form "keep X for N days, then delete/rotate/revoke" gets its own dated todo with an owner at the moment the plan is written, not at the moment the window opens. Before deciding delete-vs-label for a data copy, measure four things: newest row (the real freeze point), table-set diff and row counts against the live system (superset proof), pg_stat_activity connections (nobody reads it), and where the backup job actually points (check the archived table count in the job log). Delete when all four say so; do not export first, an export is another unmanaged copy.
Why: Unowned copies of customer data are a retention finding an auditor will ask about, and a copy that looks like production will eventually be mistaken for it (it was, for four days). The deferred-delete pattern recurs across migrations, restore tests and drills; the 8/8 restore test cleaned up same-day only because someone did it by hand.
Failure mode: SUCCESS: Claude found why a stale copy of production customer data sat unowned for 3 weeks: the 8/20 migration plan said "keep the old DB 7 days, then delete" with no owner, date or todo, so the delete step silently never fired and the incident runbook later pointed at the copy as production.
When David asks for delight (easter eggs, play, fun), the bar is a SHOW, not copy: synthesized sound (Web Audio, no files), things that move across the viewport, pop-culture scenes with a payoff. Text-only eggs are the floor, not the deliverable. Build a reusable effects layer (sound, flyers, overlays) and then write scenes on top of it.
Why: David's reference points are Tesla (sound, fart mode, Starman), Hitchhiker's, BTTF, Monty Python: multi-sensory, visual, surprising. Toasts read as safe and generic. Delight work judged by the same "is it slop" standard as design work.
Failure mode: Asked for "a lot more easter eggs" on home.sneeze.it, Claude shipped a dozen toast messages and one confetti variant. David: "you can do better" and named the bar: Tesla-style sound, things flying across the screen, Hitchhiker's Guide, Back to the Future, Monty Python.
Emoji are clip art; they cap the ceiling at "cute". For a show: draw the sprites (inline SVG with real detail, spinning wheels, flapping wings), run a canvas particle system (gravity, drag, sparks, fire, dust, fireworks), and make effects TOUCH the page (tiles get pushed by the car, squashed under the foot). Payoff beats count: fewer scenes done properly over many done cheaply.
Why: The first upgrade added scenes but kept clip-art rendering, so it still read as a toy. The rendering quality is what separates delight from a gimmick, the same way the drawing IS the argument on landing pages.
Failure mode: Second pass on home.sneeze.it easter eggs used emoji as the sprites (car, whale, foot, parrot) flying across on CSS keyframes. David: "cute but not fun, you can have greatness here, make the graphics better and funner."
When a RATE moves, check whether the DENOMINATOR moved before attributing anything to the numerator's owner. Here dials were flat (267 vs 291) and agent dials reconciled to the project total, so the rate fell because new leads nearly doubled (79 vs 40) and 57 of them were one client, HiTone, taking 51 dials across 8 locations (0.89 dials per lead). The finding is an ops capacity item, not a coaching item, and it goes to the weekly performance call rather than a caller DM.
Why: The previous day's recap draft was killed for publicly questioning a caller's numbers on a premise that was wrong (OOS L332). This is the same failure class one level up: a ratio is two numbers, and blaming the owner of one of them without looking at the other is how a data event becomes an unfair performance conversation. It also found the real problem, which is that a lead surge arrived and was not worked.
Failure mode: SUCCESS: Dan caught a rate drop that would have read as a caller-performance failure. The call center appointment rate fell 37.5% to 16.5% overnight. The reflex is to coach the callers or flag a collapse.
When a design depends on a third-party URL parameter or deep-link behaviour, open a browser and inspect the DOM rather than trusting a search-result summary. Live test 2026-09-11 found: claude.ai/new?q= and ?prompt= are both DEAD (composer empty, text absent from DOM; removed ~Oct 2025 per claude-code #8827/#19023), chatgpt.com/?q= WORKS and server-rewrites to ?prompt= which prefills but does NOT auto-submit, and perplexity.ai/search?q= WORKS and auto-submits while signed out. A web-search summary had asserted claude.ai works reliably, and it was stale. An adversarial reviewer subagent flagged the claim; the browser settled it.
Why: The whole QRPrompter design rested on claude.ai being a valid destination. Had it shipped on the search summary, the primary button on the product surface would have silently done nothing, and the failure would have looked like a phone problem rather than a dead API. Two cheap habits caught it: dispatch an adversarial reviewer that cannot see the conversation, and verify externally-owned behaviour empirically before it becomes load-bearing.
Failure mode: SUCCESS: Conatus verified LLM URL prompt parameters in a browser instead of trusting web-search summaries
(1) Never use Google customer-level metrics.conversions as leads. Pull campaign x day x segments.conversion_action_category (metrics.conversions is PROHIBITED on the conversion_action resource; select the segment from campaign or customer instead) and count only SUBMIT_LEAD_FORM, PHONE_CALL_LEAD, REQUEST_QUOTE, SIGNUP, CONTACT, BOOK_APPOINTMENT and PURCHASE. Exclude GET_DIRECTIONS, STORE_VISIT, ADD_TO_CART, BEGIN_CHECKOUT, PAGE_VIEW, OUTBOUND_CLICK, YOUTUBE views and GA4 funnel-step DEFAULT actions. On 9/14 this cut Champy's 2,003 conv/30d to 25 real calls (1,534 were directions lookups), J&K 1,448 to 96 (1,336 cart page views), and Syufy 3,535 to 84 SI iCart joins. (2) Pull Meta through /me/adaccounts with field-expanded insights.time_range(...){spend,clicks,actions} plus account_status/spend_cap/amount_spent, with one call per window per page, not per-account calls (the per-account loop hung past 10 minutes; the field expansion covered 229 accounts in under a minute). (3) Check account_status on every card: status 3 (unsettled payment) was silently stopping Fitness Quest, Fitness Quest for Women and GLO30 Short Pump, which a 7-day rollup shows only as "leads down". Put those cards on HOLD so no coach sends a momentum email about dark ads. Generator saved at ~/.claude/gen-coach-report-v3.py, pullers in ~/.claude/coach-report-pullers/.
Why: Last week's report told a client Google delivered 2,255 conversions at $1.33 each, and most of those were direction lookups. A client-facing win built on a vendor's catch-all metric is a false statement with a precise number attached. Payment holds look identical to soft demand in weekly numbers, and emailing a client about momentum while their ads are dark costs trust twice.
Failure mode: SUCCESS: Dash /coach-report 2026-09-14 found that Google Ads metrics.conversions was inflating client wins, and that Meta account_status 3 hides dark accounts inside a normal-looking weekly drop
Before asking an interview question built on a reference-file claim, check whether David was actually present for the event. Phrase it as "was this something you saw, or something you were told?" rather than assuming a scene. When David corrects a seed, fix the source ref file too, not just the story bank, because outreach copy reads from the ref.
Why: Story banks exist to stop invented stories reaching copy. A question that presumes a scene invites David to supply one. And a dramatised claim in proof-story.md ("audited and elevated") had been feeding outreach one-liners as if it were witnessed.
Failure mode: Neil built a story-bank interview question on a seed from proof-story.md ("survived PE acquisition, audited and elevated, not cut") and asked David what the audit felt like and what kept Sneeze It "when others got cut". David corrected: the audit was never made public to him, so there was no audit scene, and nobody said others were cut. The seed had turned a claim nobody witnessed into sales copy.
When a client complains in writing about bookings or follow-up, put their exact words in the callers' recap and name what it has already cost. Tie it to that day's specific uncalled leads. Point it at the process, never at one caller.
Why: Callers only see dials and rates. Softened language hides the stakes; the client's own words plus real lost revenue is what changes speed-to-lead behavior.
Failure mode: Arin's recap turned a client's written complaint (Scissors & Scotch: 18 appointments in 30 days, hesitant to add locations) into a soft line and left out that the same problem had already cost an account and a future WOA location.
A rollout-pace note on a Labs beta governs who switches it on, not whether customers hear it exists. The week's biggest shipped work goes in the weekly, usually as the pick, written as "turn it on for yourself first" to match the lab's own whyNow. Only hold a feature out of the Swamp when it is not reachable or not working in prod (the #35 tldraw-license test), never because it is a beta. Before writing a Labs pick, check every env switch it relies on and promise only what is live (e.g. tile watches ran but OLLIE_WATCH_EMAILS_LIVE was unset, so no email was promised). Also check that the Labs card the button lands on still describes the feature.
Why: The weekly exists to show customers what shipped. Holding the largest feature back because of how fast it should roll out would have sent "8 new things" with the headline missing, the same undersell as the #29-#32 undercounts.
Failure mode: Swamp issue #37: I left the whole v_next "The new OTP" Labs wave (new Ollie Home, answer tiles, Flows, company page, about 20 PRs and the biggest work of the week) out of the weekly email. My reason was David's 9/12 note that he wanted one or two people on the beta before a whole org. David at the gate: "what about the new Ollie in Labs? that should be a topic?" It became the pick.
When rebutting a client's revenue claim, lead with the full defensible match set, not the most conservative subset, and disclose the match method in the notes instead. Give the opponent's assumed conversion rate its own section: a tour close rate applied to raw web form fills is usually the second largest error after the zero-capture assumption.
Why: Over-conservatism in a dispute document reads as a weak position and concedes ground for free. The right place for match-quality caveats is the method note, where it signals rigor, not the headline, where it discounts the finding. Close rate misuse compounds every downstream figure, so it deserves to be named explicitly rather than buried in a table row.
Failure mode: In the Marietta lead misrouting analysis, the headline join count was reported as 258 (email/phone hard matches only), holding the 7 name-only matches back as a footnote. David corrected it to 265, the full match set, and pushed to lead harder on the other side's inflated close rate.
When David points at a competitor's proven UI and says copy it, spec the copy element for element first (invite link, paste box, same-domain chips, directory checklist, one CTA, auto-add toggle, other-workspaces connector bar). Put OTP-specific additions in a separate "after parity" section, never blended into v1.
Why: Fireflies has already paid for the UX learning. Blending our ideas into the first cut reintroduces the complexity David is trying to remove and delays a known-good pattern.
Failure mode: Proposed an OTP invite redesign that borrowed Fireflies' directory pick-list but then added our own variations (chart-seat matching chips, an Advanced disclosure, dropping the other-workspaces bar). David said the plan should copy exactly what Fireflies did, including the "Invite teammates from other workspaces" bar with Microsoft Teams / HubSpot / Slack Connect and Import.
When a creditor demand lands, ask David for the company's own theory of the counterparty's fault before framing leverage. Then build the causation record: what the counterparty did wrong, when, what it cost, and which loans and guarantees followed, each line sourced. Present the fee arithmetic as a check on their number, not as the case.
Why: A forensic reconstruction that only tests the creditor's number concedes the framing. David lived the sequence and knows the causal chain; the seat's job is to document it with dates and documents so counsel can use it, and to say plainly where the record is thin.
Failure mode: Alan framed the Bottom Line Concepts fee dispute as a contract-text and bankruptcy-discharge question and put the leverage in the agreement wording and ERC recapture risk. David's position is different: Bottom Line's incompetence (misfilings, wrong information, delay, the company doing the work itself) is what pushed the company into the loans and the Chapter 11, and one of those loans was personally guaranteed by Alan Wargaski as a non-owner. That causation theory is the leverage, not the fee arithmetic.
When a document is gated by a locked-figures list, anything a template prints must be pushed into that list, including free-text strings that merely contain a numeral (finding titles, denominators-in-words, ad names like "$1 Enrollment"). Add a test that composes a full report with no LLM available and asserts unverifiedNumbers(body, facts) is empty; run it before every new section ships.
Why: A gate that blocks the model but not the template produces a report that cannot be sent for reasons the coach cannot fix, and the first instinct is to switch the gate off. Making every printed figure a fact keeps the gate strict and usable at the same time.
Failure mode: SUCCESS: Coach client report rework (2026-09-16). An invariant test that a code-built report body must pass its own number guard caught two classes of figure the templates printed without a fact behind them: finding titles/evidence (rounded differently from the fact sheet) and a budget-basis label string ("$8,000/mo") that only contained a number.
When a page's buttons carry a class that sets display (inline-flex, flex, grid), the HTML hidden attribute stops working on them; add a scoped `[hidden] { display: none !important }` rule and drive the interaction (type in the search, click the chip) in a real browser before shipping. Render EJS standalone against the built CSS and serve it over http.server when no dev server is available; Playwright blocks file:// URLs.
Why: A screenshot of the first paint looked perfect; only typing a search showed "Show 24 more of 181" never hiding. Reasoning about CSS would not have found it. The design protocol's step 5 (drive the interaction, not just the render) paid for itself on the first page it was applied to after being written.
Failure mode: SUCCESS: Claude shipped the meeting templates gallery (PR #789) and caught a silent UI bug by driving the page, not by reading the CSS
A redesign of a page that feeds meetings is not done until the meetings flow itself changed: templates must be reachable from where meetings are created and scheduled (the Meetings page and its Start/Schedule modals), not only from a gallery a link away. And "looks right at first paint" is not the formatting bar: check every state (wrapped footers, long asides, narrow columns, mobile) against the reference page side by side before calling it done.
Why: David, 2026-09-16: "it could be better, seems like you got a bit lazy in formatting again and we really need to think about integration for this into the meetings." Fellow's templates page works because a template is one click from becoming a meeting on the calendar; a gallery that only runs-now is a brochure.
Failure mode: Claude shipped the meeting templates gallery (PR #789) with formatting David read as lazy (footer actions wrapping onto two lines, a long aside jammed next to the Library heading, cramped builder columns) and treated "run a template" as the integration, leaving the Meetings page and its Start/Schedule flow untouched.
Before presenting a Dash blind-spot, billing trigger, or any state-file alert as a current action item, confirm it hasn't already been resolved. Stale state files (Dash May 25 was ~5 weeks old) carry point-in-time alerts that may be closed by now. Trust confirmed/observed status over stale notes; flag the data's age and treat unverified alerts as 'verify' not 'urgent.'
Why: Re-surfacing already-resolved alerts as urgent erodes trust in the L10 briefing and spends David's attention during a low-push recovery window. Honesty about data staleness matters more than appearing comprehensive.
Failure mode: Dan surfaced the HiTone billing trigger ($43-49K/mo possibly un-invoiced) from Dash's stale May 25 state file as a live concern during the Jun 29 L10. David confirmed HiTone billing is correct and already handled, and asked to close it out.
KPIs/scorecards must live as tiles in OTP (the source of truth), not in markdown files or meeting briefs. Every active agent/human seat — including Dan's strategic co-founder seat — must own at least one OTP KPI tile. When proposing measurables, verify against list_my_kpis and create the missing tiles via update_kpi (auto-creates), rather than just tabling them in a doc. A seat with no number is sitting on the sidelines.
Why: EOS requires every seat to have a measurable. Discussing KPIs in a brief while OTP shows none of them makes the scorecard fiction and undercuts OTP as the coordination source of truth. Dan as co-founder must be measurable like everyone else.
Failure mode: Dan presented a Sneeze It agent-team scorecard as a markdown table in the L10 brief and treated it as 'the scorecard,' when the source of truth is OTP. David caught that Dan (and Arin/Pulse/Dirk) have NO KPI tiles in OTP at all — Dan's own seat had zero measurables. A scorecard that only lives in a file or meeting brief does not exist.
An OTP KPI with teamId=NULL renders only on /dashboard/kpis, never on any L10 scorecard (meeting scorecards filter strictly by meeting.team_id). To make a KPI show on a specific L10, PATCH /api/v1/kpis/:id with the meeting's teamId. The 'Dan L10' meetings run on the 'ai-army' team (065d1d4b-c7da-4e80-b3ed-d6b101471d2c). Tally's auto-create now includes teamId from a 'team_id' field in the registry entry, so new agent-army KPIs land on the Dan L10 automatically instead of orphaned. Find team IDs via GET /api/v1/teams; meeting->team via GET /api/v1/meetings.
Why: A KPI nobody can see on their meeting scorecard is functionally not on the scorecard. The owner/title is necessary but not sufficient — team scoping is what makes it report. This is a recurring gotcha for any agent creating KPIs via the API.
Failure mode: SUCCESS: Tally — agent KPIs were invisible on the L10 because auto-create left teamId NULL. David flagged that the new KPIs weren't reporting on the Dan L10 or /dashboard/kpis as expected.
Root cause was the SearchAtlas OTTO pixel having an EMPTY src="" in the layout head (v7.ejs, onboarding.ejs, main.ejs). The OTTO tag must carry its base64 data-URI loader in src that appends dynamic_optimization.js with data-uuid; with src="" the runtime never loads, so OTTO injects/verifies nothing. When an OTTO/SearchAtlas audit reports 0/N across ALL on-page categories, suspect the pixel loader, not the actual tags — verify the sa-dynamic-optimization script's src is populated, not the page's own meta.
Why: A 0/16 across every category despite visibly correct meta is the signature of a non-loading optimization runtime, not missing tags. Checking the pixel first avoids a pointless rewrite of titles/descriptions that were never the problem.
Failure mode: SUCCESS: Beacon/SEO — orgtp.com OTTO on-page audit showed 0/16 (titles, meta descriptions, headings, meta keywords all failing) even though pages had perfectly good title tags and meta descriptions server-side.
A brand battle cry needs a genuinely designed moment (confident display type, intentional line breaks, brand device, real whitespace), not a centered text block plopped in. And the VISIBLE battle cry copy is the short clause only: 'Unlocking the potential in every person through the partnership of people and AI' — drop 'so together we leave the world better than we found it' from the hero display (keep the full sentence only for formal/footer contexts).
Why: A mission line is a brand centerpiece. Long copy dilutes the punch, and an undesigned drop-in reads as filler. The payoff phrase 'partnership of people and AI' must land as the climax with design weight behind it.
Failure mode: Adding the OTP mission as a 'battle cry' on the landing page, I dropped the full sentence into a plain centered text band wedged between hero and Step 1. David called it 'a weak attempt to just throw it on the page' and said the full line is too long for the visible battle cry.
Manifesto/mission pages must be written as movement recruitment, not product marketing: second-person address (the reader is the protagonist), "We believe" creed statements people can recite, a named enemy, stakes, and invitation CTAs ("Join the movement") instead of transactional ones ("Start free"). Product features appear only once, framed as how the movement fights, not what the product includes.
Why: People join movements because they believe what the movement believes (Sinek: start with why). Copy that sells the what on a page whose job is to recruit believers reads as generic SaaS and inspires no one, no matter how good the design is.
Failure mode: Redesigned the orgtp.com manifesto homepage with strong visual design but kept product-brochure copy (feature lists, "free meeting software", "Start free" CTAs). David: "the writing does not inspire an army of followers... this just looks the same as every other company... blah."
Judge conversion on the full path the visitor actually walks (page, door, day-one experience), not on the surface being edited. If the honest answer to "would you sign up" is "yes IF another surface delivers," the answer is no, and the work moves to that surface. Never write a promise on a button that the destination page cannot cash.
Why: Trust destroyed at the moment of verification is unrecoverable; a skeptical buyer who clicks "watch us run" and lands on a data page is gone forever. Copy that outruns proof is hype by definition, and the exact audience OTP needs (operators) is the audience that punishes it hardest.
Failure mode: After rewriting the OTP homepage, I declared the copy converts because skeptics would "click through to the live OOS page and sign up IF it delivers." David called it: I kicked the can to a page I know does not deliver, and called it a win. The button promises "Watch our company run, live" but the OOS page it links to is a list of published rules, not a running company.
OTP's enemy statement is "you bought the operating system and the needle didn't move." The pitch is not better meetings; it is: the system was fine, what was missing was the workforce that runs it between the meetings. Frame all homepage/sales copy against needle-not-moving, not against meetings.
Why: This is the buyer's actual lived disappointment (paid for an operating system, company looks the same two years later) and it positions OTP against incumbents on outcomes instead of features.
Failure mode: The letter's hero framed the enemy as "the meeting" / busywork. David corrected the thesis: the real problem is that companies bought operating systems and software (Ninety, Bloom Growth, etc.) that did not move the needle. Years later the company had not grown and was not better, and they needed to change how they did things.
For any UI change, design from the user's mental model, not the data model: "my list shows my work; work I assigned to others shows under Waiting on Others." When a meeting todo is assigned to someone else, stamp the creator as delegator so it routes to the delegation view. Before shipping UI changes, run the UX lens (impeccable / web-design-guidelines skills + src/DESIGN.md), not just a minimal code patch.
Why: A technically-correct patch that ignores the user's mental model just moves the confusion. OTP's own product language already has the right home for these items (Waiting on Others); fixes should land in the model the user already understands.
Failure mode: Fixed the dashboard todo confusion (teammates' meeting todos looked like the viewer's own) by adding an owner label to the rows. David corrected: that's not thinking like a user. Labeled-or-not, other people's todos don't belong in "my to-dos" at all.
When David asks for a jaw-drop brand page, build an EXPERIENCE, not an article: full-viewport cinematic hero, scroll choreography, one idea per screen at massive scale, motifs that live in the page as motion, ruthless copy cuts, no standard nav/footer chrome breaking the spell, no section-grammar scaffolding, no FAQ accordion bolted onto a manifesto.
Why: The gap between "well-executed page" and "omg I love this" is the whole assignment on brand surfaces. Safe editorial structure is invisible at best; for-the-brave positioning demands the page itself be brave.
Failure mode: Built the /ollie manifesto page as a competent editorial layout (repeated mono eyebrow labels on every section, index rows, alternating light/dark sections, FAQ accordion at the bottom) and David rejected it outright: "this really really sucks." The brief was "reader drops on the ground saying omg I fucking love this" and the output was a safe template that reads as AI scaffolding.
When David gives a design reference URL, open it in a browser and STUDY it visually (proportions, type sizes, spacing, alignment) before designing; match its register, not just its layout skeleton. Elegant means restrained: modest type scale, centered calm hierarchy, generous whitespace, thin rules. Never hand-draw SVG artwork to imitate produced brand art; crop/reuse the actual asset or use nothing.
Why: A reference URL is the brief. Reading its HTML structure without seeing it rendered led to importing the skeleton with the wrong soul, twice. Amateur freehand art next to professional motion work destroys credibility instantly.
Failure mode: Second rejection on the /ollie page. David asked for sakana.ai/fugu: elegant, Japanese sense of design (restraint, whitespace, calm, modest type, precision). I delivered giant 9vw headlines, one shouting line per viewport, and hand-drawn SVG chevron "birds" that rendered as crude fat marker scribbles. I treated "jaw-drop" as scale and boldness when the reference was quietness and precision, and I drew freehand SVG art instead of using the actual video's artwork.
For fleet-wide spec maintenance: (1) tarball backup of ~/.claude before any agent touches specs; (2) partition files into DISJOINT clusters, one agent each, with CLAUDE.md owned by exactly one; (3) give every auditor the same stale-fact canon and the rule "verify a launchd plist exists before believing any schedule claim"; (4) auditors apply surgical edits directly for factual fixes but RETURN structural proposals for David instead of applying them; (5) synthesizer closes cross-cluster contradictions the auditors flag at each other.
Why: Agent specs rot faster than anyone audits them: this pass found live specs for a retired agent (jeff.md ending in "Go."), four phantom schedules, Todoist writes in five files, terminated employees still routed DMs, and a Bassim score-inflation bug. Periodic fleet audits with disjoint ownership are cheap insurance against agents acting on dead infrastructure.
Failure mode: SUCCESS: Claude ran a five-cluster parallel level-up of the entire agent army (80 files, ~140 surgical edits) without a single file conflict or lost spec.
Pattern for UX dead-end hunts: (1) fan out parallel read-only explorers per surface (meetings, teams/members, KPIs/todos, onboarding/settings) asking for file:line + user-visible symptom + minimal fix; (2) fix the unsatisfiable states first: any required dropdown that can render zero options must explain where its options come from and link there (owners/attendees come from the org chart, meeting membership from teams); (3) empty states must branch on WHY they are empty (org has no teams vs user not on a team need different CTAs); (4) never report an async side effect as done: invite emails now await sendEmail (which returns null on failure, never throws) and return emailSent so the UI can tell the truth; (5) a guided setup checklist computed server-side from actual data (seats/team/KPI/meeting/members exist?) beats static onboarding because it survives skipped onboarding.
Why: These are the recurring shapes of broken UX in OTP: forms with prerequisites the user cannot see, empty states that misdiagnose their cause, and optimistic success messages over fire-and-forget side effects. Fixing the shape, not just the instance, is what makes the product feel intuitive.
Failure mode: SUCCESS: Claude ran a full UX dead-end audit and fix pass across OTP (4 PRs, #113-#116, all deployed)
The sweep pattern that worked: audit by rule-cluster in parallel (fakery, insight-to-agency, jargon/states, first-meeting goal-walk), then execute severity-first. Key catches to re-check every run: (1) seeded/synthetic data leaking into numbers a reader believes are real (the is_template flag existed but was never enforced; counts now use src/shared/synthetic-orgs.ts); (2) the conversion moment must be ON the default path (end-meeting now lands on Ollie followups, not the list); (3) funnels don't exist until instrumented (insight topic: surfaced/accepted/value_delivered); (4) credentials in seed script comments (one prod DATABASE_URL scrubbed; password rotation still owed). Worklist for run 2 in otp-platform/mission-standard/WORKLIST.md.
Why: The Mission Standard is a repeatable bar, not a one-off audit. Recording the found failure classes makes run 2 start from run 1's ceiling instead of re-discovering it.
Failure mode: SUCCESS: Claude ran Mission Standard sweep run 1 (PRs #121-#123, deployed): 4 parallel rule-audits over OTP, 5 CRITICAL + 13 GAP found, all CRITICAL and 9 GAP closed same-session
When an agent-pushed OTP to-do references a document, include a clickable https link (a Google Doc), not a local file/vault path — David reviews to-dos on mobile. To update an existing to-do's description, use PUT /api/v1/todos/:id (not PATCH). otp-todo.sh has no update verb, so PUT directly with the API key.
Why: A file path in a to-do is dead weight on mobile — the reviewer can see the reference but cannot open it, which reads as "the link is missing/broken." Every agent that pushes doc-linked to-dos (Radar, Pepper, Dan) hits this.
Failure mode: Dan pushed an OTP to-do referencing a document but put a local Obsidian vault path ("2nd Brain/Agent Army/Dan/...") in the description. David opens to-dos on his phone — a vault path is not tappable, so there was "no link to click." Also used PATCH to update the to-do; the OTP todos API update method is PUT /api/v1/todos/:id (PATCH hits the marketing site and returns HTML).
To put an agent-army/IDS issue on an OTP meeting board, POST /api/v1/tickets with the team's teamId (category 'other' for strategic issues, priority low/medium/high/critical, ownerEntityType+ownerExternalId). The MCP submit_ticket tool CANNOT do this — it has no teamId param (it is the generic 'report a bug to OTP' path), which is why nobody ever got issues onto the board. Team IDs: 'AI Army' = 065d1d4b-c7da-4e80-b3ed-d6b101471d2c (the David+Dan agent-army meeting); Leadership Team = c1e1a485-414e-48d5-ae44-e81bd110b554. Update/solve via PUT /api/v1/tickets/:id (idsStatus, priorityRank, resolution).
Why: Agents could push KPIs and todos to OTP but not issues, so every L10 IDS board rendered empty and David kept discovering the hole live. Issues=tickets + teamId scoping is the missing piece; without it a meeting-readiness check would keep mislabeling a working API as absent.
Failure mode: Dan claimed 'no issues API exists in OTP' because there is no src/routes/api/issues.ts. That was wrong. OTP stores IDS issues in the TICKETS table (schema.ts: 'issues live in the tickets table'), with full IDS support (idsStatus, priorityRank, teamId, owner fields). The agent-army IDS board was empty only because our issues lived in a local markdown file and were never pushed as tickets scoped to a team.
OTP has TWO distinct features both called "Ollie Insight": (A) the per-meeting followups wizard that turns a transcript into meetings.ai_summary (src/shared/meeting-followups.ts, transcript-only), and (B) the reusable "address engine" (ollie_insights table, src/services/ollie-insight.ts + shared/ollie-insight.ts + partials/ollie-insight-block.ejs) that gathers org data (KPIs/rocks/todos/meeting-summaries) per SCOPE. When David said "KPIs shouldn't be in the meeting analysis," the fix was in System B's meeting-scope evidence gathering, NOT System A. The block partial (ollie-insight-block.ejs) is fully scope-generic (builds the API URL from data-oib-scope/scopeId client-side), so adding a brand-new 'quarter' scope end-to-end took only: add to INSIGHT_SCOPES + RULES_BY_SCOPE (shared), the scopeQuerySchema enum + a resolveInsightScope branch (api), a gatherEvidence branch + max_tokens (service) -- then just include the existing partial with scope:'quarter'. No new render/generate/receipts UI. Pattern: when a feature is "one engine, many surfaces," new surfaces are a scope + evidence branch, never new UI.
Why: The name collision hides which code to touch; picking the wrong system wastes a whole edit pass. And recognizing the scope-generic block means big-feeling asks ("a quarterly synthesis button") are small, low-risk diffs. Both are recurring shapes in OTP's Ollie work.
Failure mode: SUCCESS: Claude separated OTP's two "Ollie Insight" systems and added a whole new scope by reuse
New endpoint POST /api/v1/meetings/:id/agent-record (PR #154): an agent submits the written meeting record; OTP runs the same redaction ruleset, stores it to meetings.transcript, logs an audit baseline + agent_record event, and the existing /ai/followups generate turns it into to-dos/issues/headlines/insight unchanged. Agent path: ~/.claude/otp-meeting.sh record <meetingId> --file=<record> --source=l10dan. Wired into /l10dan conclude step 6. Verify deploy by probing the endpoint returns JSON not the marketing SPA HTML before pushing records.
Why: Ollie only reads transcripts, so agent-run meetings had no way into it — the empty-insight hole David hit live. This closes it: any agent meeting can now produce Ollie follow-ups. Also a general UI rule captured: buttons reflect capability/state (action taken -> button disappears).
Failure mode: SUCCESS: Dan shipped the Ollie agent-record path so agent-facilitated meetings (David + AI L10, no audio transcript) can feed Ollie Insights.
Before treating a KPI/data source as blocked, re-read the LIVE source, not the note about it. The Havok "non-client %" KPI was marked blocked for ~14 weeks on "the timesheet has no client column" — a 105-day-old memory. The live sheet (1VPlH5ZqTowOe2nJDFwOjuvXrCXUnOzeOb3aO-M936Xo) had since grown per-person tabs with a full Client/Time/Date schema; one read unblocked it. Pattern for wiring a messy sheet into Tally: (1) get_spreadsheet_info to list tabs, (2) read a person tab to learn the real schema, (3) confirm recency by reading the tail (last date), (4) add a focused extract mode to tally.py rather than reshaping the sheet — here `client_attribution_since` (reads ALL valueRanges via a new _values_2d_all, parses h:mm via _parse_hhmm, added dotted DD.MM.YYYY to _parse_date, classifies internal by an `internal_contains` substring), (5) `tally.py --dry-run --kpi "<title>"` to prove the number off live data before pushing. Human-owner + a 1:1-team KPI (owner HUM_BOGDANTABAKA, team = David-Bogdan 1:1) pushes fine via find_or_create_kpi.
Why: Blocked-status notes rot silently while the underlying source improves; a KPI can sit "pending" for a quarter when it was buildable weeks ago. Re-reading the live source first is the cheap unblock. And the tally.py extract-mode pattern makes any timesheet/sheet a live KPI without asking a human to restructure their doc.
Failure mode: SUCCESS: Dan/Tally shipped the Havok client-attribution KPI live in one session after a 14-week "blocked" note turned out stale
Before editing any OTP view to fix an on-screen bug, grep unique visible strings from the screenshot (e.g. "WAITING ON OTHERS", "always only yours") across src/views to confirm WHICH template renders that exact surface. Multiple pages can render similar-looking todo lists (me-todos.ejs vs dashboard-daily.ejs). Verify the rendering route (reply.view target) too.
Why: Two round-trips and two merged PRs produced zero visible change because the edits were on the wrong template, which read as "nothing is fixed" and eroded trust. A 10-second grep on the screenshot text would have pointed to the right file immediately.
Failure mode: Fixing an OTP todos UI bug, I edited src/views/pages/me-todos.ejs twice and shipped two PRs, but the surface David actually uses is the dashboard-daily "Waiting on others" widget (src/views/pages/dashboard-daily.ejs). Nothing he saw changed.
Two reusable patterns: (1) Before building any OTP email/engagement feature, grep src/services for existing infrastructure -- re-engagement.ts, lifecycle-scheduler.ts, and user_engagement_log already carried cadence caps, suppression, logging, and a daily cron, so the todo-aware upgrade was ~350 lines instead of a new subsystem. (2) When resolving "which org does this Clerk user belong to", organizations.clerkOrgId only knows the org CREATOR; invited teammates must be resolved through org_members.clerkUserId + claimedEntityIds. This gap is why per-user personalization (open todos) missed non-creator members like Nate.
Why: One engagement channel with shared caps is what keeps daily utilization pressure from becoming annoying double-mailing, and the creator-vs-member resolution gap will bite any future per-user feature (digests, notifications, billing seats) that starts from organizations.clerkOrgId.
Failure mode: SUCCESS: Claude shipped the smart engagement email engine (PR #186) by upgrading the existing re-engagement service instead of building a parallel system
Frame help/success/onboarding call copy positively: state plainly that the call is there to help and guide them, and describe what will actually happen on it (we'll walk through your setup, get you unstuck, answer your questions). Never say "this is not a sales call" or "no pitch" — describe the help, don't disclaim the sell.
Why: Defensive "not a sales" language triggers the exact suspicion it tries to defuse and undercuts a genuine help offer. David flagged this immediately.
Failure mode: Wrote a "customer success call" Calendly description that leaned on "no pitch, no slides" / not-a-sales-call framing. Protesting that it isn't a sales call makes it sound like one.
Any script that answers "is anything missing / is everything covered?" must fail LOUD, never return an empty set as reassurance. Two rules: (1) assert the expected top-level key exists (`if 'data' not in resp: raise`) before computing a result; a zero/empty answer from a health check is a claim that must be proven, not a default. (2) Cross-check a zero result against one known-positive case before reporting it -- here, one direct call to a single account would have shown $57 of spend and exposed the lie instantly. Also: `~/.claude/meta-ads.sh accounts` exits 0 and prints nothing; do not build on it. Sweep with `/{business_id}/adaccounts?fields=name,account_status&limit=500` then batch `/{act_id}/insights` 50 at a time.
Why: "Nothing is wrong" is the single most dangerous output an audit can produce, because nobody investigates it. A silent empty result on a billing sweep means real revenue is never invoiced and nobody ever finds out. The failure mode is not a crash, it is confident silence.
Failure mode: SUCCESS: Claude caught a silent false-negative in a Meta billing sweep. Querying the Graph API adaccounts edge with nested field expansion (`fields=name,insights.date_preset(this_month){spend}`) returned an error payload with NO `data` key. The sweep script read it as an empty account list and confidently reported "0 accounts, $0.00 unbilled MTD spend" -- a clean bill of health that was entirely fabricated. Direct per-account queries then revealed 7 unbilled accounts spending $5,673 MTD.
When you own a strong long-form asset (a landing page that argues the WHY), the cold email must NOT re-argue it in miniature. Cut the product explanation entirely. The email's only job is to earn one click: proof you read their work, one line naming THEIR problem in THEIR language, the link, out. No feature, no price, no call ask. Match the page to the reader's specific pain rather than sending everyone to the same URL. And assign/log an A/B arm on every send, because an untested signature shipping at 100% for months is not a decision, it is a habit.
Why: A cold email is the worst possible venue for explaining a product: no trust, no attention, no context. A page is the best. Using the email to sell the READ instead of the PRODUCT plays each asset to its strength. The deeper failure was cheaper: months of sends with no arm logged and no reply column filled, which means the volume produced zero learning. Volume without measurement is just noise you paid for.
Failure mode: SUCCESS: Crafter cold-email rework -- the email body was trying to explain the product to a stranger in three sentences, while the website spent full pages arguing the why. The two assets contradicted each other and the email was losing. Outreach log had been dead since Apr 26 with effectively no replies.
When compiling guru influences into an Ollie persona, push voice fidelity hard: the dominant influence's cadence, vocabulary, and signature moves should LEAD the writing, not decorate it. Channel the style ("in the room, guiding in their voice") while keeping the hard never-impersonate line: never claim to BE the person, never claim endorsement. Also: differentiation surfaces (the Lab) must let users edit each guru's forked principles inline and must show WHERE a principle or voice move landed in the output (highlights on the insight, influence tags on todos/issues/headlines) so voices can actually be compared.
Why: The entire value of per-team Ollie voices is that the voice audibly changes what the team hears. If influences only shift content, not voice, the differentiation is inaudible and the Lab proves nothing. Attribution marks are what make the difference legible.
Failure mode: Ollie Lab guru influences rendered as principles-with-a-hint-of-style: the output read like Ollie citing a thinker, not like the thinker's voice guiding the room. David: "it should be as if they were speaking with their voice guiding us in their voice... It should be like they are in the room."
Voice fidelity lives in structure, never slogans. Ban catchphrase quoting explicitly in the persona compile ("never quote their slogans; that is imitation's cheapest form"). Give each library guru a hand-written VOICE DNA block: sentence rhythm and length, how they open, how they build an argument, what they notice first, how they land a point, emotional register, what they never do. For custom gurus, instruct the model to reconstruct the person's published voice from structure and register, not taglines. Method instruction: "before writing, ask how NAME would structure this and what they would notice first; write from inside that mind."
Why: Catchphrases signal imitation and break trust instantly; structural voice makes the reader feel the thinker in the room without a single borrowed phrase. This is the difference between a costume and a mind, and it is the entire premium of per-team Ollie voices.
Failure mode: Ollie's guru voice channeling produced surface mimicry: dropping the thinker's catchphrases ("Start with why") instead of writing from inside their rhetorical DNA. David: "not cheap tricks... it needs to be in the DNA of what Ollie is saying. Dig deep, make it real."
Essay-page hero art is a first-class deliverable, not a wireframe: match the family's craft bar (a real scene that tells the page's story, Ollie present, warm accent treatment, gradients/glow, staged ignition-style animation, reduced-motion complete state). Before drawing, read the sibling page's actual SVG to absorb its techniques, then design a scene, not a diagram.
Why: The hero art IS the argument at first glance on these pages; the candle page's flames are the thesis made visible. A schematic undercuts an inspiration page precisely where it must inspire.
Failure mode: The /the-voice-in-the-room hero art shipped as a lazy schematic (dot, four thin lines, grey boxes with circles for people) while its sibling pages (/one-candle, /mission-to-the-moon) carry hand-crafted narrative scenes with the mascot, gradients, glows, and staged animation. David: "you got lazy there for sure."
When running counterfactual/what-if simulations, make the events endogenous: model causal mechanisms (innovation rate, demographics, conflict outcomes) as functions of the changed variable and run Monte Carlo over branching timelines, rather than overlaying new participation rates on the fixed historical record. Show a distribution of divergent timelines, not one re-skinned version of real history.
Why: Fixed-event counterfactuals smuggle in the answer (the world converges because you forced it to). Branching simulation is what actually answers "would things be different" questions, and it's the difference between a re-labeled chart and genuine out-of-the-box analysis.
Failure mode: Built a counterfactual history simulation that held all real-world events fixed (same Industrial Revolution, same wars, same tech timeline) and only varied participation rates within them. David flagged that this assumes the conclusion: if the labor/care allocation changes, the events themselves change — maybe industrialization comes later (or earlier), wars resolve differently, the whole timeline branches.
In live L10/Delta Meeting facilitation, sequence is: surface ALL signals first (grouped, neutral, no recommendation attached), let David react and pick what to IDS, and only then frame decisions. Decisions come after shared context, not before. Also: open every meeting with Ollie's read of the PRIOR meeting's record (pulled from the OTP followups/insight), as a standing first section — David explicitly values it.
Why: A decision framed before the signal review railroads the meeting toward Dan's framing and skips the part where David's pattern-recognition works on the raw material. The facilitator's job is to lay out the board, not to compress it into a pre-picked fork. And the Ollie prior-meeting read is the continuity loop that makes each meeting compound on the last.
Failure mode: Dan facilitated the live L10 by pushing straight to the rock-set timing decision immediately after the scorecard, without first walking through all the signals (client wins, CC trends, attribution reading, unverified tiles, churn signals) so David could see the whole board. David: "we need to review all the signals before moving so quickly, you are trying to skip ahead too fast."
L10 prep MUST include a live scan of OTP itself: the team's rocks/priorities board, the issues (tickets) board, todos, and KPIs, pulled fresh at prep time. OTP is the source of truth; local files are mirrors that go stale the moment David works in the product directly (which is the whole point of OTP). Additionally, per David's 7/13 ruling: local shared-state files are now formally HISTORICAL unless a live consumer reads them; staleness flags on superseded files are noise, but staleness in the OTP scan is a real miss.
Why: David increasingly works inside OTP directly (weekend rock-setting), so any prep that skips the live product misses his most recent decisions and re-surfaces solved or stale items. The mirror-drift failure has now happened twice (June 15, July 13); the fix is structural, scan the product, not the mirror.
Failure mode: Dan's L10 prep read local mirror files and this morning's Tally/KPI pipeline but never scanned the live OTP boards (rocks/priorities and their attached issues) before the meeting. Result: Q3 rocks David added over the weekend were missing from the prep, and stale issues sitting on the rocks board went unnoticed. David caught it live, again (first time was June 15).
Every KPI on any board must pass the needle test before it earns a tile: one sentence stating the causal chain from this number to the company goal (margin, retention, revenue, rock completion). Dan owns running this test — on every existing tile quarterly and on every proposed tile before creation. Activity metrics (emails drafted, projects counted, pushes made) are health checks at best; they live in the readiness script, not on the scorecard. Wiring a dead tile is worthless if the tile measures the wrong thing.
Why: The strategic co-founder seat exists to hold the big picture David cannot hold while operating. A perfectly-wired scorecard of needle-irrelevant numbers is worse than an empty one, because it manufactures the feeling of accountability without the substance. This is the second-order version of "every seat owns a number": every number must own a reason.
Failure mode: Dan treated the scorecard as a plumbing problem (are tiles wired, do values push) instead of a strategy problem (does each KPI move the needle toward the goal). David: "we/I make all of these changes thinking you are looking at the big picture, that does not seem to be the case. Your job is to make sure that we reach our goal, and the KPIs should support that. How many emails Pepper reads does not move the needle. Each KPI needs to answer how it moves the needle, and how."
L10 prep must WALK THE ACTUAL MEETING before the meeting: open/fetch the exact meeting David will see (via API), verify every section renders with real data (scorecard snapshot has values, rocks board current across ALL teams incl. corporate, issues/todos loaded), run the needle-test on every tile, and FIX or stage fixes for everything found - all before 8am. The brief reports what was already repaired, not what will be discovered. Prep = simulate the meeting end-to-end; the meeting itself is only for decisions the human must make.
Why: A meeting that debugs itself live burns the scarcest resource (David's attention) on work an agent could have done at 7am. "The work happens between the meetings" is the entire operating philosophy of the meeting cadence - prep that only compiles data without verifying the meeting surfaces is half a prep. This is the root cause behind L059/L060/L061; fixing it structurally prevents all three recurring.
Failure mode: The 7/13 L10 scored 4/10. David's reason: prep ran in the morning but the meeting still spent most of its time discovering and fixing things live (weekend rocks missed, empty scorecard render, dead tiles, needle-less KPIs, corporate rocks invisible) - "we are fixing the meeting within the meeting with an absence of information. The work happens BETWEEN the meetings and this is not the case here."
When auditing whether an invariant holds across a codebase, verify it per STATEMENT, not per file. A per-file grep count is an aggregate, and aggregates hide the exact case you are hunting: the mixed file. Then encode the audit as a test that scans every call site, and mutation-test that scanner by reintroducing a known bug to confirm it actually fails. A scanner that only agrees with current code proves nothing. Related pattern seen the same day: when a data model gains a concept (rock levels, agent-owned KPIs), the model and the dashboard get wired up and the OTHER surfaces silently do not. Ask which surfaces read this table, not just which one is broken.
Why: The three holes the per-file count missed included the worst one: the blueprint serializer baked a private rock into a shareable template, which would carry it into another account. The failure mode of an aggregate check is a false clean bill of health, which is more dangerous than no check at all because it stops the search.
Failure mode: SUCCESS: Dan audited a privacy invariant (shadow rocks are owner-only) and found 7 holes, but the FIRST pass counted guards per FILE and missed 3 of them, because a file can contain one guarded query and one unguarded query and still look guarded in aggregate.
Invert it. Give the user ONE copy-paste block (MCP connection + a self-registration prompt) that they drop into Claude Code, Claude Desktop, ChatGPT, Cursor, or any MCP client. The AGENT then connects to OTP and registers ITSELF: it reads its own system prompt, calls a register/enroll MCP tool with its name, role, what it owns, what it does not own, and its KPIs, and OTP creates the seat and KPIs automatically. The user's only job is copy, paste, done. No forms, no parsing, no filling anything in.
Why: The agent already knows what it is -- making a human retype it is redundant work and a competence gate. Any flow that requires the user to know how to describe or configure their agent is not idiot-proof and will lose the non-technical user. The agent is the most reliable source of truth about itself, and it is already sitting on the other end of the MCP connection, so let it do the work. Design rule: when an AI is on the other side of the pipe, push the setup work to the AI, not the human.
Failure mode: Built OTP's "Connect an agent" flow as a human-driven form: the user pastes their CLAUDE.md into a textarea, OTP parses it, and the user hand-fills name / role / owns / does-not-own / KPIs before a seat is created. It made onboarding an existing agent the USER's clerical job and assumed the user knows how to describe their own agent.
Before any L10, dry-run tally.py and for every regex_in_file KPI check the source file's mtime against the KPI's time grain; if older than one grain, re-pull the number from the live system (Accelo, Search Atlas, Sheets) before pushing. New registry entries must use kind/regex/group (never type/pattern) and carry pending:true until their emit line exists. The unknown-kind branch now honors pending.
Why: A KPI pushed from a stale mirror is worse than a missing one: Crystal's tile would have said 32 when reality is 44, and the failure alert noise from mis-schema'd pending entries erodes trust in the one agent whose whole job is keeping the scorecard honest. Live-source-first is the same lesson as L042/L060 applied to Tally's own pipeline.
Failure mode: SUCCESS: Tally pre-L10 KPI sweep found and fixed three silent scorecard rot points: (1) registry entries added at the 7/13 L10 used type/pattern keys but the runner requires kind/regex, so their pending:true flag was never honored and they reported as failures; (2) the runner's unknown-source-kind branch ignored pending entirely; (3) two regex_in_file KPIs (Crystal 32, Beacon 0) were feeding from stale files (Jun 8 and Jun 22) while the live sources (Accelo: 44 projects; Search Atlas: 8 keywords tracked, not 43) had moved.
Any browser feature that accumulates unrecoverable state in page memory (MediaRecorder audio, unsent drafts) must (1) block/defer every programmatic self-reload while active, (2) be flushed by navigation-triggering handlers (End meeting awaits OTPAudioRecord.finish() before navigating), (3) guard beforeunload, (4) never drop data on a failed upload — keep the blob and offer retry. Long-term: stream chunks to the server (timeslice) so the page is never the only copy.
Why: A full Delta Meeting's recording/transcript was permanently lost — the audio never left the browser. Guard rails shipped in audio-record.ejs + l8-leadership.ejs; chunked streaming upload is the queued follow-up (touches billing-metered meeting-audio.ts, needs billing lock).
Failure mode: SUCCESS: Conatus — root-caused OTP meeting recording loss (2026-07-16): the browser recorder holds all audio in tab memory until Stop, while the live meeting page self-reloads on routine actions (reloadKeep, SSE scheduleReload, section-refresh fallbacks) and End Meeting navigates away — any of these silently killed a live MediaRecorder with zero warning.
Durable pattern for any browser feature holding unrecoverable state: (1) layer the fix — guard rails first (cheap, same-day, stops the active bleeding), durable streaming/persistence second; (2) keep the old one-shot path untouched as an automatic fallback so degradation can never be worse than before; (3) money-path parity — finalize calls the exact same precheck/ingest/charge sequence as the one-shot path, charge only after successful transcribe+ingest; (4) hold the billing lock across the whole build, release only after merge; (5) verify with a harness that runs the REAL shipped script (44 assertions) so old behavior is provably byte-identical with the new feature off; (6) cap every attacker-spinnable counter (segment count was a finalize-loop DoS lever).
Why: Completes L067's open loop: the "never lose a recording" guarantee is now structural, not procedural. The layered-fix + fallback-preserving pattern is reusable for every OTP feature that buffers user work in the browser (draft notes, offline edits), and the billing-parity discipline is how streaming touched the money path with zero semantic change. Both PRs merged and confirmed live on orgtp.com 2026-07-16 evening.
Failure mode: SUCCESS: Conatus — full resolution of the 2026-07-16 meeting-recording loss, shipped to prod same day in two layers: PR #205 guard rails (meeting page defers all self-reloads while recording; End Meeting flushes the recorder before navigating; beforeunload guard; pinned REC pill; failed uploads keep the blob with retry) and PR #206 streaming upload (recorder streams ~10s chunks to a server recording session; crash/reload loses ≤10s; resume banner stitches segments into one transcript; phone QR flow covered).
When invoking Steve Jobs as a design standard, treat him as holding BOTH axes to one bar: visual craft (typography, proportion, detail) and end-to-end experience (defaults, subtraction of steps, invisible mechanism). Never frame him as the "UX half" opposite a visual system; frame the visual system as one half of the single Jobs-level standard.
Why: David's design north star for OTP is the full Jobs standard. Splitting it wrongly would let screens pass a visual checklist while the flow, or the craft, gets held to a lower bar. The correct frame keeps one bar over both layers.
Failure mode: When framing the design-standard marriage (Fugu + Jobs), Conatus split it as "Fugu = visual craft, Jobs = journey/UX", understating that Jobs was also a master of visual design (Reed calligraphy class, typography on the original Mac interface).
For /coach-report and any Dash run: do MCP-dependent pulls (Search Atlas, Google Sheets, Calendar) in the MAIN session; delegate only file/CLI/analysis work to subagents. Check a subagent's tool access assumption before waiting on it. Treat CCM STL as unusable until the CloudCRM timezone offset is fixed platform-wide, and never report STL from July data.
Why: Two full delegation rounds were wasted waiting on pullers that could never succeed; the report would have shipped without SEO (repeat of L057) and without CCM if the main session had not redone the pulls. The STL corruption finding upgrades the known Villa-only issue (June) to system-wide, which changes every STL-based alert and coaching metric until fixed.
Failure mode: SUCCESS (with lesson): Dash /coach-report 2026-07-19 — delegated Search Atlas, CCM sheet, and calendar pulls to three subagents; SEO and CCM pullers were fully blocked because spawned subagents do NOT inherit the session's MCP servers (search-atlas and google-workspace tools were absent from their toolsets). Main session had both and pulled everything directly. Also: CCM July speed-to-lead timestamps are corrupt SYSTEM-WIDE (large negative timezone artifacts on every project, not just Villa Sport).
Two mechanics to remember: (1) .gitleaksignore fingerprints are commit-hash-bound, so ANY commit touching a line with secret-shaped placeholder text (like 'Bearer YOUR_API_KEY' in docs) re-mints the fingerprint and re-triggers the scanner — squash merges guarantee this recurs. The durable fix is neutralizing the placeholder so the rule can't match (angle brackets: 'Bearer <your-api-key>'), plus fingerprinting the immutable history. (2) CI checkouts with fetch-depth:0 fetch ALL refs and gitleaks scans all of them — one bad commit on an unmerged branch fails every branch's CI simultaneously; the ignore entry must reach each scanning checkout's .gitleaksignore, which means pushing it to the branch being scanned AND to main.
Why: Symptom (lint-and-type-check job failing everywhere at once) looks like a code regression but is actually the secret scanner; without knowing the two mechanics, the obvious fix (add one fingerprint) only patches one branch and the mole pops up on the next touch of the file.
Failure mode: SUCCESS: Conatus diagnosed a repo-wide CI outage caused by gitleaks fingerprint whack-a-mole — every branch's CI (including main pushes) went red at once from ONE unmerged branch's commit.
Before a swamp send, reconcile the changelog against the week's real PRs (git log origin/main --since since the last issue, filter to feat/ and customer-facing) and write entries for anything unlogged -- do not assume changelog.ts is complete. When the shared repo is contested by a concurrent session, do all changelog authoring in an ISOLATED git worktree (git worktree add off origin/main, symlink node_modules to reuse deps), PR it, merge via gh after CI is green, and run the REAL send from the worktree -- never edit or send from the shared working tree. Date drop-wave entries to the issue's MONDAY anchor (the sender's window upper bound), NOT to the send day: entries dated the Tuesday send day render as future and hide, and getRecentEntries (OS-today) masks this in preflight.
Why: The changelog is the single source of truth for both /whats-new and the email; an unmaintained changelog silently undersells the product to every subscriber. On a machine with concurrent agent sessions, the shared working tree is not safe for a multi-step author+send; a worktree makes the work deterministic and collision-proof. The Monday-anchor date rule is invisible until a dated-Tuesday entry silently disappears from the send.
Failure mode: SUCCESS: Swamp #29 -- two non-obvious operational wins. (1) The digest only reflects changelog.ts, so a big shipping week (~14 features) went out as "2 things" because most PRs never got changelog entries; David caught it. (2) Running the swamp send while another session was actively branch-switching in ~/otp-platform stranded an early commit on a feature branch and made a direct push to protected main a no-op.
For any EJS page whose logic lives in an inline script, add a test that parses every inline script with new Function() and fails the build on a syntax error, then prove the tripwire by running it against the broken version before keeping it. tsc cannot see inside a template and EJS renders a broken string happily.
Why: It was caught only by driving the actual rendered page in a headless browser, not by tsc, lint, or 1204 passing tests. First-run screens are where paying customers land, and a dead one is invisible from the server side.
Failure mode: SUCCESS: Conatus found OTP onboarding Door 4 (Give Ollie everything, PR #232, shipped 2026-07-19) had been completely dead in production since Sunday. onboarding-import.ejs line 214 had an apostrophe inside a single-quoted JS string, so the browser discarded the whole inline script and the file drop, analyze and commit buttons did nothing, silently, for every new customer who picked that door.
David: A2P pages do not allow form fills. Remove the lead form from A2P review landing pages entirely. Keep the SMS disclosure, the Privacy Policy and Terms links, and the registered business name and address on the page, because those are what carriers actually read. Route the quote path to phone plus the GHL chat widget. Reword any disclosure copy referencing "check the box above" so the page does not describe a mechanism that no longer exists.
Why: The form was the architecture of these pages, so this is not cosmetic: it orphans the consent audit trail and the POST /:slug/lead route, and it changes what the A2P campaign registration can declare as its opt-in method. Getting it wrong in either direction risks 10DLC rejection, which blocks SMS for the client entirely. The pattern repeats for every future A2P location page.
Failure mode: Built the Dryer Vent Squad A2P landing pages (Katy, DFW) around a web lead form with an optional SMS consent checkbox, treating that form as the opt-in proof mechanism for A2P/10DLC review, plus a consent audit trail behind it (consent.js and sms_consent_text/timestamp/IP/user-agent captured on POST /:slug/lead).
The unbilled-spend sweep (billing-report step 3b) reads ~/.claude/billing/sweep-exclusions.json and drops those account IDs from the Review tab entirely. Use it for accounts that spend on our Google/Meta but are NOT billed on % of ad spend (white-label / flat-fee). Added 2026-07-23 per David: Phillip Jeffries, M.V. Parker Law, Champy's Chicken (+Nashville), Emily Shalant, Jet City Blinds, J&K Engines, Meyer Law, True Path, GettaMeeting, Lazzara Law, Studstill Firm. Only add IDs here on David's explicit instruction. Separate from the Clients-tab dont_bill mode (which still shows a DO NOT BILL row on the Billing output).
Why: Without a persistent list these 12 accounts (~$24.4K/mo) resurface as "confirm arrangement" in every monthly sweep, wasting David's review time. The file makes the exclusion durable across sessions.
Failure mode: SUCCESS: Billing sweep now has a persistent white-label exclusion list
Three patterns for the OTP recorder. (1) "No way to record a second time" is TWO bugs: the missing UI affordance AND a server ingest that overwrites (ingestTranscriptOneShot did set({transcript})). Fix both or the feature silently destroys data; recording ingests now pass mode:'append', idempotent on the tail so worker retries cannot double-append, while paste/import stays 'replace'. (2) "Mic recorded silence after a long pause" on a phone is a dead MediaStreamTrack (readyState 'ended' or muted): MediaRecorder.resume() succeeds and records nothing. Recover by swapping in a fresh getUserMedia stream, and ALWAYS flush the retired recorder's final chunk BEFORE bumping the server segment number, or the old container's tail bytes land in the new segment and corrupt it. Auto-recover only when the track is provably dead; when it is alive but silent, warn and offer a button, since a genuinely quiet room looks identical. (3) Check git log before rebuilding from a support ticket: background transcription had already shipped that morning, so ticket 3 needed a sequential-to-concurrent R2 upload fix plus a crash fix, not a rebuild.
Why: The recorder holds the customer's only copy of a meeting until upload completes, so every bug in it costs unrecoverable audio, and two shipped in one day. The widget is inline JS in an EJS partial with no import path, which is why it went untested; it can now be tested by rendering the partial, extracting the script, and running it in a vm against a fake MediaRecorder (src/views/partials/meeting/audio-record.test.ts). Use that harness for future recorder changes.
Failure mode: SUCCESS: Claude (OTP dev) fixed three meeting-recording tickets (PR #315) at root cause, and caught that PR #309 from earlier the same day left savedStatus() calling itself on the non-pending branch (a stack overflow that swallowed the save confirmation on the R2-off path) because the recorder widget's inline JS had no tests.
When editing OTP trust/security claims, edit src/config/trust.ts (the file the /trust route imports and ships in the dist image). trust.yaml at the repo root is only the audit copy carrying `# source:` code citations for legal; nothing reads it at runtime. The two had already drifted (legalEntity OTP,LLC vs OrgTP,LLC; a stale lastUpdate; and — critically — trust.ts shipped a prohibited EOS mark "L10" that trust.yaml did not). Always mirror any claim change into BOTH files, and treat trust.ts as authoritative for what the public actually sees.
Why: A trademark-compliance violation (EOS "L10" mark) was live on the lawyer-facing trust page for weeks because the de-EOS pass only fixed trust.yaml, which never ships. Editing the audit copy feels like fixing the page but changes nothing a visitor sees. This is a recurring drift trap worth a permanent CI equivalence check between the two files.
Failure mode: SUCCESS: the public /trust page renders src/config/trust.ts, NOT trust.yaml — the "source of truth" file never loads at runtime
Never position OTP (or Ollie) as an employee. A founder, and even an employee, does not want another employee; they want to know the organization is held. "Employee" also imports employee mentality: waits to be told, owns a lane not the whole. OTP guides the organization. Frame it as the guide/what holds the company, positioned between employee (too small), guru (too big), and co-founder (not really).
Why: Positioning error at the core-promise level: hiring an employee increases a founder's load (managing, explaining, checking); OTP's promise must reduce the holding. Wrong noun poisons every downstream page, price, and demo.
Failure mode: Framed OTP's destination as "the first employee a company hires whose job is to remember" in the product discovery brief.
When a vendor was chosen specifically to absorb an operational burden, do not propose solutions that hand that burden back, even when the vendor documents them. Default order for the consent-screen concern: (1) own the story in product copy (tell users they will see Composio, frame it as the security vault), (2) at most white-label the two or three marquee providers if branding ever matters commercially, (3) never the whole catalog.
Why: Vendor docs happily describe features that shift work back onto the customer. The right frame is the original build-vs-buy decision: Composio IS the OAuth team. Recommendations that quietly re-hire that job in-house waste the subscription and David's time.
Failure mode: Asked how to get OTP branding on the Composio OAuth consent screen, Claude recommended white-labeling via custom auth configs, which means creating and maintaining an OTP-owned OAuth app per provider (Slack app review, Google verification, secret rotation, scope upkeep). David rejected it: the reason OTP uses Composio at all is ONE managed OAuth surface across hundreds of products; a per-provider OAuth app pipeline recreates the exact burden Composio was chosen to eliminate.
The company boundary applies to every artifact that feeds Ollie, not just what is said in the room. When writing a meeting record for the Sneeze It L10, describe mechanisms in company-neutral operational terms (the meeting pipeline, the prep gate, the record path) and keep OTP product identifiers (PR numbers, endpoints, feature ship dates) out of the record entirely. Before pushing any record via agent-record, scan it with the same company-mismatch check the preflight applies to the board.
Why: Ollie's insight renders inside the Sneeze It meeting, so a record contaminated with OTP content produces a contaminated insight automatically, one week later, with no human in the loop. The boundary check must move upstream to where the source is authored or the violation recurs on autopilot.
Failure mode: Dan wrote the 7/20 agent-record for the Sneeze It L10 full of OTP product identifiers (PR #154, the agent-record endpoint, ship dates), so the Ollie Insight generated from it reads as OTP product narrative inside the Sneeze It meeting. David caught it live on 7/27: "Ollie still thinks OTP and Sneeze It are one." The context-bleed boundary was enforced in live speech but not at record-writing time, and the record is the insight's source.
Prep does not end at the Slack brief. Any signal the brief nominates for IDS gets pushed to the OTP board as a ticket (otp-issue.sh, with teamId, AI Army = 065d1d4b) BEFORE the meeting, so David opens the meeting with the issues already loaded and workable in-product. The Slack brief is the narrative; the board is the working surface.
Why: The meeting runs inside OTP, so an issue that exists only in Slack is invisible at the moment of solving. This is the same source-of-truth lesson as L060 applied to prep outputs, not just prep inputs.
Failure mode: Dan surfaced pre-meeting signals (Pulse dark, pipeline shape, HiTone, overdue queue items) only in the Slack prep brief. David at the 7/27 L10: "you should have wrote those signals in OTP, too late now." Signals that deserve IDS never landed as tickets on the AI Army board, so the meeting could not work them in-product.
When David closes an alert as handled, retire the RULE that generates it, not just the instance. For HiTone: edit the CLAUDE.md trigger line and coach-report spec to read "billing confirmed active since Jul 2026, do not flag on spend." In general: trace any recurring alert to the config line that emits it and fix that line, or the alert regenerates forever.
Why: A closed instance with a live trigger is an alert factory. It spends David's attention on the same resolved question weekly and erodes trust in real billing flags.
Failure mode: HiTone billing keeps re-surfacing to David even though he confirmed billing is correct and active (Jun 29, and again 7/27: "HiTone is being billed, you ask about that a lot"). Root cause: the BILLING TRIGGER rule still lives in CLAUDE.md's active-clients list and in the coach-report spec, so any agent that reads the config re-fires the flag whenever HiTone spend appears. The Jun 29 closure was recorded in the rocks file, but the upstream trigger rule was never retired.
Prep must scan for WON/signed revenue explicitly, not just open pipeline: GHL won opportunities across ALL pipelines since the last meeting, plus Proposify signed events, categorized new/expansion/reactivation. Signed expansion revenue is the Q3 headline metric; it leads the scorecard section, and its dollar values get pushed to the manual Expansion tile the same morning. A tile with no automated source still gets its value entered at prep time from the won-deal scan; manual source does not mean no value.
Why: The board exists to catch exactly this: revenue proving or disproving the quarterly thesis. A prep that inventories dead tiles but misses live signed money reports the plumbing and skips the water. David finding revenue wins that his facilitator missed inverts the entire point of the seat.
Failure mode: Prep missed two signed expansion deals (Glo30 and WOA franchise additional revenue) that David saw on the board himself and had to point out at the 7/27 meeting: "you skip that a lot, did you not see them?" The prep scan read open opportunities in one GHL pipeline and the Expansion KPI tile (manual, no values), so signed/won expansion revenue had no path into the brief. The single most strategy-relevant signal of the quarter, expansion revenue from existing accounts, was invisible to prep while five dead tiles got named in detail.
Same-session capture rule for ALL agents: when you witness David build or ship something real (a dashboard, an integration, a signed deal, a process), record it THAT session: a headline line in the daily note, and a todo or KPI value in OTP if money or a rock is touched. Do not wait for the weekly meeting; the builder remembering to report is not a capture mechanism. Dan additionally runs a Shipped This Week sweep every Monday prep as the backstop.
Why: Work done outside meetings is systematically invisible to a meeting-based OS, and the founder's most valuable hours happen outside meetings. Three instances surfaced in one meeting (dashboard, prep signals, signed revenue). Invisible work costs real money: David unknowingly built part of Bogdan's open reporting-cost rock.
Failure mode: David built a WOA dashboard for iCart (central database, less Zapier, better reporting) and the only witness was the AI in that session; no headline, todo, or record reached the operating system. David at the 7/27 meeting: "the only one that knows is you, this is a true failure of the operating system that needs to be reconciled." Same morning, two signed expansion deals (Glo30, WOA franchise) were also absent from every system of record.
Facilitation = operating the product live. The moment a section starts, its artifact moves: an issue under discussion is verified rendering on the board before discussing it; the moment David decides, the ticket is solved with its resolution, the todo is created, the KPI is pushed, in that minute, not at conclude. After every state change, verify the rendered surface. Conclude should be a read-back of changes already made, never a batch of pending writes.
Why: A meeting inside OTP is only real if the product state changes while the humans watch. Deferred writes recreate the mirror-drift problem inside a single meeting, and David cannot trust a board that lags the conversation.
Failure mode: Dan facilitated IDS discussion in chat but did not move the meeting's product surfaces in real time: the Pulse issue being discussed was not visible in the meeting's IDS section, and the already-solved invisible-work issue still sat open on the board. David 7/27: "you should be moving the meeting along like a human, doing things as we work."
The needle test has a step zero: before wiring, fixing, or reporting any KPI, confirm the program/offer it measures still exists in the business, by asking David or checking recent revenue/activity, not config files. A dead program's tile is retired, not wired, and its language gets swept from all agent config at source (L099 pattern) so no agent rebuilds it. Guarantee/T20 program: DEAD as of 2026-07-27.
Why: Config outlives strategy. Wiring effort spent on a dead program's metric is worse than a dead tile, it would have shipped a number that misrepresents the business as still running an offer it killed, and every agent reading the tile would have inherited the fiction.
Failure mode: Dan was one step from wiring the guarantee-clients-retained KPI (emit line written, Tally about to fire) when David said the guarantee program is dead and no longer offered. The wiring work treated the tile's PROGRAM as alive because the config said so; nobody had asked whether the business still runs the thing the tile measures.
When a media element silently stalls (networkState LOADING, error null, no console output), suspect CSP: a 302 redirect from a same-origin playback route to a cross-origin storage URL violates media-src (falling back to default-src 'self') and Chrome blocks it with zero surfaced errors. Diagnose live by attaching a securitypolicyviolation listener and probing a known cross-origin media URL. Fix: add the presigned bucket origin to media-src, computed from storage config at boot, never hardcoded.
Why: This failure mode is invisible by design (no element error, no console noise) and the natural debugging paths (storage probes, presigned URL tests, range requests) all pass, sending you everywhere except the CSP header. Cost about an hour across two sessions; the probe technique turns it into a 2-minute check.
Failure mode: SUCCESS: Claude/Conatus diagnosed why OTP meeting recordings never played in the browser (player stuck at 0:00, click did nothing) while server-side storage checks all passed.
Never invent a person's first name from an email address or initial. If the name is not stated, refer to them by the email address or ask David, and only record a name once confirmed.
Why: Guessed names propagate into client-facing artifacts (emails, dashboards, docs) and getting a client's name wrong damages trust; an email initial is not evidence of a name.
Failure mode: Claude inferred the Drybar Ballston client's first name as "Jodi" from the email address jsterling@sterlingcapitalllc.com and used it in the summary, credentials file, and memory. The client is Julie Sterling.
`ghl.sh update-opp <oppId> <stage>` with the optional value argument omitted hits an inline Python syntax error (`data['monetaryValue'] = ` with nothing after it), so the stage payload is never built: the opp's updatedAt bumps but the stage does NOT change, and nothing errors loudly. Workaround until ghl.sh is fixed: always pass the value explicitly, e.g. `update-opp <id> sql 0` (confirm the opp's current monetaryValue first so you do not overwrite a real value). Always verify stage changes by re-reading the opportunity after the write.
Why: A stage move that silently no-ops corrupts the pipeline of record without any error signal. Both /otp-sales and /sneeze-sales route stage moves through this command; without the verify-after-write habit the bug would have shipped 2 phantom SQL bumps today.
Failure mode: SUCCESS: Sneeze-Sales found and worked around a silent ghl.sh write failure
Before bumping a stage or queueing a task on any reply, check the PERSON and company against the active client list (CLAUDE.md), pepper-clients.md, and known client people, not just the cold-load exclusion at contact creation. Jordan Anderson is a client (Workout Anytime / Proof Fitness): never prospect-touch him on any domain. Replies inside threads a team member already owns (for example Zeynep scheduling) get no David task; the owner handles it. A reply landing in the prospect book is a signal to verify WHO it is, not proof they are a prospect.
Why: The prospects-only law fails at the edges: client people reply from domains that sit in the prospect book because franchisee outreach and client domains overlap (WOA). A wrong SQL bump plus a David task double-touches a client relationship and burns David's queue on non-sales work.
Failure mode: Sneeze-Sales treated Jordan Anderson (Workout Anytime) as a prospect: bumped his opp MQL to SQL on an inbound reply and queued David a respond-within-24h task. David corrected: Jordan Anderson is a client. Also queued David a confirm task on Lindsey (Fitness Factory) when Zeynep already owns that thread.
When an OAuth integration fails silently, verify each layer with direct probes instead of reasoning from app behavior: (1) print env var names with cat -v, since a pasted quote becomes part of the variable NAME and the app reads undefined (Railway kv showed "GOOGLE_CALENDAR_OAUTH_CLIENT_ID with a literal leading quote); (2) test client credentials against the provider token endpoint with a bogus auth code, since the error distinguishes exactly: invalid_client means bad id/secret, invalid_grant Malformed auth code means credentials are VALID; (3) treat the database as ground truth for whether a flow completed, since users saying connected can mean a different surface (Composio integrations page vs the calendar card).
Why: Three probe layers turned what could have been hours of guessing into minutes: no callback in logs proved the flow died at Google, cat -v exposed the quote character, and the bogus-code token probe verified the replacement secret BEFORE the user retried, avoiding another failed round trip. Reusable for every OAuth integration OTP adds (Microsoft, Zoom, future providers).
Failure mode: SUCCESS: Claude shipped Recall calendar auto-join and debugged two invisible OAuth config failures the same afternoon (quoted env var name, invalid client secret)
When graduating a feature out of Labs, grep the whole repo for the feature key and isFeatureEnabledForOrg calls before merging; every gate (page, API, scheduler, MCP tools) must come off in the same PR. A fail-closed flag check on a deleted key silently disables the feature for everyone, which reads as random 404s, not as a flag problem.
Why: Fail-closed gating is correct security posture, but it means catalog removal IS a kill switch. The bug shipped invisible because the page worked while the API did not, and the generic 404 body hid the cause; the honest error message shipped hours earlier (names the workspace and feature) is what made the real diagnosis possible.
Failure mode: Projects went GA (PR #347 removed it from the Labs catalog and the page gate) but the API routes kept gating on the removed key; isFeatureEnabledForOrg fails closed on unknown keys, so all project CRUD 404'd platform-wide for two days until Dawson and David hit it
Any otp-platform endpoint that scopes data or authz by getAuth(request).userId is broken under impersonation. The rule: gate and scope by the EFFECTIVE viewer (request.impersonation.as when active, else auth.userId), return/audit with the RAW session id, and for context-pinning actions (org switch) re-issue the impersonation cookie via startImpersonation rather than moving the admin's own cookies. When reviewing or writing any new route, grep for getAuth(request).userId used in a WHERE clause — each one is a latent impersonation bug.
Why: Impersonation is how David supports customers (view-as Tom, Kristen, etc.). Every raw-session usage silently shows the admin's data under the customer's banner or 403s the customer's own surfaces — a privacy leak in one direction and a support dead-end in the other. Fixed instances: dashboard (2026-06-02), PRs #381, #382, #384 (2026-07-28).
Failure mode: SUCCESS: Dan identified a recurring defect class in otp-platform — four separate surfaces broke under super-admin impersonation in one day (portfolio pages listing the admin's portfolios, portfolio API 403ing "Could not load team", the sidebar org label showing the admin's org, and the org switcher 403ing "Could not switch organization"), all with the same root cause.
Do not treat a missing or small wallet balance as a signal of anything. An org with no wallet is simply not an active OTP user, so its balance says nothing about product readiness. New orgs are seeded with $25 of credit to incentivise starting, so a funded wallet is the default going forward rather than a hurdle. When assessing whether a metered feature is usable, filter to orgs with real activity and check the NON-wallet prerequisites, since those are the ones that actually gate anyone.
Why: Reporting wallet balances as blockers manufactures work out of the ordinary shape of the user base: most rows in that table are dormant signups, not stuck customers. It also buries the prerequisite that does bite, because a real blocker listed next to four fake ones reads as one item in a list instead of the single thing to fix.
Failure mode: Flagged orgs with low or missing wallet balances as a readiness problem for OTP scheduling, treating wallet funding as a live blocker worth David's attention.
Before retiring a scarcity, cohort or badge claim, count each cohort in the database and check whether one phrase names several cohorts. Shared wording is not shared meaning. If counts contradict the instruction's premise, return to the human with the numbers instead of executing literally.
Why: Broad approval rests on an assumed premise. When the premise is partly false, literal execution silently destroys value: a deleted live offer throws no error, it just yields fewer signups. One query per cohort is cheaper than a loss nobody detects.
Failure mode: SUCCESS: Claude verified cohort counts before executing an approved codebase-wide sweep of OTP's "first 50" claim. It was closed for signups (50/50) but live for Founding Publishers (45/50) and Founding Partners (6/50). Literal execution would have deleted two accurate live offers.
For multi-agent feature builds: (1) research agents return structured briefs before any code, and briefs override the spec when they conflict (two detectors were impossible as specced: meetings have no booked-duration column, decisions are not rows). (2) Put all shared-file wiring (server.ts) in ONE sequential final task so parallel agents never collide. (3) Always run an independent fresh-context diff review before committing: it caught two honesty blockers the per-task verifications missed (ratified moves netting costs away; org-wide gains summed over subtree-scoped costs). (4) Discovery worth acting on: subscriptions.plan_rate is never written by any code path in otp-platform, so any revenue/cost feature reading it ships dark until billing populates it.
Why: The per-task agents were all green individually; only the cross-seam review found the invariant violations. Repo guard tests (private-issue leak scan, blueprint coverage) also fired exactly as designed, proving lint-style guard tests catch what unit tests cannot.
Failure mode: SUCCESS: Claude shipped OTP Impact Phase 1 (PR #427) via 11-agent build: parallel research briefs, wave execution with pure-function cores, then adversarial diff review before commit
When adding a NEW page to an existing app, open a sibling page that already ships (for OTP admin surfaces, /admin/support) and copy its outer container, top padding and control classes verbatim before writing any markup. Do not hand-roll spacing from DESIGN.md tokens alone -- the tokens do not tell you the page-level offsets that keep content clear of the fixed header. For any default that is a money amount, confirm the number rather than inferring it from the option list order.
Why: A new page laid out from first principles looks subtly wrong in ways the author cannot see without loading it: header collisions and control scale only show up in a browser, not in typecheck, lint or design-lint, all of which passed. Copying a shipped sibling inherits every page-level decision already made and reviewed.
Failure mode: Built /admin/join-link with hand-rolled Tailwind layout (max-w-3xl, custom padding, custom input classes) instead of copying the container and control classes from an existing admin page. Result: the page header collided with the fixed top nav so the title was unreadable, and the form controls were oversized versus OTP's 32px control scale. Also picked $50 as the default starting credit without asking; David wants $25.
When promoting any OTP Labs feature from beta to live, do THREE things, not one: (1) flip `status` in src/shared/lab-features.ts; (2) grep src/views for the feature's `surfaceUrl` -- if the ONLY link is the Labs-injected rail item, add a permanent entry to layouts/main.ejs in BOTH the `_sbItems` array and the mobile settings menu; (3) grep the page for stale "this is a Labs feature" banner copy pointing at a /settings/labs toggle that graduation just removed. Also verify any second, independent gate (e.g. an env check like recall-calendar.ts calendarIntegrationEnabled) and make the registry copy match what is actually configured in production -- check `railway variables --kv` rather than trusting the existing description.
Why: Graduation looks like a one-line status change and is not. The rail item, the page's own Labs banner, and any env-based second gate all key off the old state, so a naive flip can make a feature LESS reachable than it was in beta while appearing to ship it. Confirmed live: PR #437 shipped the flag plus the nav entry together, and the promoted page rendered correctly with the Calendar section visible and no Labs opt-in.
Failure mode: SUCCESS: Claude caught that graduating an OTP Labs feature from `beta` to `live` silently DELETES its left-rail nav item, which would have shipped calendar auto-join into being unreachable. `getOrgLabNavItems` (src/services/lab-features.ts) filters on `f.status === 'beta'`, so only beta features get a rail item injected. /settings/meeting-presence had no other link anywhere in src/views, so flipping the flag alone would have removed the only way to navigate to it.
Four reusable rules for /coach-report and any Dash run. (1) When `meta-ads.sh token-check` returns Valid:False, do not report the portfolio as quiet: the CCM Stats sheet "Ad Spend" column carries the Meta-side spend for every call-centre project, so it is a working fallback for spend and leads. Label the source on the card. (2) `mcp__google-workspace__get_events` silently caps at max_results and truncates the NEWEST events, not the oldest. A 90-day pull capped at 250 returned nothing after Jun 30 and would have reported "no client meetings in July." Always check the max date in the response against time_max and re-pull in narrower windows. (3) Never average Google cost-per-conversion together with CCM cost-per-booked-appointment in a franchise network benchmark. Group the cost metric by source, keep only the largest comparable group, suppress the table when fewer than 2 comparable peers, and name the metric explicitly on the card. (4) When CCM shows appointments booked but zero shows, that is unconfirmed data, not a zero show rate. Do not report show rate; ask for confirmation instead.
Why: Each of these silently produces a confident wrong number in a client-facing artifact. The calendar cap fabricates a churn signal, the mixed-metric average makes a healthy location look 60x worse than a peer, and a missing Meta token makes a $136K/month portfolio look dead. Cards go to clients, so a wrong number costs trust directly.
Failure mode: SUCCESS: Dash /coach-report 2026-08-03 — shipped 50 cards with Meta Ads fully down, and caught two silent data traps that would have produced wrong client-facing numbers.
Refines the CCM-fallback rule captured in L125. The CCM "Ad Spend" column matches Meta actuals closely (Rockstars Frisco $453.23 vs $452.54, Okeechobee $150.07 vs $151.09, China Grove $384.79 vs $388.21) but ONLY when the sheet has a complete row for every day in the window. WOA Winder read $133.98 against Meta's $303.08 because the project stopped appearing in the sheet after Jul 30. So: before using CCM as a Meta proxy, count the daily rows per project across the window and flag any project with fewer rows than days. A project that silently drops out of the sheet reads as a spend and lead collapse when it is a recording gap. Second lesson: do not attribute a lead decline to an ad-platform outage without checking delivery. Meta REPORTING was dark to us from Jul 27, but Meta DELIVERY was fine (portfolio leads only -9% week over week, spend flat), so the much larger per-project CCM declines were a recording or routing artefact, not an ad problem.
Why: The fallback is genuinely good enough to save a report, but only with the completeness check. Without it, a project that falls out of the sheet produces a fabricated churn signal that a coach would take to a client. And blaming a visible outage for an invisible decline is the easy wrong answer that stops the real investigation.
Failure mode: SUCCESS: Dash — Meta token restored 2026-08-03, and the CCM fallback used during the outage was measured against real Meta data once it came back. The fallback was accurate to within 1% on 3 of 4 spot-checked projects but understated WOA Winder by 56%.
Upload the coach report to Drive exactly ONCE per run, at the very end, after the validation sweep passes. Never upload an intermediate build, even when the intent is to re-upload a better one later, because the shared folder is read by coaches the moment a file lands. If a rebuild is genuinely needed after an upload, trash the superseded file in the same action rather than leaving both. Use update_drive_file with trashed=true (recoverable); neither Drive MCP exposes a hard delete. Verify by file size against what was generated locally before trashing anything, and get David's approval first since the folder is shared.
Why: A shared folder is a publishing channel, not a working directory. Every extra file is an opportunity for a coach to open the wrong numbers and take them to a client. The same duplicate pattern already exists on 2026-07-11, 2026-05-25, 2026-04-06 and 2026-03-10, so this is a recurring habit rather than a one-off.
Failure mode: Dash uploaded three files to the shared Coach Reports Drive folder in one morning, as the data improved from Meta-blind to full to corrected. Zeynep has reader access, so two of the three were coaches' paths to stale client numbers until David approved trashing them.
Never read a green tile as evidence that the agent named on it is alive. Decoupling a KPI from its agent protects the number but removes the number's ability to report the agent's death, so agent liveness needs its own signal: check the shared-state file mtime alongside the tile before making any seat decision. When a seat review comes up, present tile value and agent liveness as two separate lines, never one.
Why: A seat can look healthy and be vacant. Here the 32.6% measured Erica and Amanda's human performance, not agent output, so the tile would have argued against repurposing a seat that had already stopped running. Seat decisions made on decoupled tiles are made on the wrong evidence, and the error is invisible because the metric is genuinely accurate about the thing it actually measures.
Failure mode: Dan argued in the 8/3 meeting that Arin's 32.6% appointment rate showed the Arin seat was working, so repurposing it was not urgent. That inference was wrong. The Arin KPI was deliberately decoupled from Arin-the-agent in June (source kind composio_action, reading the CCM sheet directly) so it would survive a repurpose. Meanwhile arin-latest.md had been stale for 287 hours, about 12 days. The agent was dark and its tile was green the entire time.
Before escalating any data problem to a vendor, query the vendor's live API for the object in question and reconcile it against our own source of truth. Here the API showed the project had held exactly 8 keywords since creation (first_keyword_added_at), so nothing was ever lost, and our own cluster map defined 28 pillars rather than the 43 the KPI asserted. The genuine vendor issue turned out to be a different and much larger one: 19 of 27 projects silently blocked on NO_QUOTA, including six client projects. Split the ticket: send the vendor only what the vendor actually owns.
Why: A false claim to a vendor costs credibility and burns the support cycle you need for the real issue. It also hides the internal defect behind a vendor excuse, so it never gets fixed. In this case verifying first both protected the vendor relationship and surfaced a client-delivery problem nobody had noticed.
Failure mode: An item routed to a vendor (Search Atlas) as "the vendor dropped our 43 tracked keywords to 8" was accepted at face value from an L10 without checking live vendor state first. Sending it would have been a false data-loss claim against the vendor.
When a measurement pipeline is broken, check whether the metric itself is the defect before repairing the plumbing. Beacon's KPI was "pillar keywords in top 10 (of 43)" on a domain whose pillar pages had launched six weeks earlier. That number reads 0 for quarters regardless of whether the work is excellent or abandoned, so repairing the vendor feed would have restored a tile that still said nothing. Test any KPI with two questions: can this number move within one review cycle, and if it moved would I trust what it means? Beacon failed both (its actual top-10 rankings were unrelated junk queries). The fix was to change the instrument to Google Search Console, which is free and already authorized, and the metric to pillar-cluster impressions, which moves weekly. Rename the existing tile in place via PATCH /api/v1/kpis/{id} rather than creating a new one, so the seat keeps its position and no dead tile lingers on the chart.
Why: A KPI that cannot move is not accountability, it is decoration, and it quietly consumes the attention a real metric would earn. This one also hid a genuine finding for six weeks: the pages were being surfaced 3,479 times and earning one click, which is a titles and intent problem no top-10 counter would ever reveal. For OTP specifically this is a Constitution matter, since the axiom is that an org reconciles what it says with what it does; running a fake KPI on OTP's own chart dogfoods the disease the product exists to cure.
Failure mode: SUCCESS: Beacon rewired from a vendor-dependent KPI that had never once produced a real number to a free first-party one that works.
After editing any .ejs or CSS in otp-platform, run `node scripts/design-lint.mjs` as well as tsc/tests/smoke:render. It has no npm script, so the standard local verify passes while CI's lint-and-type-check job fails. It counts DESIGN.md violations per file against scripts/design-lint-baseline.json and fails on any increase. When it fires, fold the new selector into the existing rule instead of running --update on the baseline.
Why: A view change can look fully verified locally and still red CI, costing a full round trip. On PR #459 a focus ring (not a resting shadow) tripped css-resting-shadow 6 -> 7; folding the picker into the existing input and focus rules fixed it and was what DESIGN.md wanted anyway, so the gate caught real design debt rather than noise.
Failure mode: SUCCESS: Claude found the otp-platform verify recipe is incomplete for any view change
A bounce is never proof an account is fake. Before deleting any account, check three things: (1) does an auth-provider user exist, (2) when did they last sign in, (3) what data and pending invites hang off them. Then read the bounce SHAPE: a hard bounce in ~3 seconds means the domain or mailbox does not resolve, which usually points to a typo in an otherwise real address; a bounce 12-14 hours after send is a soft bounce from retry exhaustion (full mailbox, suspended account, reputation deferral) and the person is real. Report the split and get explicit confirmation before deleting anything that is not provably fake.
Why: Bounces cluster on real customers with mistyped addresses, not on fabricated signups. A bad email is a recoverable lead: One Jump's owner mistyped a subdomain and became unreachable for seven weeks, while the colleague he invited was reachable at the correct corporate domain the whole time. Treating "bounced" as "fake" deletes paying-customer-shaped signups and destroys the only evidence needed to win them back. Deletion is irreversible; suppression achieves the actual goal (stop the bouncing) at zero cost.
Failure mode: SUCCESS: Claude caught that a "delete these fake bounced accounts" request included a live customer org. Of three addresses flagged as fake, only one was: test@abctest.com had no Clerk user at all. business@mail.onejumpinc.com was Jessup Jong, owner of the live "One Jump" org, last signed in 3 weeks earlier, with a chart, team, meeting, and a pending invite to a colleague expiring in 5 days. Deleting as asked would have destroyed a real signup mid-onboarding, irreversibly.
Never write a file that other processes read with open(path,'w') plus write. Write to a temp file in the SAME directory via tempfile.mkstemp, then os.replace(), which is atomic on POSIX so readers see the old file or the new one and never a half-written one. Wrap it so the temp file is unlinked if the write throws. When you find one writer of a shared file doing this, grep for every other writer of the same file and fix them all; fixing one leaves the race intact. Verify with a concurrency test that reproduces the failure on the old code and shows zero on the new, rather than assuming. Note that mkstemp yields 0600, which tightens a credential file from the usual 0644 and is an improvement.
Why: A truncate-then-write race on a shared credential file is invisible in normal use and only appears under concurrency, so it presents as a flaky, unreproducible alarm on a load-bearing data source. That trains the operator to dismiss the monitoring. Worse, the failure mode is indistinguishable from a genuinely revoked token, so the real alarm and the false one look identical. Atomic replace removes the entire class of failure at the source rather than papering over it with retries.
Failure mode: SUCCESS: Radar traced an intermittent phantom "Google Ads token invalid" alarm to a non-atomic credential write, not to the token or the network. Three separate processes write ~/.claude/mcp-google-ads/google_ads_token.json (google-ads.sh, google_ads_server.py, auth_setup.py) and all three used open(path,'w') followed by write. That truncates the file first, so any concurrent reader json.load()s a partial document and throws. Because get_access_token ends in 2>/dev/null, the exception surfaced as an empty token and got reported as a dead credential. A hammer test measured 178 partial reads out of 400 writes, a 44% failure rate under contention.
Before repainting any utility class in a view, grep it in src/styles for an !important clamp selector and move that hook in the same commit. Verify repaints with a pixel diff against a before-screenshot, never on a green lint run alone.
Why: A clamp keyed on a class name is an invisible coupling no test or linter can see, and cleaning up that class name is exactly the change that breaks it.
Failure mode: SUCCESS: Claude found OTP dashboard-daily hairline-row grammar was produced by an !important clamp in input.css keyed on the row wash class name. Repainting that class onto tokens silently detached every row from the clamp, restoring borders and radius and shifting padding. Tests and design-lint both stayed green.
For the Sneeze It AGENCY lane, outreach is account-based, not volume-based. The outreach that produced actual revenue (Beem, now a multi-location client) was individually researched and drafted into David's Gmail, one account at a time. Build the big list for SELECTION, not for sending: a master Google Sheet where every row carries enough context to be decidable in ten seconds, David marks an X on the rows he approves, and each approved account then gets real research and a bespoke Gmail draft. GHL becomes the record of the resulting deal, not the send engine. Do not templatize, do not sequence, do not meter.
Why: Sneeze It accounts are worth $44k to $136k a year each (HiTone is 8 locations at $5,472/mo; WOA is 13 club accounts at ~$136K/yr expansion). At that account value, thirty minutes of research per email is trivially correct economics, and 500 templated sends optimize the wrong variable entirely. The data agrees: 221 contacts sit parked at MQL against 11 that ever advanced, and the templated 19-contact batch sent 7/27 produced nothing measurable in nine days, while bespoke research-led outreach produced a real client. IMPORTANT SCOPE NOTE: this INVERTS the existing learning that says default to a multi-touch sequence rather than hand-personalized copy. That rule was derived from OTP coach outreach at list scale and remains correct there. It does NOT transfer to Sneeze It agency prospecting, where the ICP is small (roughly 800 franchisors) and each account is large. Check which lane you are in before choosing the artifact shape.
Failure mode: I diagnosed the Sneeze It cold-outreach problem as insufficient volume and proposed waves of 500 templated emails metered out through GHL at 25/day. I was optimizing for throughput when the evidence in front of me said throughput is exactly what has never worked.
Do not treat an open GHL opportunity in the Sneeze It Sales Funnel as evidence of a live relationship. David confirmed 2026-08-05 it is a graveyard: 221 of 234 open opportunities are parked at MQL and have never moved a stage. Only a WON opportunity (they became a client) is a hard block. Open, lost and abandoned opportunities are just a dated touch and should fall through to the COOLING or REACTIVATION tier based on how long ago that touch was. Separately, out-of-ICP companies (equipment manufacturers like The Abs Company, food, dental) belong in a dedicated ICP-exclusion file, not in do-not-blast.md, which exists for compliance and unsubscribe obligations.
Why: Blocking on open opportunities inverted the meaning of the pipeline. A stage that nothing ever leaves is a record of who we once imported, not who we are talking to, so using it as a suppression signal removes the exact brands most worth re-approaching. Reading a dead pipeline as a live one is the same class of error as reading a capped API response as a complete one: the data is technically present but means something different from what its name suggests. Check whether records in a stage actually move before you let that stage gate behaviour.
Failure mode: The suppression builder treated any OPEN opportunity in the Sneeze It Sales Funnel as a hard block, on the assumption that an open opportunity means a live deal that must not be cold-emailed.
David explicitly authorized DIRECT SENDING for the master-prospect-sheet lane on 2026-08-05, after I raised the risk and he reaffirmed. The human gate moved rather than disappeared: it is now the X he types in column A of the Sneeze It Master Prospect List, which is a per-company approval made before any email exists. Hard limits he set: max 30 sends per day, from dsteel@sneeze.it, one scheduled run per day, with a manual override command to run on demand. Every send stamps the sheet with sent status and date so a no-reply follow-up can fire ~30 days later. L096 is NOT repealed anywhere else: the `/sneeze-sales` GHL harvest still never applies a sequence tag and still never sends. Check which lane you are in.
Why: The reason behind L096 was never "an agent must not send" as a principle; it was that no cold email should reach a client or someone David is already talking to. A per-row human X satisfies that intent more directly than sequence enrollment did. The risk that remains is different and worth naming to a future operator: the X approves the COMPANY, not the SENTENCE, so nobody reads the email before it goes. That makes research accuracy and suppression freshness load-bearing in a way they were not when David hand-sent drafts. Rebuild the suppression list before any run, and never send to a BLOCKED domain, a bounced address, or an unsubscribe, regardless of what the sheet says.
Failure mode: Standing rules said the Sneeze It outreach agent never sends email: L096 required David to personally vet each contact and enroll them in the sequence himself, and the morning-pass floor says "prepare drafts, never fire them." Under the new sheet-driven process those rules would block the whole loop.
In-home care and senior care FRANCHISORS are IN ICP for Sneeze It as of 2026-08-05 (David: "low target and worth a try"). The qualifying trait is a multi-location franchisor with a real lead-generation budget, not a membership billing model. Do not hold or flag them. Because David framed it as a try rather than a conviction, tag these sends so their reply rate is measurable separately from fitness instead of being blended into one number: an experiment you cannot read the result of is not an experiment. The residential end (nursing homes, assisted living facilities) is still untested and was not part of this ruling.
Why: The ICP was written as a description of existing clients, who happen to be membership gyms, and I applied it as a boundary on who could ever be a client. Those are different things. What Sneeze It actually sells is paid lead generation plus a call center that works the leads fast, and any franchisor buying leads for local operators has that problem regardless of how they bill their customers. Watch for this shape generally: an ICP inferred from the current book will keep reproducing the current book, and the operator is usually the one who can see past it.
Failure mode: I flagged seven in-home care franchisors (Home Instead, Assisting Hands, Always Best Care, Comfort Keepers, Senior Helpers, Synergy HomeCare, FirstLight) as out of ICP and held them from sending, reading the Sneeze It ICP as membership-model fitness and wellness only and treating the inherited Nick-era "senior living" exclusion as covering them.
When a resource has an access rule on its primary page, extract that rule into ONE shared helper (canReadMeeting in services/meeting-read-access.ts) and apply it at every read surface, list AND single-row, instead of letting each endpoint hand-roll a partial check. Also: Ollie chat executes tools via app.inject with the user's own session (makeSessionOtpFetch), so fixing the HTTP endpoints automatically scopes Ollie's chat answers — no separate AI-context fix needed. Org-API-key callers carry no member row and stay org-wide by design.
Why: Access-control drift is a leak class, not a one-off: every new read surface (followups, exports, recordings, captures panel) shipped without the team gate because the rule lived inline in one page handler. A single source of truth makes the next surface safe by default, and knowing Ollie chat rides the user's session means endpoint-level authz is the one place to fix AI data exposure too.
Failure mode: SUCCESS: Conatus — root-caused critical R3V ticket 45f74cff (users could open any meeting): /l8/meeting/:id had a team/attendee/creator gate for months, but the followups page, all per-meeting read APIs (transcript, exports, recordings, agenda, headlines, share, SSE), and the meeting-captures panel only checked the 'restricted' flag — the gate existed but was never shared.
Adding a scheduled command to run-claude.sh takes THREE edits, not one: (1) APPROVED_COMMANDS, (2) a BUDGET case entry, (3) a RESOLVED_PROMPT case entry that expands the slash command into a literal instruction. The file contains SEVERAL separate `case "$PROMPT" in` blocks, so never anchor an insert on a command name alone; anchor on something unique to the target block (for the resolver, the line `RESOLVED_PROMPT="$PROMPT"` immediately above its `case`). Verify by extracting the resolver block and running it with the prompt as input, rather than by reading the diff: after the edit, `/outreach` must echo a real instruction string, and at least two pre-existing commands must still resolve to prove nothing was clobbered. `bash -n` passing proves only syntax, not that the edit landed in the right block.
Why: Both errors share a shape: a change that looks complete because the part you touched is correct, while the part you did not know about is untouched. The whitelist edit read as done because the command appeared in the file. The anchored insert read as done because the diff showed the right text. Neither was checked against behaviour. The saving grace was that run-claude.sh fails LOUD on an unresolved command, writing FATAL to the run log and an entry to alerts.log, so the job did nothing and said so rather than reporting success. That is the pattern worth copying into every scheduled job: an unconfigured job must be distinguishable from an idle one.
Failure mode: I added /outreach to APPROVED_COMMANDS in run-claude.sh and declared the scheduled job ready. It fired at 10:12 on 2026-08-06 and did nothing: approving a command is only half the wiring, and headless bare mode cannot resolve a slash command from the commands directory. Then, fixing it, I anchored the insert on the string ' "/otp-sales")' and landed the resolver inside the BUDGET case statement instead of the RESOLVED_PROMPT one, which would have left the original bug in place while also giving the job no budget entry.
A blocker on one item never blocks the others. Work every item that can be worked, exhaust every avenue on the ones that cannot, and only then report. Park the genuinely undecidable ones and keep going. A question for David is fine and should be raised, but it is raised ALONGSIDE completed work, never instead of it. Concretely for any queue-processing run: partition the queue into workable and blocked at the start, finish the workable set completely, and treat "I found a bug in my own tooling" as work to do inside the same run rather than a finding to report at the end of it.
Why: David's scarcest resource is his attention, and a report that says "here is why nothing happened" spends it without buying anything. The blockers were real but they applied to 2 of 12 rows; I let them set the pace for all 12. The deeper error is that stopping felt like diligence: I had genuinely found real problems, so surfacing them felt like the responsible act. It was not, because surfacing a problem is only half the job when the other ten items were sitting there workable the whole time. Exhaust the possible before you escalate the impossible.
Failure mode: The outreach run hit two items needing David's ruling (Powerhouse Gym, Fastest Labs) and a bug of mine blocking three more, and I reported all of it and stopped. Nine rows sat unsent while I wrote a summary. Worse, I had Clay-verified addresses for two of those companies (Annie Long at Senior Helpers, Jennifer Chasteen at Synergy HomeCare) already in hand from the previous session and did not use them. I presented blockers as a reason the batch could not proceed, when they were only reasons those specific items could not proceed.
When David asks for a report he can distribute, assume the audience is internal leadership, not the people being measured. Per-agent performance comparisons belong in that document at full detail with no warning attached. Do not volunteer to redact, soften, or produce a second sanitized version of performance data unless David names an external or team-wide audience. If the audience would genuinely change the content, ask which audience up front before building, never as a caveat appended at the end.
Why: Manager-level reporting exists to name who is converting and who is not. Attaching a sensitivity warning to that treats normal management reporting as a risk, and hands David an extra decision he did not ask for while he is trying to walk into a meeting. It also spends the final impression of the deliverable on a hypothetical instead of the findings. General shape: resolve audience before writing, not after.
Failure mode: Arin built the 14-day call center review Google Doc for David, then closed by flagging the Amanda vs Erica per-caller comparison as sensitive and offering to cut a sanitized second version before distribution. David corrected: the doc is internal and was never going to the call team. Both the hedge and the offer of a redacted variant were wasted.
A franchisee client does NOT block the franchisor, and the two are different companies on different domains (powerhousegymbridgeport.com vs powerhousegym.com). David 2026-08-06: "we do powerhouse bridgeport one location not corporate so if this is the corporate location good to go." Reverse the instinct entirely: an existing franchisee relationship is the strongest possible ASSET in a franchisor email, because it is proof delivered rather than claimed. Lead with it. The direction that matters is one-way: never cold-email a franchisee of a brand whose CORPORATE relationship we are mid-conversation with, and never email the specific client location, but corporate remains open and warmer than cold. Check which entity a domain actually belongs to before assuming brand-family contamination.
Why: I generalised one true fact (do not email a client) into a rule that would quietly delete the addressable market. Franchising is precisely the structure where one brand contains many independent buyers, so brand-level blocking is the wrong shape for this ICP. The asymmetry is worth remembering: blocking a franchisor to protect a franchisee costs a six-figure account and protects nothing, because the franchisee is not the one receiving the email. It also throws away the single best proof point available, which is that the brand already works with us somewhere.
Failure mode: I held Powerhouse Gym corporate (powerhousegym.com) as a policy risk because Powerhouse Gym Bridgeport/Stratford is an active Accelo client on that brand, and I proposed hard-blocking any brand domain where an active client sits anywhere in the system. That rule would have blocked every franchisor whose franchisee we already serve, which is most of the best targets we have.
Separate first-party link tracking from email-provider tracking. A tokenized link pointing at our own domain (orgtp.com/join/sales/<token>) is fully click-trackable no matter which mailbox sent it -- the prospect's browser hits our server and we stamp it, no provider cooperation needed. Manual personal-mailbox sending costs only the steps BEFORE the click: email opens (no pixel) and bounce/delivery visibility. When someone says "it came from a personal email so we can't track it", check whether the link destination is ours before agreeing.
Why: The two are routinely conflated, and the conflation kills features that would have worked. Here it nearly closed a ticket whose core ask (who clicked, who signed up, success rate) was fully buildable. The residual limitation is real but narrow, and worth stating precisely rather than as a blanket "not trackable": a dead address renders identically to a live address that ignored you, and the mint timestamp is not the send timestamp.
Failure mode: SUCCESS: Claude -- a join-link analytics ticket was nearly dropped on the false premise that sending from a personal mailbox makes clicks untrackable.
Before treating an OTTO queue as work to approve, check three things: (1) autopilot_ai_settings limits, where 0 means nothing is ever generated regardless of autopilot_is_active; (2) whether pending rows actually carry a non-empty recommended_value; (3) is_active as the approval flag, since the API's status field describes the current value's condition (compliant / invalid_length) and the status query param is accepted then silently ignored. When changing autopilot settings, always read-modify-write the complete settings object, because a partial PATCH risks dropping the other issue types.
Why: Counting pending tasks as available SEO wins overstates the work by orders of magnitude and invites a blanket approve. On orgtp only 171 of 779 pending titles were genuine defects; approving all of them would have overwritten 600+ pieces of good human copy with generic AI copy, damaging the pillar pages the Beacon impressions KPI depends on. The zero-limit config also means every client site, including Workout Anytime at 2,233 pages with issues, has an OTTO that has never run.
Failure mode: SUCCESS: Beacon found that OTTO's "pending task" count is not a backlog of approvable SEO fixes. On orgtp.com ~7,100 pending tasks all had an empty recommended_value, because autopilot_ai_settings carried limit 0 for every issue type on all 19 Search Atlas projects. autopilot_is_active reported true the whole time, so the config looked healthy while generating nothing.
When calculating any client credit or make-good on misspent budget, net out the value of what was actually delivered before proposing an amount. Formula: credit = spend under review minus (conversions delivered x a defensible cost per conversion). State which benchmark rate is used and why. Never default to crediting 100% of spend when the spend produced results.
Why: Crediting gross spend overpays the client and understates the work that did land. It also sets a precedent that any misallocation equals a full refund regardless of outcome. Netting delivered value is both fairer to Sneeze It and more defensible to the client, because it shows the math instead of a round apology number.
Failure mode: Drafted a client make-good credit at 100% of the misdirected spend ($2,501.53), treating the entire amount as a total loss. The spend was not a total loss: it delivered 6 real conversions to the client, and the draft gave that value away for free.
When changing the arity or shape of a function that other code passes a hand-rolled structural stub into, grep every caller for stubs BEFORE trusting typecheck. In otp-platform, registerOtpTools is called with `as unknown as Parameters<typeof registerOtpTools>[0]` in three places (the HTTP route, Ollie's collector in src/services/ollie-tool-registry.ts, and the parity test). That cast makes any missing method invisible to tsc: migrating otp-tools.ts from server.tool() to server.registerTool() typechecked clean while Ollie's collector, which only implemented .tool(), would have collected zero tools and removed every OTP tool from the chat box at runtime. Rule: a structural stub behind an `as unknown as` cast is an untypechecked interface. Treat it like a second implementation and update it in the same PR. Also: when two hand-maintained lists answer the same question (Ollie's WRITE_TOOLS/READ_TOOLS vs the MCP readOnlyHint/destructiveHint annotations), tie them together with a test rather than trusting them to be edited in step; ours had already drifted twice (discover_intelligence POSTs and INSERTs but was classified read, so it ran with no confirm card; sync_rules_to_file only rendered text but was classified write). Finally, verify a guard by breaking the invariant and watching it fail, not just by watching it pass.
Why: The failure mode is invisible to every automated check: tsc passes, all 3232 tests passed before the shim was fixed because the shim's own test used the same stale stub. It only surfaces as "Ollie can suddenly do nothing" in production. The generalizable form is that casts convert compile-time contracts into runtime hopes, and a codebase with structural stubs has as many implementations of an interface as it has stubs.
Failure mode: SUCCESS: Claude migrated all 56 OTP MCP tools to registerTool with directory annotations, and caught a shim that would have silently emptied Ollie's tool registry in prod.
Before diagnosing an OTP support ticket, identify the exact surface the reporter was on, then verify the reported cause can even occur in that state. From the 8/7 R3V batch: (1) "reassign meeting to another team" read as a feature request but was a creation bug — the UI offers "No team (personal)" and POST /meetings silently substitutes the Leadership Team; reassignment already worked via PUT /meetings/:id. (2) "Ask Ollie can't file a ticket" was mistaken identity — OTP has TWO assistants: Ask AI (corpus-only, no tools, so its refusal was truthful) and /ollie-chat (full MCP registry, has submit_ticket). (3) A plausible cause for a typing-freeze was ruled out by a precondition check: the transcribing banner only renders when the meeting has NO transcript, and the reporter had already generated Ollie insights. Say so honestly rather than shipping a fix under a false claim.
Why: Users describe symptoms, not causes. A plausible cause that survives no precondition check produces a fix that fixes nothing while closing the ticket. Checking which surface and which state rules candidates in and out cheaply, and turns "feature request" into "bug" often enough to change what actually gets built.
Failure mode: SUCCESS: Claude — three of six R3V support tickets had root causes different from what their titles said, and only reading the actual surface found them
Before designing any adoption, activation or gamification feature, query production for the funnel FIRST and count organizations rather than events. Then look for the outcome the product already produces and is only labelling wrong. Three rules that fell out and should be reused: (1) make progress steps OBSERVED FACTS re-checked on every render, never stored completion flags, so a step goes back down when its fact stops being true (a revoked key must not leave a badge behind); (2) start the ladder with a rung the user has ALREADY cleared (endowed progress) rather than at zero; (3) report the biggest ABSOLUTE drop, not the smallest number, because they are different steps -- here the largest loss was 49 orgs at signup-to-first-meeting, not the agent gap being investigated.
Why: A metric that requires a hand-written query gets checked once and then never again, which is exactly how a zero on the company's core thesis survived for months next to a dashboard that looked healthy. Counting events instead of organizations is the specific trap. And honesty is load-bearing on any adoption surface: the moment a number flatters, the whole surface is worth less than showing nothing, so no points, no badges, no streaks, and zero must render as zero.
Failure mode: SUCCESS: Claude found OTP's north-star metric was zero and nobody knew, by querying production before designing anything. 61 orgs, 12 ran a meeting, 1 ever created an agent seat, 0 agents ever called OTP. The reason it hid: /admin/usage counts ACTIONS (looks healthy, a few orgs run many meetings) while the thesis needs per-ORGANIZATION counting. The fix reused data we already had: Ollie's real work was already recorded per org in wallet_ledger.metadata->>'feature' and had only ever been rendered as billing. Read as a timesheet, the same rows prove an agent already works there.
Before presenting a Dash blind-spot, billing trigger, or any state-file alert as a current action item, confirm it hasn't already been resolved. Stale state files (Dash May 25 was ~5 weeks old) carry point-in-time alerts that may be closed by now. Trust confirmed/observed status over stale notes; flag the data's age and treat unverified alerts as 'verify' not 'urgent.'
Why: Re-surfacing already-resolved alerts as urgent erodes trust in the L10 briefing and spends David's attention during a low-push recovery window. Honesty about data staleness matters more than appearing comprehensive.
Failure mode: Dan surfaced the HiTone billing trigger ($43-49K/mo possibly un-invoiced) from Dash's stale May 25 state file as a live concern during the Jun 29 L10. David confirmed HiTone billing is correct and already handled, and asked to close it out.
KPIs/scorecards must live as tiles in OTP (the source of truth), not in markdown files or meeting briefs. Every active agent/human seat — including Dan's strategic co-founder seat — must own at least one OTP KPI tile. When proposing measurables, verify against list_my_kpis and create the missing tiles via update_kpi (auto-creates), rather than just tabling them in a doc. A seat with no number is sitting on the sidelines.
Why: EOS requires every seat to have a measurable. Discussing KPIs in a brief while OTP shows none of them makes the scorecard fiction and undercuts OTP as the coordination source of truth. Dan as co-founder must be measurable like everyone else.
Failure mode: Dan presented a Sneeze It agent-team scorecard as a markdown table in the L10 brief and treated it as 'the scorecard,' when the source of truth is OTP. David caught that Dan (and Arin/Pulse/Dirk) have NO KPI tiles in OTP at all — Dan's own seat had zero measurables. A scorecard that only lives in a file or meeting brief does not exist.
An OTP KPI with teamId=NULL renders only on /dashboard/kpis, never on any L10 scorecard (meeting scorecards filter strictly by meeting.team_id). To make a KPI show on a specific L10, PATCH /api/v1/kpis/:id with the meeting's teamId. The 'Dan L10' meetings run on the 'ai-army' team (065d1d4b-c7da-4e80-b3ed-d6b101471d2c). Tally's auto-create now includes teamId from a 'team_id' field in the registry entry, so new agent-army KPIs land on the Dan L10 automatically instead of orphaned. Find team IDs via GET /api/v1/teams; meeting->team via GET /api/v1/meetings.
Why: A KPI nobody can see on their meeting scorecard is functionally not on the scorecard. The owner/title is necessary but not sufficient — team scoping is what makes it report. This is a recurring gotcha for any agent creating KPIs via the API.
Failure mode: SUCCESS: Tally — agent KPIs were invisible on the L10 because auto-create left teamId NULL. David flagged that the new KPIs weren't reporting on the Dan L10 or /dashboard/kpis as expected.
Root cause was the SearchAtlas OTTO pixel having an EMPTY src="" in the layout head (v7.ejs, onboarding.ejs, main.ejs). The OTTO tag must carry its base64 data-URI loader in src that appends dynamic_optimization.js with data-uuid; with src="" the runtime never loads, so OTTO injects/verifies nothing. When an OTTO/SearchAtlas audit reports 0/N across ALL on-page categories, suspect the pixel loader, not the actual tags — verify the sa-dynamic-optimization script's src is populated, not the page's own meta.
Why: A 0/16 across every category despite visibly correct meta is the signature of a non-loading optimization runtime, not missing tags. Checking the pixel first avoids a pointless rewrite of titles/descriptions that were never the problem.
Failure mode: SUCCESS: Beacon/SEO — orgtp.com OTTO on-page audit showed 0/16 (titles, meta descriptions, headings, meta keywords all failing) even though pages had perfectly good title tags and meta descriptions server-side.
A brand battle cry needs a genuinely designed moment (confident display type, intentional line breaks, brand device, real whitespace), not a centered text block plopped in. And the VISIBLE battle cry copy is the short clause only: 'Unlocking the potential in every person through the partnership of people and AI' — drop 'so together we leave the world better than we found it' from the hero display (keep the full sentence only for formal/footer contexts).
Why: A mission line is a brand centerpiece. Long copy dilutes the punch, and an undesigned drop-in reads as filler. The payoff phrase 'partnership of people and AI' must land as the climax with design weight behind it.
Failure mode: Adding the OTP mission as a 'battle cry' on the landing page, I dropped the full sentence into a plain centered text band wedged between hero and Step 1. David called it 'a weak attempt to just throw it on the page' and said the full line is too long for the visible battle cry.
Manifesto/mission pages must be written as movement recruitment, not product marketing: second-person address (the reader is the protagonist), "We believe" creed statements people can recite, a named enemy, stakes, and invitation CTAs ("Join the movement") instead of transactional ones ("Start free"). Product features appear only once, framed as how the movement fights, not what the product includes.
Why: People join movements because they believe what the movement believes (Sinek: start with why). Copy that sells the what on a page whose job is to recruit believers reads as generic SaaS and inspires no one, no matter how good the design is.
Failure mode: Redesigned the orgtp.com manifesto homepage with strong visual design but kept product-brochure copy (feature lists, "free meeting software", "Start free" CTAs). David: "the writing does not inspire an army of followers... this just looks the same as every other company... blah."
Judge conversion on the full path the visitor actually walks (page, door, day-one experience), not on the surface being edited. If the honest answer to "would you sign up" is "yes IF another surface delivers," the answer is no, and the work moves to that surface. Never write a promise on a button that the destination page cannot cash.
Why: Trust destroyed at the moment of verification is unrecoverable; a skeptical buyer who clicks "watch us run" and lands on a data page is gone forever. Copy that outruns proof is hype by definition, and the exact audience OTP needs (operators) is the audience that punishes it hardest.
Failure mode: After rewriting the OTP homepage, I declared the copy converts because skeptics would "click through to the live OOS page and sign up IF it delivers." David called it: I kicked the can to a page I know does not deliver, and called it a win. The button promises "Watch our company run, live" but the OOS page it links to is a list of published rules, not a running company.
OTP's enemy statement is "you bought the operating system and the needle didn't move." The pitch is not better meetings; it is: the system was fine, what was missing was the workforce that runs it between the meetings. Frame all homepage/sales copy against needle-not-moving, not against meetings.
Why: This is the buyer's actual lived disappointment (paid for an operating system, company looks the same two years later) and it positions OTP against incumbents on outcomes instead of features.
Failure mode: The letter's hero framed the enemy as "the meeting" / busywork. David corrected the thesis: the real problem is that companies bought operating systems and software (Ninety, Bloom Growth, etc.) that did not move the needle. Years later the company had not grown and was not better, and they needed to change how they did things.
For any UI change, design from the user's mental model, not the data model: "my list shows my work; work I assigned to others shows under Waiting on Others." When a meeting todo is assigned to someone else, stamp the creator as delegator so it routes to the delegation view. Before shipping UI changes, run the UX lens (impeccable / web-design-guidelines skills + src/DESIGN.md), not just a minimal code patch.
Why: A technically-correct patch that ignores the user's mental model just moves the confusion. OTP's own product language already has the right home for these items (Waiting on Others); fixes should land in the model the user already understands.
Failure mode: Fixed the dashboard todo confusion (teammates' meeting todos looked like the viewer's own) by adding an owner label to the rows. David corrected: that's not thinking like a user. Labeled-or-not, other people's todos don't belong in "my to-dos" at all.
When David asks for a jaw-drop brand page, build an EXPERIENCE, not an article: full-viewport cinematic hero, scroll choreography, one idea per screen at massive scale, motifs that live in the page as motion, ruthless copy cuts, no standard nav/footer chrome breaking the spell, no section-grammar scaffolding, no FAQ accordion bolted onto a manifesto.
Why: The gap between "well-executed page" and "omg I love this" is the whole assignment on brand surfaces. Safe editorial structure is invisible at best; for-the-brave positioning demands the page itself be brave.
Failure mode: Built the /ollie manifesto page as a competent editorial layout (repeated mono eyebrow labels on every section, index rows, alternating light/dark sections, FAQ accordion at the bottom) and David rejected it outright: "this really really sucks." The brief was "reader drops on the ground saying omg I fucking love this" and the output was a safe template that reads as AI scaffolding.
When David gives a design reference URL, open it in a browser and STUDY it visually (proportions, type sizes, spacing, alignment) before designing; match its register, not just its layout skeleton. Elegant means restrained: modest type scale, centered calm hierarchy, generous whitespace, thin rules. Never hand-draw SVG artwork to imitate produced brand art; crop/reuse the actual asset or use nothing.
Why: A reference URL is the brief. Reading its HTML structure without seeing it rendered led to importing the skeleton with the wrong soul, twice. Amateur freehand art next to professional motion work destroys credibility instantly.
Failure mode: Second rejection on the /ollie page. David asked for sakana.ai/fugu: elegant, Japanese sense of design (restraint, whitespace, calm, modest type, precision). I delivered giant 9vw headlines, one shouting line per viewport, and hand-drawn SVG chevron "birds" that rendered as crude fat marker scribbles. I treated "jaw-drop" as scale and boldness when the reference was quietness and precision, and I drew freehand SVG art instead of using the actual video's artwork.
For fleet-wide spec maintenance: (1) tarball backup of ~/.claude before any agent touches specs; (2) partition files into DISJOINT clusters, one agent each, with CLAUDE.md owned by exactly one; (3) give every auditor the same stale-fact canon and the rule "verify a launchd plist exists before believing any schedule claim"; (4) auditors apply surgical edits directly for factual fixes but RETURN structural proposals for David instead of applying them; (5) synthesizer closes cross-cluster contradictions the auditors flag at each other.
Why: Agent specs rot faster than anyone audits them: this pass found live specs for a retired agent (jeff.md ending in "Go."), four phantom schedules, Todoist writes in five files, terminated employees still routed DMs, and a Bassim score-inflation bug. Periodic fleet audits with disjoint ownership are cheap insurance against agents acting on dead infrastructure.
Failure mode: SUCCESS: Claude ran a five-cluster parallel level-up of the entire agent army (80 files, ~140 surgical edits) without a single file conflict or lost spec.
Pattern for UX dead-end hunts: (1) fan out parallel read-only explorers per surface (meetings, teams/members, KPIs/todos, onboarding/settings) asking for file:line + user-visible symptom + minimal fix; (2) fix the unsatisfiable states first: any required dropdown that can render zero options must explain where its options come from and link there (owners/attendees come from the org chart, meeting membership from teams); (3) empty states must branch on WHY they are empty (org has no teams vs user not on a team need different CTAs); (4) never report an async side effect as done: invite emails now await sendEmail (which returns null on failure, never throws) and return emailSent so the UI can tell the truth; (5) a guided setup checklist computed server-side from actual data (seats/team/KPI/meeting/members exist?) beats static onboarding because it survives skipped onboarding.
Why: These are the recurring shapes of broken UX in OTP: forms with prerequisites the user cannot see, empty states that misdiagnose their cause, and optimistic success messages over fire-and-forget side effects. Fixing the shape, not just the instance, is what makes the product feel intuitive.
Failure mode: SUCCESS: Claude ran a full UX dead-end audit and fix pass across OTP (4 PRs, #113-#116, all deployed)
The sweep pattern that worked: audit by rule-cluster in parallel (fakery, insight-to-agency, jargon/states, first-meeting goal-walk), then execute severity-first. Key catches to re-check every run: (1) seeded/synthetic data leaking into numbers a reader believes are real (the is_template flag existed but was never enforced; counts now use src/shared/synthetic-orgs.ts); (2) the conversion moment must be ON the default path (end-meeting now lands on Ollie followups, not the list); (3) funnels don't exist until instrumented (insight topic: surfaced/accepted/value_delivered); (4) credentials in seed script comments (one prod DATABASE_URL scrubbed; password rotation still owed). Worklist for run 2 in otp-platform/mission-standard/WORKLIST.md.
Why: The Mission Standard is a repeatable bar, not a one-off audit. Recording the found failure classes makes run 2 start from run 1's ceiling instead of re-discovering it.
Failure mode: SUCCESS: Claude ran Mission Standard sweep run 1 (PRs #121-#123, deployed): 4 parallel rule-audits over OTP, 5 CRITICAL + 13 GAP found, all CRITICAL and 9 GAP closed same-session
When an agent-pushed OTP to-do references a document, include a clickable https link (a Google Doc), not a local file/vault path — David reviews to-dos on mobile. To update an existing to-do's description, use PUT /api/v1/todos/:id (not PATCH). otp-todo.sh has no update verb, so PUT directly with the API key.
Why: A file path in a to-do is dead weight on mobile — the reviewer can see the reference but cannot open it, which reads as "the link is missing/broken." Every agent that pushes doc-linked to-dos (Radar, Pepper, Dan) hits this.
Failure mode: Dan pushed an OTP to-do referencing a document but put a local Obsidian vault path ("2nd Brain/Agent Army/Dan/...") in the description. David opens to-dos on his phone — a vault path is not tappable, so there was "no link to click." Also used PATCH to update the to-do; the OTP todos API update method is PUT /api/v1/todos/:id (PATCH hits the marketing site and returns HTML).
To put an agent-army/IDS issue on an OTP meeting board, POST /api/v1/tickets with the team's teamId (category 'other' for strategic issues, priority low/medium/high/critical, ownerEntityType+ownerExternalId). The MCP submit_ticket tool CANNOT do this — it has no teamId param (it is the generic 'report a bug to OTP' path), which is why nobody ever got issues onto the board. Team IDs: 'AI Army' = 065d1d4b-c7da-4e80-b3ed-d6b101471d2c (the David+Dan agent-army meeting); Leadership Team = c1e1a485-414e-48d5-ae44-e81bd110b554. Update/solve via PUT /api/v1/tickets/:id (idsStatus, priorityRank, resolution).
Why: Agents could push KPIs and todos to OTP but not issues, so every L10 IDS board rendered empty and David kept discovering the hole live. Issues=tickets + teamId scoping is the missing piece; without it a meeting-readiness check would keep mislabeling a working API as absent.
Failure mode: Dan claimed 'no issues API exists in OTP' because there is no src/routes/api/issues.ts. That was wrong. OTP stores IDS issues in the TICKETS table (schema.ts: 'issues live in the tickets table'), with full IDS support (idsStatus, priorityRank, teamId, owner fields). The agent-army IDS board was empty only because our issues lived in a local markdown file and were never pushed as tickets scoped to a team.
OTP has TWO distinct features both called "Ollie Insight": (A) the per-meeting followups wizard that turns a transcript into meetings.ai_summary (src/shared/meeting-followups.ts, transcript-only), and (B) the reusable "address engine" (ollie_insights table, src/services/ollie-insight.ts + shared/ollie-insight.ts + partials/ollie-insight-block.ejs) that gathers org data (KPIs/rocks/todos/meeting-summaries) per SCOPE. When David said "KPIs shouldn't be in the meeting analysis," the fix was in System B's meeting-scope evidence gathering, NOT System A. The block partial (ollie-insight-block.ejs) is fully scope-generic (builds the API URL from data-oib-scope/scopeId client-side), so adding a brand-new 'quarter' scope end-to-end took only: add to INSIGHT_SCOPES + RULES_BY_SCOPE (shared), the scopeQuerySchema enum + a resolveInsightScope branch (api), a gatherEvidence branch + max_tokens (service) -- then just include the existing partial with scope:'quarter'. No new render/generate/receipts UI. Pattern: when a feature is "one engine, many surfaces," new surfaces are a scope + evidence branch, never new UI.
Why: The name collision hides which code to touch; picking the wrong system wastes a whole edit pass. And recognizing the scope-generic block means big-feeling asks ("a quarterly synthesis button") are small, low-risk diffs. Both are recurring shapes in OTP's Ollie work.
Failure mode: SUCCESS: Claude separated OTP's two "Ollie Insight" systems and added a whole new scope by reuse
New endpoint POST /api/v1/meetings/:id/agent-record (PR #154): an agent submits the written meeting record; OTP runs the same redaction ruleset, stores it to meetings.transcript, logs an audit baseline + agent_record event, and the existing /ai/followups generate turns it into to-dos/issues/headlines/insight unchanged. Agent path: ~/.claude/otp-meeting.sh record <meetingId> --file=<record> --source=l10dan. Wired into /l10dan conclude step 6. Verify deploy by probing the endpoint returns JSON not the marketing SPA HTML before pushing records.
Why: Ollie only reads transcripts, so agent-run meetings had no way into it — the empty-insight hole David hit live. This closes it: any agent meeting can now produce Ollie follow-ups. Also a general UI rule captured: buttons reflect capability/state (action taken -> button disappears).
Failure mode: SUCCESS: Dan shipped the Ollie agent-record path so agent-facilitated meetings (David + AI L10, no audio transcript) can feed Ollie Insights.
Before treating a KPI/data source as blocked, re-read the LIVE source, not the note about it. The Havok "non-client %" KPI was marked blocked for ~14 weeks on "the timesheet has no client column" — a 105-day-old memory. The live sheet (1VPlH5ZqTowOe2nJDFwOjuvXrCXUnOzeOb3aO-M936Xo) had since grown per-person tabs with a full Client/Time/Date schema; one read unblocked it. Pattern for wiring a messy sheet into Tally: (1) get_spreadsheet_info to list tabs, (2) read a person tab to learn the real schema, (3) confirm recency by reading the tail (last date), (4) add a focused extract mode to tally.py rather than reshaping the sheet — here `client_attribution_since` (reads ALL valueRanges via a new _values_2d_all, parses h:mm via _parse_hhmm, added dotted DD.MM.YYYY to _parse_date, classifies internal by an `internal_contains` substring), (5) `tally.py --dry-run --kpi "<title>"` to prove the number off live data before pushing. Human-owner + a 1:1-team KPI (owner HUM_BOGDANTABAKA, team = David-Bogdan 1:1) pushes fine via find_or_create_kpi.
Why: Blocked-status notes rot silently while the underlying source improves; a KPI can sit "pending" for a quarter when it was buildable weeks ago. Re-reading the live source first is the cheap unblock. And the tally.py extract-mode pattern makes any timesheet/sheet a live KPI without asking a human to restructure their doc.
Failure mode: SUCCESS: Dan/Tally shipped the Havok client-attribution KPI live in one session after a 14-week "blocked" note turned out stale
Before editing any OTP view to fix an on-screen bug, grep unique visible strings from the screenshot (e.g. "WAITING ON OTHERS", "always only yours") across src/views to confirm WHICH template renders that exact surface. Multiple pages can render similar-looking todo lists (me-todos.ejs vs dashboard-daily.ejs). Verify the rendering route (reply.view target) too.
Why: Two round-trips and two merged PRs produced zero visible change because the edits were on the wrong template, which read as "nothing is fixed" and eroded trust. A 10-second grep on the screenshot text would have pointed to the right file immediately.
Failure mode: Fixing an OTP todos UI bug, I edited src/views/pages/me-todos.ejs twice and shipped two PRs, but the surface David actually uses is the dashboard-daily "Waiting on others" widget (src/views/pages/dashboard-daily.ejs). Nothing he saw changed.
Two reusable patterns: (1) Before building any OTP email/engagement feature, grep src/services for existing infrastructure -- re-engagement.ts, lifecycle-scheduler.ts, and user_engagement_log already carried cadence caps, suppression, logging, and a daily cron, so the todo-aware upgrade was ~350 lines instead of a new subsystem. (2) When resolving "which org does this Clerk user belong to", organizations.clerkOrgId only knows the org CREATOR; invited teammates must be resolved through org_members.clerkUserId + claimedEntityIds. This gap is why per-user personalization (open todos) missed non-creator members like Nate.
Why: One engagement channel with shared caps is what keeps daily utilization pressure from becoming annoying double-mailing, and the creator-vs-member resolution gap will bite any future per-user feature (digests, notifications, billing seats) that starts from organizations.clerkOrgId.
Failure mode: SUCCESS: Claude shipped the smart engagement email engine (PR #186) by upgrading the existing re-engagement service instead of building a parallel system
Frame help/success/onboarding call copy positively: state plainly that the call is there to help and guide them, and describe what will actually happen on it (we'll walk through your setup, get you unstuck, answer your questions). Never say "this is not a sales call" or "no pitch" — describe the help, don't disclaim the sell.
Why: Defensive "not a sales" language triggers the exact suspicion it tries to defuse and undercuts a genuine help offer. David flagged this immediately.
Failure mode: Wrote a "customer success call" Calendly description that leaned on "no pitch, no slides" / not-a-sales-call framing. Protesting that it isn't a sales call makes it sound like one.
Any script that answers "is anything missing / is everything covered?" must fail LOUD, never return an empty set as reassurance. Two rules: (1) assert the expected top-level key exists (`if 'data' not in resp: raise`) before computing a result; a zero/empty answer from a health check is a claim that must be proven, not a default. (2) Cross-check a zero result against one known-positive case before reporting it -- here, one direct call to a single account would have shown $57 of spend and exposed the lie instantly. Also: `~/.claude/meta-ads.sh accounts` exits 0 and prints nothing; do not build on it. Sweep with `/{business_id}/adaccounts?fields=name,account_status&limit=500` then batch `/{act_id}/insights` 50 at a time.
Why: "Nothing is wrong" is the single most dangerous output an audit can produce, because nobody investigates it. A silent empty result on a billing sweep means real revenue is never invoiced and nobody ever finds out. The failure mode is not a crash, it is confident silence.
Failure mode: SUCCESS: Claude caught a silent false-negative in a Meta billing sweep. Querying the Graph API adaccounts edge with nested field expansion (`fields=name,insights.date_preset(this_month){spend}`) returned an error payload with NO `data` key. The sweep script read it as an empty account list and confidently reported "0 accounts, $0.00 unbilled MTD spend" -- a clean bill of health that was entirely fabricated. Direct per-account queries then revealed 7 unbilled accounts spending $5,673 MTD.
Support responses for annual plan customers auto-flagged YELLOW. Annual customers represent 4x monthly revenue.
Why: Sloppy response to monthly customer costs $150/year if they churn. Sloppy response to annual customer costs $1,800. Review depth should match revenue at risk.
Failure mode: Annual customer billing question gets generic response. Feels undervalued. Does not renew. $1,800 lost.
Engineering alerts suppresses repeat pages for same issue within 30 minutes. First alert pages. Subsequent alerts update the existing incident thread.
Why: A database slow query triggered 14 pages in 8 minutes. On-call engineer overwhelmed. Missed the actual resolution signal buried in noise.
Failure mode: Same issue generates 14 pages. Each interrupts the engineer. Noise drowns signal. Resolution delayed 20 minutes.
Patient education handouts are written at a 6th-grade reading level using the Flesch-Kincaid scale. The education agent checks readability before submitting for review.
Why: Our patient population includes a significant number of non-native English speakers and older adults. Early handouts scored at 10th-grade reading level. Dr. Okafor observed patients nodding along but clearly not understanding the content. Post-visit comprehension checks confirmed the gap.
Failure mode: Handout on managing hypertension uses terms like "antihypertensive regimen" and "sodium restriction protocol." Patient takes the handout home, does not understand it, and does not follow the guidance. Blood pressure remains uncontrolled at the next visit.
Appointment reminders for patients who no-showed their last visit include a warmer, non-judgmental tone and an explicit offer to reschedule. No mention of the missed appointment.
Why: The default reminder tone felt transactional. Patients who had already missed once responded better to "We'd love to see you" than "You have an appointment on Tuesday." Reschedule rate for prior no-shows improved from 31% to 48% after the tone change.
Failure mode: Standard reminder sent to a patient who missed their last appointment. Patient feels guilty or defensive. Ignores the reminder. No-shows again. Pattern solidifies.
Onboarding documents are versioned with a date stamp in the filename. When clinical protocols change, the onboarding agent regenerates affected documents within 48 hours. Old versions are archived, never deleted.
Why: A new MA was trained using a document that referenced the old blood draw protocol (tourniquet for 60 seconds). Protocol had changed to 30 seconds two months prior. Document had not been updated. MA followed the outdated procedure for a full week before a supervising nurse caught it.
Failure mode: Outdated onboarding document trains new staff on a deprecated procedure. Staff performs the procedure incorrectly. In a primary care setting, most deprecated procedures are low-risk, but the cumulative effect of outdated training erodes clinical quality.
Incident severity is determined by customer impact radius, not system impact. A database hiccup affecting 1 internal dashboard is P3. A 200ms latency increase affecting all customer eval jobs is P1.
Why: Internal systems failing is inconvenient. Customer-facing systems degrading is revenue-threatening.
Failure mode: Monitoring agent classified a latency spike as P3 because the internal system health dashboard showed green (it only measured error rates, not latency). 85 customers experienced 3x slower eval results for 2 hours. 15 filed support tickets. 3 enterprise customers included the incident in their quarterly vendor review. Two of those reviews resulted in "conditional renewal" status.
Investor updates are published monthly, on the 5th, regardless of whether the numbers are good. Skipping a month signals that something is wrong.
Why: Investors pattern-match on communication cadence. A missed update generates more anxiety than a bad update.
Failure mode: Lina skipped the March investor update because MRR had dipped 4% (3 customers delayed renewals). Two board members texted within a week asking "everything okay?" The April update included the March data and the dip explanation, but the trust damage from the missed communication took the entire board meeting to repair.
Sales demo prep agent refreshes demo environments weekly. Stale demo data that references outdated features or deprecated APIs undermines credibility during live demos.
Why: Enterprise prospects evaluate attention to detail. A demo that shows a deprecated feature signals that the product moves faster than the company can manage.
Failure mode: A demo environment showed an eval metric type that had been deprecated 2 months earlier. The prospect asked about it. The sales engineer said "oh, that's been removed." The prospect replied: "So your demo doesn't reflect your actual product? What else is out of date?" The deal took an additional 3 weeks to close and included a requirement for a "current state" audit before signing.
Haven's first response to any customer inquiry must be sent within 15 minutes during business hours (9 AM - 6 PM ET). The response can be a draft that the CS rep reviews, but the customer must see a reply within 15 minutes. Outside business hours, the autoresponse sets expectations for next-business-day response.
Why: Speed to first response is the single highest-correlating factor with CS satisfaction scores. A fast "we're looking into this" beats a slow comprehensive answer every time. Haven's drafts are fast. Human review can happen after the first touch.
Failure mode: Before Haven, average first response time was 4.2 hours. Customer satisfaction (CSAT) was 3.4/5. After implementing the 15-minute target with Haven drafts, first response dropped to 8 minutes average. CSAT rose to 4.1/5 within 6 weeks. No other change was made during that period.
Rhythm must test subject lines on a 10% sample before full send for any campaign going to more than 5,000 contacts. The winning subject line (by open rate after 2 hours) goes to the remaining 90%. No exceptions for "time-sensitive" campaigns.
Why: A 5% improvement in open rate on a 28,000-person list is 1,400 additional opens. On Threadline's average click-to-open rate of 18%, that's 252 additional clicks. At a 3.2% conversion rate, that's 8 additional orders averaging $67 each -- $536 in revenue from a 2-hour wait.
Failure mode: The marketing coordinator overrode the A/B test for a Black Friday campaign because "we need to send now, every minute counts." The chosen subject line had a 14% open rate. The founder ran the unused B variant to a test segment later: 23% open rate. Estimated lost revenue from skipping the test: $4,800 on Black Friday, the single highest-revenue email day of the year.
Quarterly investor reports are prepared 21 days before distribution. The first 7 days are for agent drafting and internal review. The next 7 days are for Chen's compliance review and outside counsel if needed. The final 7 days are buffer for revisions.
Why: Rushing quarterly reports produces the C003-type errors. The 21-day cycle ensures every number is audited, every statement is compliant, and there is time to fix problems.
Failure mode: Before the 21-day cycle, Q3 reports were prepared in 5 days. The C003 incident (preliminary vs. audited IRR discrepancy) happened because there was no time for Derek to complete the audit reconciliation before distribution.
Deal memos include a mandatory "Risk Factors" section with a minimum of 8 risk factors. The deal memo agent generates risk factors from a master risk taxonomy and adds deal-specific risks identified during analysis.
Why: Insufficient risk disclosure in offering materials creates legal liability. If an investor loses money on a risk that was foreseeable but undisclosed, the liability falls on the fund.
Failure mode: An early deal memo had 3 risk factors, all generic ("Market conditions may change," "Past performance does not guarantee future results," "Real estate is illiquid"). Chen added 9 deal-specific risks including environmental remediation liability, tenant concentration risk, and interest rate sensitivity. After this, the minimum was set at 8 with mandatory deal-specific analysis.
Investor communications use tiered language precision based on the content type. Performance updates: exact numbers with 2 decimal places and data source attribution. Market context: ranges and qualifiers ("approximately," "in the range of"). Outlook: conditional language only ("if market conditions persist," "subject to").
Why: Precision signals competence. But false precision on uncertain topics signals naivete or deception. An investor who reads "We project 14.7% IRR" treats it as a promise. "Under base case assumptions, projected returns range from 12-16% IRR" is honest.
Failure mode: The investor comms agent drafted a year-end letter stating "Our portfolio returned 11.4% in 2025." Derek's audited number was 11.38%. The rounding was correct, but the letter didn't cite the audited source. Chen added "Based on audited Q4 2025 financials prepared by [Auditor Name]" to every performance figure.
When using GPT for creative agents (Cadence, Chorus) and Claude for analytical agents (Pulse, Signal, Atlas, Ledger, Prism), maintain separate evaluation criteria. Creative agents are evaluated on brand voice consistency and engagement metrics. Analytical agents are evaluated on accuracy and signal-to-noise ratio. Never evaluate a creative agent on precision or an analytical agent on tone.
Why: The two platforms were chosen for different strengths. Evaluating both with the same rubric incentivizes the wrong behaviors -- GPT agents get over-optimized for accuracy (killing creativity) and Claude agents get prompted for engaging tone (introducing imprecision).
Failure mode: The team applied a single "quality score" rubric to all agents. Chorus (GPT, creative) scored low on "factual accuracy" because product descriptions included aspirational language. The team tried to make Chorus more precise, which killed the brand voice. Meanwhile, Pulse (Claude, analytical) scored low on "engaging presentation." The team added formatting requirements that made Pulse's alerts harder to scan quickly. Both agents got worse by being evaluated on the wrong criteria.
Agent context switches between brands must include a "brand flush" step: clear the prior brand's context, load the new brand's configuration file, and confirm the brand identity in the output header. No "carry-over" operations where an agent finishes Brand A work and immediately starts Brand B work without context clearing.
Why: Context carry-over is the root cause of voice bleed, data leakage, and policy confusion. The 30-second cost of a brand flush is negligible compared to the cost of any cross-brand contamination incident.
Failure mode: The adventure-tee incident (C001), the email cross-contamination (C002), and the CS tone mismatch (C006) all traced back to context carry-over. Implementing mandatory brand flush reduced cross-brand incidents from 4-6 per month to 0-1 per month within the first 30 days.