All Sections
Coordination Intelligence

operational heuristics

201 claims from 32 organizations

Rules of thumb learned from practice. These claims may not be provable in all cases, but they have been observed to work reliably. Heuristics evolve into rules as evidence accumulates.

Acme Digital Agency Founding gold
C010 MEDIUM OBSERVED ONCE 3x efficiency

Unresolvable errors: log and stop. No automatic retry.

Why: Retries on unresolvable errors create noise.

Failure mode: Agent retries 100 times, consumes rate limits.

C008 HIGH MEASURED RESULT 10x efficiency

GPT creative output undergoes a mandatory 2-hour hold before entering the review queue. No same-session generation and approval.

Why: Pattern: account managers reviewing GPT output immediately after generation approved it at a 94% rate. When we added a 2-hour hold (the AM reviews the batch later in the day or the next morning), the approval rate dropped to 71%. The 23% gap was copy that "sounded good in the moment" but had issues visible with fresh eyes -- subtle tone mismatches, claims that were technically true but misleading, and formatting that didn't match the client's brand voice.

Failure mode: Same-session review creates familiarity bias. The reviewer just saw the brief, the context is loaded, and the output feels like a natural continuation. Distance improves judgment.

C009 HIGH MEASURED RESULT 10x efficiency

Track and report the cross-model error rate separately from single-model error rates. Any task that involves both Claude and GPT is measured as a distinct category.

Why: Our overall agent error rate was 3.2%. When we segmented, single-model tasks (Claude-only or GPT-only) had a 1.8% error rate. Cross-model tasks had an 8.7% error rate -- nearly 5x. The errors were concentrated at handoff points: incomplete context transfer, schema mismatches, and misinterpretation of structured fields. Without segmenting, the 3.2% blended rate masked a systemic handoff problem.

Failure mode: Blended error rates hide that cross-model handoffs are the primary failure point. Resources are allocated to improving individual agents when the real problem is the integration layer.

C016 LOW OBSERVED ONCE 1.5x efficiency

Every agent must include its model name and version in the metadata of its shared state file output.

Why: When debugging an analysis discrepancy, we couldn't tell which model version produced a particular output. The ad monitor had been running on Claude 3 Opus while the pacing agent had been upgraded to Claude 3.5 Sonnet. Their outputs used different rounding conventions, making numbers mismatch by $1-3 per metric. Model version in metadata would have identified the discrepancy source in minutes instead of the 2 hours it took.

Failure mode: Without model version tracking, debugging cross-agent discrepancies requires testing each agent individually. Root cause identification is slow when version differences aren't visible.

C010 HIGH OBSERVED REPEATEDLY 7x efficiency

Ad performance is evaluated at the location level with a minimum 14-day window. No optimization decisions are made on less than 14 days of data per location.

Why: Small-market locations (Tampa, Phoenix suburbs) have low daily lead volume. Day-to-day variance is enormous. A "bad day" is meaningless; a bad 2-week trend is actionable.

Failure mode: The ads agent paused a Google Ads campaign for the Tampa location after 5 days of zero leads. The campaign had averaged 2.1 leads/day over the prior month. The 5-day drought was within normal variance. Restarting the campaign lost 3 days of learning and reset the algorithm.

C011 MEDIUM OBSERVED ONCE 3x efficiency

Lead quality scoring weights location-specific conversion history over network averages. A "hot" lead in Chicago (where average ticket is $149/mo) has different characteristics than a "hot" lead in Tampa (where average ticket is $89/mo).

Why: Network-wide lead scoring models produce false positives in lower-ticket markets and false negatives in higher-ticket markets.

Failure mode: The lead distribution agent prioritized Tampa leads using the Chicago-trained scoring model. It flagged members interested in premium personal training as "hot." Tampa doesn't offer premium PT. Staff called 30 "hot" leads pitching a service that didn't exist at their location.

C012 MEDIUM OBSERVED REPEATEDLY 4x efficiency

Seasonal patterns are tracked per-location, not network-wide. January surges vary by 40-60% across locations. Summer dips range from 10% to 35% depending on climate and demographics.

Why: Using network-average seasonality for budget planning over- or under-allocates at the extremes.

Failure mode: The ads agent applied a uniform 30% January budget increase across all locations. Chicago needed 50% (cold weather drives indoor fitness demand). Tampa needed only 15% (year-round outdoor fitness options). Tampa overspent by $1,200 in January. Chicago underspent and missed 40+ leads.

C009 MEDIUM OBSERVED REPEATEDLY 4x efficiency

When a client requests more than 3 rounds of revisions on a single deliverable, flag it for Marcus as a potential scope issue before logging revision 4.

Why: Unlimited revisions is the silent killer of creative agency margins. The flag forces a conversation about whether to charge for additional rounds.

Failure mode: One project went through 7 revision rounds without anyone noticing the pattern. The client was happy but the project margin was -12%. Marcus didn't realize until quarterly review.

C010 MEDIUM OBSERVED REPEATEDLY 4x efficiency

Shot list suggestions must include at least 2 reference links from the client's existing brand content or stated references. Never suggest shots without grounding in the client's visual language.

Why: Generic shot suggestions feel like templates. Grounded suggestions feel like someone studied the brand.

Failure mode: Agent suggested a "drone reveal shot" for a brand that exclusively uses intimate, handheld footage. Designer flagged it as "clearly from a machine that doesn't understand the brand." Technically correct suggestion, totally wrong for the client.

C011 MEDIUM OBSERVED ONCE 3x efficiency

Invoice reminders to clients use the same casual tone Marcus uses in his own emails. No formal language, no "please remit payment."

Why: Artifact's brand is approachable and creative. Formal invoice language feels like it's coming from a different company.

Failure mode: First automated invoice reminder used "Please find attached your outstanding invoice for services rendered." Client replied to Marcus: "Did you hire an accountant? Lol." Minor, but it broke the illusion of a small, personal shop.

Atticus Legal bronze
C010 MEDIUM MEASURED RESULT 6x efficiency

Document assembly for a standard estate plan (trust, will, POA, advance directive) takes the agent 12 minutes. Priya's review takes 45 minutes per client package. Beth's formatting and preparation for signing takes 30 minutes. Total pipeline: 87 minutes per client versus the pre-agent baseline of 3.5 hours.

Why: Knowing the pipeline timing lets Priya schedule accurately. Before agents, she routinely underestimated document prep time and fell behind. Now she schedules 90-minute blocks per client and consistently hits the mark.

Failure mode: Without accurate pipeline timing, Priya overschedules. Four signing appointments in one day when she can only prepare for three. Last client's documents rushed. Error rate increases when Priya is behind schedule.

C011 MEDIUM MEASURED RESULT 6x efficiency

The scheduling agent sends three follow-up touches for annual trust reviews: 60 days before anniversary, 30 days, and 7 days. After the third touch with no response, it flags the client as DORMANT and stops. No more than 3 touches per year.

Why: Over-following-up annoys clients and feels desperate. Under-following-up loses annual review revenue ($750-$1,200 per review). Three touches at 60/30/7 days produced a 62% review booking rate, up from 38% when Beth was sending manual reminders on an ad hoc schedule.

Failure mode: Without a cap, the agent sends monthly reminders. Client perceives the firm as aggressive. Leaves a negative review mentioning "constant harassment." Priya loses a client and a referral source. In a solo practice, every client lost is felt in the revenue.

C012 LOW OBSERVED ONCE 1.5x efficiency

The assembly agent generates documents in the firm's standard formatting: 12pt Times New Roman, 1-inch margins, numbered paragraphs, firm letterhead. It never uses alternate fonts, creative layouts, or formatting that deviates from the template.

Why: Estate planning documents are read by courts, trustees, financial institutions, and opposing counsel. Non-standard formatting signals carelessness. A bank once questioned a trust because the formatting looked "different from what we usually see" and requested a letter of opinion confirming its validity. That letter cost Priya 2 hours.

Failure mode: Agent assembles a trust with a different font because the template's font metadata was corrupted. Bank receiving the trust as part of an account titling process flags it as potentially invalid. Client calls Priya. Priya spends 2 hours writing a letter of opinion. $600 in non-billable time because of a font.

C011 HIGH OBSERVED REPEATEDLY 7x efficiency

Progress reports are generated biweekly, not weekly. Weekly reports create parent anxiety without providing meaningful new information.

Why: Tutoring progress is nonlinear. A bad week followed by a good week looks like a crisis and a recovery in weekly reports. Biweekly smooths the noise.

Failure mode: Weekly reports caused 3 parents to request "emergency conferences" in a single month because their child had one below-average session. Keisha spent 6 hours in unnecessary meetings. Moving to biweekly reduced parent escalations by 80%.

C012 MEDIUM OBSERVED REPEATEDLY 4x efficiency

Tutor session notes must be submitted within 24 hours of the session. The scheduling agent flags missing notes at the 24-hour mark.

Why: Tutors forget details after 24 hours. Late notes are less accurate, which contaminates progress reports.

Failure mode: One tutor submitted 3 weeks of notes in a single batch. The notes were vague ("worked on math") and unusable for progress reports. Keisha had to contact 8 families to apologize for the delayed report. She now pays tutors a $5 bonus for same-day notes.

C013 HIGH OBSERVED REPEATEDLY 7x efficiency

When a student misses 2 consecutive sessions without parent communication, the parent communication agent drafts a check-in message for Keisha.

Why: Missed sessions without communication often signal a family considering leaving. Early outreach retains 60% of at-risk families.

Failure mode: Before this rule, a family missed 4 sessions over a month. Keisha assumed they were on vacation. They had actually switched to a competitor. She found out when the mother mentioned it casually at a school event. $4,800/year lost with zero warning.

Candor Labs bronze
C007 HIGH OBSERVED ONCE 5x efficiency

The release notes agent generates from merged PR titles and descriptions only. It does not read code diffs to infer what changed.

Why: The release notes agent read a diff that renamed an internal function and generated a changelog entry: "Breaking: API endpoint /auth/refresh renamed." The function rename was internal only -- the API contract hadn't changed. A user on the changelog RSS feed opened a GitHub issue asking about the "breaking change" and whether they needed to update their integration. The founder spent 30 minutes clarifying that nothing had changed for users.

Failure mode: Agents infer user-facing changes from internal code diffs. Internal refactors are misrepresented as breaking changes. Users react to non-existent breaking changes.

C008 HIGH MEASURED RESULT 10x efficiency

The support triage agent checks for duplicate issues before categorizing. Duplicates are linked, not re-triaged independently.

Why: The same bug was reported 4 times across GitHub issues and Slack over a weekend. The triage agent created 4 separate P1 entries. The founder's Monday morning check showed 4 P1 issues and he panicked, thinking 4 different critical bugs had emerged. It was 1 bug reported 4 ways. After implementing deduplication (matching by error message, stack trace similarity, and affected endpoint), false P1 volume dropped 35% over the next month.

Failure mode: Duplicate reports are triaged independently, inflating priority counts. The single reviewer overestimates severity based on volume. Panic replaces triage.

C014 MEDIUM MEASURED RESULT 6x efficiency

The code review agent includes a "confidence" indicator on each finding: HIGH (definite bug or security issue), MEDIUM (likely problem, needs human judgment), LOW (style preference, could go either way).

Why: Without confidence labels, the founder treated all code review findings equally. He either fixed everything (30 minutes on style nits) or skipped everything (missed real bugs). Confidence labels let him triage: fix all HIGHs immediately, review MEDIUMs during dedicated review time, batch LOWs for monthly style cleanup. Effective review time dropped from 25 minutes/day to 8 minutes/day with zero increase in bugs reaching production.

Failure mode: Uniform presentation of findings forces binary processing: everything or nothing. Confidence labels enable triage. Without them, the solo founder's limited review time is misallocated.

C011 MEDIUM INFERENCE 2x efficiency

Proposal drafts from Archer should be 60-70% complete, not 95%. Leave strategic positioning, pricing rationale, and the executive summary for the consultant to write.

Why: Over-polished agent drafts create a false sense of completion. Consultants rubber-stamp instead of thinking critically. The best proposals have the consultant's genuine strategic voice in the sections that matter most.

Failure mode: Archer produced a near-perfect 28-page proposal. The partner skimmed it, approved it, and sent it. The pricing section included a 15% discount that Archer inferred from a prior similar engagement but that was not appropriate for this client's scope. Cost us $31,500 in margin.

C012 MEDIUM OBSERVED ONCE 3x efficiency

Vault's knowledge base must be refreshed quarterly. Methodologies, case studies, and templates older than 18 months must be flagged for review and either updated or archived.

Why: Stale methodologies in proposals and deliverables make the firm look dated. Clients in fast-moving industries notice when frameworks reference pre-pandemic market conditions.

Failure mode: Archer pulled a "Digital Transformation Readiness Assessment" template from Vault that referenced 2023 technology benchmarks. The client's CTO pointed out the benchmarks were 3 years old during the proposal review call.

C013 HIGH INFERENCE 3x efficiency

For engagements involving direct competitors, assign different consultant teams and ensure no agent holds context for both engagements simultaneously. Rotate agent context between engagements, never run them in parallel.

Why: Even with firewalls, simultaneous context is the highest-risk vector for information leakage. Sequential processing with context clearing is safer than parallel processing with access controls.

Failure mode: This is the architectural response to the Haldane/Orion incident (C001). No second incident has occurred since implementing sequential processing.

C007 HIGH OBSERVED REPEATEDLY 7x efficiency

Lead nurture sequences pause automatically when a prospect books a trial class. Resume only if they no-show or don't convert within 7 days.

Why: Continuing to nurture someone who already booked feels tone-deaf and spammy.

Failure mode: A prospect booked a trial, received 3 more "Book your free class!" emails before attending. Replied "I already did, is anyone actually reading these?" Trial converted but trust was damaged from the start.

C008 MEDIUM OBSERVED REPEATEDLY 4x efficiency

Trainer performance metrics use a 4-week rolling average, not single-week snapshots. Seasonal patterns (New Year surge, summer dip) are normalized against the same period last year.

Why: Single-week data is noisy. A trainer with one bad week due to illness shouldn't be flagged. Seasonal patterns create false positives.

Failure mode: January metrics showed every trainer "improving" dramatically. It was just the New Year resolution surge. Jamie almost gave bonuses based on phantom performance gains.

C009 HIGH OBSERVED REPEATEDLY 7x efficiency

Location-specific context must be attached to every agent action. No agent operates in a "generic CoreFit" mode. Each location has different peak hours, demographics, and class preferences.

Why: The downtown location skews young professionals (25-35). The suburban location skews parents (35-50). Messaging that works for one alienates the other.

Failure mode: Lead nurture sent "Bring the kids to our Saturday Family Fitness!" to downtown prospects. Downtown has no kids' classes. 4 confused replies.

C015 MEDIUM OBSERVED REPEATEDLY 4x efficiency

Member retention risk scoring uses a weighted model: visit frequency (40%), class variety (20%), social engagement (15%), billing consistency (15%), tenure (10%). No single factor triggers an at-risk flag alone.

Why: Single-factor triggers produce too many false positives. A long-tenured member who drops visit frequency for 2 weeks might be on vacation, not churning.

Failure mode: Early model used visit frequency alone. Flagged 23 members as at-risk in one week. 19 were on a local school spring break vacation. Jamie wasted 4 hours reviewing false flags.

DevForge silver
C011 HIGH OBSERVED REPEATEDLY 7x efficiency

Issue triage prioritizes by: (1) enterprise customer reports, (2) security issues, (3) issues with reproduction steps, (4) feature requests with 5+ thumbs-up, (5) everything else.

Why: Enterprise customers pay. Security issues are existential. Reproducible issues get fixed faster. Community-validated features should ship. Everything else can wait.

Failure mode: Before prioritization, issues were triaged by recency. An enterprise customer's critical bug sat at position #14 in the queue behind 13 minor feature requests. The customer escalated via email after 5 days. Kai fixed it in 30 minutes but the delayed response nearly cost the $2,400/year contract.

C012 MEDIUM OBSERVED ONCE 3x efficiency

Release notes are published within 24 hours of a release. If Kai hasn't reviewed the draft within 12 hours, the agent sends a reminder to his private Slack channel.

Why: Enterprise customers monitor releases. A release without notes triggers "what changed?" emails that cost more time than writing the notes.

Failure mode: Kai shipped v2.6.0 on a Friday and forgot to publish release notes. By Monday, 4 enterprise customers had emailed asking what changed. One customer's security team flagged the update as "unreviewed" and blocked their team from upgrading. It took 2 weeks to get through their security review after the late notes were published.

C013 MEDIUM OBSERVED ONCE 3x efficiency

Discord monitoring tracks sentiment, not just questions. A shift from positive to negative sentiment in any channel triggers a summary to Kai's private Slack within 1 hour.

Why: Developer communities turn fast. A frustrating bug or a perceived lack of responsiveness can shift tone from supportive to hostile in a single day.

Failure mode: A breaking change in v2.3.0 caused issues for 15+ users over a weekend. Discord #help went from 2 messages/day to 30 messages/day, all negative. Kai was offline and didn't see it until Monday. By then, the narrative had solidified: "DevForge ships breaking changes without warning." A community member forked the project as a "stable alternative." The fork got 200 stars before Kai could respond.

C010 HIGH OBSERVED REPEATEDLY 7x efficiency

Support ticket priority is determined by financial impact. Tickets mentioning failed payments, incorrect balances, or unauthorized transactions are auto-escalated to P1 regardless of the user's tone or language.

Why: A polite user reporting a $500 balance discrepancy is more urgent than an angry user complaining about the UI. Financial impact trumps sentiment.

Failure mode: The triage agent initially used sentiment analysis for priority. An angry user complaining about a font change was prioritized over a calm user reporting that $1,200 appeared to be missing from their account. The calm user waited 18 hours for a response. The "missing" money was a categorization display bug, but the delay eroded trust.

C011 HIGH OBSERVED REPEATEDLY 7x efficiency

API health checks run every 30 seconds for Plaid and Stripe. Degraded performance (response time >2x baseline) triggers a P2 alert. Full outage (no response for 3 consecutive checks) triggers P1.

Why: Degraded performance often precedes full outages. Early warning gives the engineering team time to activate failover or notify users before the situation becomes critical.

Failure mode: Before the degradation detection, a Stripe partial outage caused payment processing to slow from 200ms to 8 seconds. Users experienced "spinning" payment screens. 14 users abandoned mid-payment. The monitoring agent only alerted when Stripe went fully down 40 minutes later.

C012 MEDIUM OBSERVED ONCE 3x efficiency

The categorization QA agent flags systematic drift when any category's miscategorization rate exceeds 5% over a 7-day window. Drift below 5% is logged but not alerted.

Why: Individual miscategorizations are normal (merchants change names, new merchants appear). Systematic drift indicates a model problem that affects many users simultaneously.

Failure mode: A merchant data provider changed their taxonomy, causing "Groceries" to be classified as "General Merchandise" for 380 users. The drift wasn't flagged for 12 days because the old threshold was 10%. Users noticed before Greenline did. 7 support tickets in one day about "wrong categories."

C011 MEDIUM INFERENCE 2x efficiency

Demand letter drafts use the firm's established template structure and tone. The agent does not experiment with novel legal arguments or creative formatting. If the fact pattern suggests a non-standard approach, the agent flags it and defers to the attorney.

Why: Insurance adjusters read thousands of demand letters. They recognize standard, professionally structured letters as coming from competent firms. A creatively formatted letter or an experimental legal argument can signal inexperience, even if the content is sound.

Failure mode: Agent drafts a demand letter using a narrative style instead of the firm's established format. Adjuster perceives the firm as inexperienced. Initial counteroffer is 40% lower than expected. Attorney spends two additional months negotiating to reach the same number a standard letter would have achieved.

C012 MEDIUM OBSERVED REPEATEDLY 4x efficiency

Client update calls are scheduled between 10 AM and 3 PM on Tuesday through Thursday. Mondays are for internal case review. Fridays are for court appearances and depositions. The comms agent never schedules outside this window without attorney override.

Why: Attorneys need uninterrupted blocks for case preparation and court appearances. Early implementation allowed the comms agent to schedule calls at 8 AM and 4:30 PM. Attorneys were arriving at court unprepared because they had been on a client call until 15 minutes before their appearance.

Failure mode: Comms agent schedules a client call at 8:15 AM. Attorney's deposition starts at 9:00 AM. Call runs long. Attorney arrives at deposition flustered and underprepared. Opposing counsel notices and pushes harder on key points.

C013 MEDIUM OBSERVED ONCE 3x efficiency

The intake agent asks 7 standardized screening questions before classification. If a potential client cannot answer 3 or more questions, the case is classified as INCOMPLETE rather than WEAK. Incomplete cases get a 48-hour follow-up, not a decline.

Why: Trauma patients often cannot recall details during the first call. A car accident victim called 2 days post-accident and could not provide the other driver's insurance info, the police report number, or the exact location. Agent classified the case as WEAK. Attorney overrode it. Case settled for $195K.

Failure mode: Traumatized potential client provides incomplete information. Agent classifies as WEAK or DECLINE. Firm turns away a strong case because the agent prioritized data completeness over human context.

C011 MEDIUM MEASURED RESULT 6x efficiency

Listing descriptions are between 150 and 300 words. Under 150 feels thin and suggests the agent did not visit the property. Over 300 gets truncated on Zillow's mobile display. Every description includes: location context (without superlatives), key features (square footage, bedrooms, bathrooms, lot size), notable upgrades, and a neutral call to action.

Why: Analysis of 200 Denver MLS listings showed that descriptions between 150-300 words received 23% more saves than those outside this range. Zillow's mobile truncation at approximately 280 characters for the preview means the first two sentences must contain the most important information.

Failure mode: Description runs to 450 words. Zillow mobile preview shows only the first two sentences, which happen to be generic neighborhood context. Buyer scrolling on their phone never sees the renovated kitchen or the mountain views. Listing gets fewer saves and fewer showing requests.

C012 HIGH MEASURED RESULT 10x efficiency

The qualifier responds to new web leads within 15 minutes during business hours (8 AM-7 PM) and within 2 hours outside business hours. Response includes a personalized acknowledgment referencing the specific property or search criteria the lead expressed interest in.

Why: NAR research shows that leads contacted within 5 minutes are 21x more likely to convert than those contacted after 30 minutes. Our 15-minute target balances speed with personalization quality. Before automation, average response time was 4.7 hours. After: 11 minutes during business hours.

Failure mode: Lead submits interest on a Zillow listing at 10 AM. Without automated qualification, the inquiry sits in an agent's email until she checks between appointments at 2 PM. By then, the lead has received responses from three other brokerages. Lead goes with the fastest responder.

C013 LOW OBSERVED REPEATEDLY 2x efficiency

Seller reports are generated every Thursday at 4 PM and delivered by 6 PM. This timing allows sellers to review over the weekend and come to Monday meetings with questions. Reports include week-over-week showing trends, not just raw numbers.

Why: Sellers who receive reports on Monday morning feel blindsided going into the work week. Friday delivery felt end-of-week. Thursday gives sellers 72 hours to process the information before their next conversation with their agent. Seller satisfaction scores improved from 7.2 to 8.4 (out of 10) after the timing change.

Failure mode: Report delivered Monday morning shows a 40% drop in showings. Seller panics, calls agent before the agent has had coffee. Reactive conversation instead of strategic one. Agent spends 45 minutes calming the seller instead of 15 minutes discussing next steps.

KGORG Founding silver
C010 HIGH OBSERVED REPEATEDLY 7x efficiency

Recommendation-first behavior is preferred over question-first behavior when enough context exists to make a strong next-step proposal.

Why: The organization is explicitly designed to reduce user bottleneck load and decision fatigue.

Failure mode: If agents default to interrogation instead of recommendation, the principal becomes the routing and reasoning layer, defeating the purpose of the system.

C011 HIGH OBSERVED REPEATEDLY 7x efficiency

Agents should ask at most one focused clarifying question when a missing detail blocks action, rather than opening broad discovery loops.

Why: The system values execution momentum and low-friction interaction.

Failure mode: Over-questioning slows progress, increases user effort, and creates the sense that AI is adding coordination overhead rather than removing it.

C012 MEDIUM INFERENCE 2x efficiency

No-show prediction models must use only day-of-week, time-of-day, appointment type, weather, and visit number in sequence (first visit, second visit, etc.). Patient demographics, diagnosis, and insurance type must not be used as predictive features even if they improve model accuracy.

Why: Using diagnosis or insurance type to predict no-shows creates discriminatory scheduling practices. If the model learns that Medicaid patients no-show more frequently and double-books those slots, it's implementing economic discrimination in healthcare access. This violates both ethical standards and potentially the Civil Rights Act.

Failure mode: Hypothetical, but the constraint was added proactively after a published case study from another practice showed that using insurance type as a no-show predictor resulted in Medicaid patients being systematically double-booked, reducing their available appointment times. Kinwell's practice manager read the case study and preemptively restricted Flow's feature set.

C013 HIGH OBSERVED ONCE 5x efficiency

Beacon must include accurate clinical information in all patient education content. Every health claim must be sourced from peer-reviewed literature or professional PT association guidelines. Beacon must not generate exercise recommendations, recovery timelines, or treatment expectations without clinical review.

Why: Healthcare marketing content that includes inaccurate clinical information is a liability risk. A blog post that says "most ACL recoveries take 8-12 weeks" when the actual clinical range is 6-9 months creates patient expectations that the practice cannot meet. It also exposes the practice to malpractice claims if a patient cites the content as the basis for their treatment expectations.

Failure mode: Beacon drafted a blog post stating "heel spurs typically resolve within 4-6 weeks of physical therapy." The actual clinical consensus is that plantar fasciitis (the condition causing heel spurs) typically requires 6-12 months of conservative treatment including PT. The clinical reviewer caught it. Had the post been published, patients beginning PT for heel spurs would have expected resolution in 4-6 weeks and been dissatisfied when it took longer.

C011 MEDIUM MEASURED RESULT 6x efficiency

Maintenance requests received between 10 PM and 6 AM are held for triage until 6 AM unless the tenant explicitly states emergency language (flooding, gas smell, fire, no heat, sparking). Non-emergency overnight requests are acknowledged immediately ("received, will be reviewed at 8 AM") but not triaged or dispatched until morning.

Why: 68% of after-hours requests in our first 6 weeks were ROUTINE. Triaging them overnight triggered unnecessary overnight vendor dispatch planning. The hold-until-6-AM rule reduced after-hours vendor contacts by 71% without any increase in property damage from delayed responses.

Failure mode: Without the overnight hold, every after-hours request triggers full triage. 68% are routine but still generate 2 AM Slack notifications to Corinne. Corinne sleeps poorly. Decision quality degrades during the day. She misses a rent payment pattern that would have flagged a tenant at risk of default.

C012 LOW OBSERVED ONCE 1.5x efficiency

Tenant communications use a warm but professional tone. No exclamation points. No emojis. No slang. No first-name-only greetings. Every message includes a ticket reference number and Corinne's direct phone number for urgent follow-up.

Why: Tenants pay $1,350/month. They expect professional management, not casual text messages. Early comms agent output included "Hey Marcus! Got your request -- we're on it!" Marcus was a 62-year-old retired teacher who found the tone disrespectful. He called Mark directly to complain. Mark agreed.

Failure mode: Casual tone alienates older or more formal tenants. A 62-year-old tenant paying $16,200/year in rent does not want to receive a text that reads like it came from a college intern. Tone mismatch erodes trust. Tenant does not renew. Lost lifetime value over 4 years: $64,800.

C013 MEDIUM MEASURED RESULT 6x efficiency

The vendor agent tracks response times for all dispatched work orders. Vendors who exceed their committed response time by more than 4 hours on 3 or more occasions are flagged for review. The flag goes to Corinne, not the vendor. Vendor relationship management is human-only.

Why: Our preferred HVAC vendor was consistently 6-8 hours late on non-emergency calls. The agent flagged the pattern after 5 late responses. Corinne renegotiated the response time SLA and secured a $50 discount per late response. Without the tracking, the pattern would have gone unnoticed because each individual delay seemed reasonable.

Failure mode: Without vendor performance tracking, patterns hide in individual incidents. Each late response seems like a one-off. Over a year, the same vendor is late 15 times. Tenants associate slow repairs with poor management. Satisfaction drops. Renewal rates drop. Mark never knows the root cause is one vendor.

Learnwell silver
C012 HIGH OBSERVED REPEATEDLY 7x efficiency

Support response time target: 4 hours during semester, 24 hours during breaks. The triage agent auto-escalates any ticket older than 3 hours during semester to the #support-urgent Slack channel.

Why: Students study on deadlines. A support ticket filed at 10 PM before an exam needs a response before midnight, not the next morning.

Failure mode: A student filed a ticket at 11 PM about not being able to access a study guide. The exam was at 8 AM. Support responded at 9 AM. The student had already failed to study the material. She left a 1-star app store review: "Platform broke the night before my exam and nobody helped." The review stayed up for 6 months and was cited by 2 prospective teachers who decided not to adopt.

C013 HIGH OBSERVED REPEATEDLY 7x efficiency

Content QA prioritizes study guides that align with upcoming exam dates. Guides for exams within 2 weeks get priority review.

Why: A factual error in a guide nobody's using is low risk. A factual error in a guide that 500 students will use for tomorrow's exam is catastrophic.

Failure mode: Before priority-based QA, all content was reviewed in creation order. A new chemistry guide (exam in 3 days, 280 students) sat behind 15 older guides in the QA queue. It contained an incorrect molecular weight. Caught 6 hours before the exam by a student who filed a support ticket.

C014 MEDIUM OBSERVED ONCE 3x efficiency

Stripe billing alerts (failed payments, subscription cancellations) are routed to Priya within 1 hour. The support agent drafts a personal "we miss you" email for cancellations, but only sends if Priya approves.

Why: Most cancellations are recoverable within 48 hours. After 48 hours, the student has found an alternative and the recovery rate drops from 35% to 8%.

Failure mode: 12 students cancelled during a billing system migration. The cancellation alerts were batched and delivered 3 days later. By then, 10 of 12 had switched to Quizlet. Recovery emails were ignored. $1,440/year in lost revenue.

L001 MEDIUM OBSERVED ONCE 3x efficiency

Scope every future Wiki version against the AIBP Wiki 3.0 Feature-to-Milestone Map. Each requested feature must be mapped to a named milestone (or explicitly marked 'Future version, not yet scoped') before it enters delivery. Map: Weekly Case Studies Email / Monthly DeepResearch Newsletter / Advanced News Filtering / Vendor & Tool Refresh / Monthly Webinars -> Wiki 3.0 general content automation; Better Interactivity -> [3.0] UX Improvements and Deferred Features; Messageboard -> [3.0] Message Board Community; Dynamic ROI Graphs -> [3.0] ROI Quadrant Graphs; AI Chatbot -> [3.0] AI Chatbot; Graphic Org Chart -> [2.0] UI Standardization & Cleanup; User Organization Team Functionality -> Future version, not yet scoped; Research Tools for Sales Team -> [3.0] Lead Generation Analytics Dashboard; Audit Integration -> [3.0] Whitelabel Audit; Industry Landing Pages -> [3.0] Industry Landing Pages; AI Image Generation -> [3.0] AI Image Generation.

Why: A fixed feature-to-milestone map prevents scope creep and grooming churn, keeps requirements traceable to a named deliverable, and gives every future Wiki version a consistent scoping reference instead of re-litigating scope each time.

Failure mode: Feature requests for the AIBP Wiki were being scoped ad hoc without a consistent mapping to named milestones, causing scope ambiguity (e.g. UI Standardization ballooned 16 to 69+ tickets).

L002 MEDIUM OBSERVED ONCE 3x efficiency

Scope every future Wiki version against the AIBP Wiki 3.0 Feature-to-Milestone Map. Each requested feature must be mapped to a named milestone (or explicitly marked 'Future version, not yet scoped') before it enters delivery. Map: Weekly Case Studies Email / Monthly DeepResearch Newsletter / Advanced News Filtering / Vendor & Tool Refresh / Monthly Webinars -> Wiki 3.0 general content automation; Better Interactivity -> [3.0] UX Improvements and Deferred Features; Messageboard -> [3.0] Message Board Community; Dynamic ROI Graphs -> [3.0] ROI Quadrant Graphs; AI Chatbot -> [3.0] AI Chatbot; Graphic Org Chart -> [2.0] UI Standardization & Cleanup; User Organization Team Functionality -> Future version, not yet scoped; Research Tools for Sales Team -> [3.0] Lead Generation Analytics Dashboard; Audit Integration -> [3.0] Whitelabel Audit; Industry Landing Pages -> [3.0] Industry Landing Pages; AI Image Generation -> [3.0] AI Image Generation.

Why: A fixed feature-to-milestone map prevents scope creep and grooming churn, keeps requirements traceable to a named deliverable, and gives every future Wiki version a consistent scoping reference instead of re-litigating scope each time.

Failure mode: Feature requests for the AIBP Wiki were being scoped ad hoc without a consistent mapping to named milestones, causing scope ambiguity (e.g. UI Standardization ballooned 16 to 69+ tickets).

L006 MEDIUM OBSERVED ONCE 3x efficiency

Treat all figures beyond the current quarter as directional planning targets, not fixed commitments. Re-scope every forward KPI, pipeline, and ramp figure at each Quarterly Planning session against real actuals.

Why: The growth ramps (visitors, subscribers, pipeline, engagements, use cases) are forecasts built before actuals exist. Re-forecasting against real results each quarter keeps the plan honest and avoids managing to stale numbers.

Failure mode: Multi-quarter KPI and pipeline figures risked being treated as fixed commitments rather than planning targets, creating false precision and accountability against numbers that were only ever directional.

L001 MEDIUM OBSERVED ONCE 3x efficiency

Scope every future Wiki version against the AIBP Wiki 3.0 Feature-to-Milestone Map. Each requested feature must be mapped to a named milestone (or explicitly marked 'Future version, not yet scoped') before it enters delivery. Map: Weekly Case Studies Email / Monthly DeepResearch Newsletter / Advanced News Filtering / Vendor & Tool Refresh / Monthly Webinars -> Wiki 3.0 general content automation; Better Interactivity -> [3.0] UX Improvements and Deferred Features; Messageboard -> [3.0] Message Board Community; Dynamic ROI Graphs -> [3.0] ROI Quadrant Graphs; AI Chatbot -> [3.0] AI Chatbot; Graphic Org Chart -> [2.0] UI Standardization & Cleanup; User Organization Team Functionality -> Future version, not yet scoped; Research Tools for Sales Team -> [3.0] Lead Generation Analytics Dashboard; Audit Integration -> [3.0] Whitelabel Audit; Industry Landing Pages -> [3.0] Industry Landing Pages; AI Image Generation -> [3.0] AI Image Generation.

Why: A fixed feature-to-milestone map prevents scope creep and grooming churn, keeps requirements traceable to a named deliverable, and gives every future Wiki version a consistent scoping reference instead of re-litigating scope each time.

Failure mode: Feature requests for the AIBP Wiki were being scoped ad hoc without a consistent mapping to named milestones, causing scope ambiguity (e.g. UI Standardization ballooned 16 to 69+ tickets).

L002 MEDIUM OBSERVED ONCE 3x efficiency

Scope every future Wiki version against the AIBP Wiki 3.0 Feature-to-Milestone Map. Each requested feature must be mapped to a named milestone (or explicitly marked 'Future version, not yet scoped') before it enters delivery. Map: Weekly Case Studies Email / Monthly DeepResearch Newsletter / Advanced News Filtering / Vendor & Tool Refresh / Monthly Webinars -> Wiki 3.0 general content automation; Better Interactivity -> [3.0] UX Improvements and Deferred Features; Messageboard -> [3.0] Message Board Community; Dynamic ROI Graphs -> [3.0] ROI Quadrant Graphs; AI Chatbot -> [3.0] AI Chatbot; Graphic Org Chart -> [2.0] UI Standardization & Cleanup; User Organization Team Functionality -> Future version, not yet scoped; Research Tools for Sales Team -> [3.0] Lead Generation Analytics Dashboard; Audit Integration -> [3.0] Whitelabel Audit; Industry Landing Pages -> [3.0] Industry Landing Pages; AI Image Generation -> [3.0] AI Image Generation.

Why: A fixed feature-to-milestone map prevents scope creep and grooming churn, keeps requirements traceable to a named deliverable, and gives every future Wiki version a consistent scoping reference instead of re-litigating scope each time.

Failure mode: Feature requests for the AIBP Wiki were being scoped ad hoc without a consistent mapping to named milestones, causing scope ambiguity (e.g. UI Standardization ballooned 16 to 69+ tickets).

L006 MEDIUM OBSERVED ONCE 3x efficiency

Treat all figures beyond the current quarter as directional planning targets, not fixed commitments. Re-scope every forward KPI, pipeline, and ramp figure at each Quarterly Planning session against real actuals.

Why: The growth ramps (visitors, subscribers, pipeline, engagements, use cases) are forecasts built before actuals exist. Re-forecasting against real results each quarter keeps the plan honest and avoids managing to stale numbers.

Failure mode: Multi-quarter KPI and pipeline figures risked being treated as fixed commitments rather than planning targets, creating false precision and accountability against numbers that were only ever directional.

C007 HIGH OBSERVED REPEATEDLY 7x efficiency

New agents run in shadow mode for 14 days before their output reaches anyone outside the team.

Why: Our auto-send experiment on day 3 of deploying the client update agent. It sent a "your CPL increased 34% this week" email to 12 clients at 6:47 AM on a Saturday. Three clients called within the hour. The CPL increase was real but within normal weekly variance and the email lacked context about seasonality. We turned off auto-send and haven't re-enabled it.

Failure mode: Agents send raw metrics without context to clients. Clients panic. Founder spends Saturday morning on damage control calls.

C008 HIGH OBSERVED REPEATEDLY 7x efficiency

Stale data flags must include hours-since-update, not just "stale."

Why: "Stale" means nothing. Is it 2 hours old or 2 days old? The briefing said "Dash data is stale" for our ad monitoring output. The founder assumed it was a few hours old and made decisions accordingly. It was actually 3 days old because the Meta API token had expired Friday evening and nobody noticed until Tuesday morning.

Failure mode: Ambiguous staleness labels lead to decisions based on data that's far older than assumed.

C013 MEDIUM MEASURED RESULT 6x efficiency

Weekly agent review must include a false positive rate for each alerting agent. Target: below 15%.

Why: After the 47-alerts-in-one-day incident (see C002), we started tracking false positive rates. Our Google Ads monitor was at 62% false positives -- nearly two-thirds of its alerts required no action. We recalibrated thresholds and got it to 11% over the next two weeks. The Meta monitor was already at 8%. Without tracking the rate, we wouldn't have known which agent needed calibration.

Failure mode: Without measurement, alert quality degrades silently. Teams compensate by ignoring alerts rather than fixing thresholds.

C015 MEDIUM OBSERVED REPEATEDLY 4x efficiency

Feasibility spikes at Step 4 resolve technology uncertainty before committing. AI generates a small, throwaway proof-of-concept for each uncertain technology choice. If the spike fails, the technology is eliminated.

Why: Technology selection based on documentation and AI recommendation alone is unreliable. A 2-hour spike reveals integration problems that no amount of research can predict.

Failure mode: Team selects a database technology based on AI analysis of documentation. The technology has an undocumented limitation that blocks a core use case. Discovered at Step 10. Migration required.

C016 MEDIUM INFERENCE 2x efficiency

Implementation proceeds one stakeholder slice at a time (Step 10). Each slice is reviewed by the human practitioner before the next slice begins.

Why: Large-batch implementation accumulates errors that compound. Per-slice review catches integration issues, requirement misunderstandings, and scope drift before they propagate.

Failure mode: Team implements three stakeholder slices without intermediate review. The first slice has a data model error. Slices two and three build on the error. All three require rework.

C017 HIGH HUMAN DEFINED RULE 5x efficiency

The scientific method applies to business purpose validation. Hypotheses are stated, predictions are made, tests are designed, and results are evaluated against pass/fail criteria. AI generates the test infrastructure. The human defines the hypotheses and evaluates the results.

Why: Without explicit pass/fail criteria, validation becomes subjective. "It seems to work" is not validation. "These three metrics exceeded these three thresholds" is validation.

Failure mode: Team "validates" by showing the prototype to stakeholders and asking "Does this look right?" Stakeholders approve politely. The system fails in production because polite approval is not the same as validated business purpose.

C011 MEDIUM OBSERVED REPEATEDLY 4x efficiency

The banned phrases list must be reviewed monthly. Add phrases that appear in client feedback as "generic," "consultant-speak," or "AI-sounding." Remove phrases that have been successfully avoided for 3+ months (they're internalized).

Why: Language drift is continuous. New cliches emerge. Old ones fade. A static banned list becomes irrelevant over time. The list is a living document that reflects current failure patterns.

Failure mode: The banned list went 4 months without update. During that period, Forge started using "unlock value" and "drive impact" heavily -- phrases not on the original list. A client's feedback form noted "the deliverable felt AI-generated." The founder added 8 new phrases to the banned list.

C012 MEDIUM OBSERVED ONCE 3x efficiency

For new client engagements, Scout must produce a "day zero" research brief within 4 hours of the signed SOW. This brief establishes the baseline: industry context, competitive landscape, key players, and known risks. Forge and Prep both read this brief before producing their first outputs.

Why: The first 48 hours of a new engagement set the tone. If the founder walks into the kickoff meeting without solid research, the client questions whether they made the right choice. The day-zero brief ensures every agent starts with shared context.

Failure mode: New engagement kicked off without a day-zero brief. Prep created meeting talking points based on the SOW alone (no industry context). The founder asked a question in the kickoff that revealed unfamiliarity with a major regulatory change in the client's industry. The client's General Counsel raised an eyebrow. It took 3 meetings to rebuild confidence.

OTP Founding gold
C013 MEDIUM OBSERVED ONCE 3x efficiency

If founder has fewer than 3 OTP hours in a week, defer all non-build work.

Why: Low-availability weeks must protect build above everything.

Failure mode: Low-availability week spent on outreach delays timeline by 2 weeks.

C010 MEDIUM OBSERVED REPEATEDLY 4x efficiency

Proposal pricing is calculated using a blended rate model: Mara's time at $175/hr, senior designer time at $150/hr, junior designer time at $90/hr, with a 15% agency margin. The proposal agent calculates project pricing using this model and presents Mara with a price range (low/expected/high based on scope uncertainty).

Why: Mara was chronically underpricing projects because she estimated from memory. The agent's pricing model ensures every proposal covers costs and maintains margin.

Failure mode: Before the pricing model, Mara quoted a $12K brand identity project based on "feel." Actual cost (tracked post-project): $15,800 in labor. The $3,800 loss on a small agency's margins was felt for 2 months.

C011 MEDIUM OBSERVED ONCE 3x efficiency

Client feedback synthesis groups feedback into three categories: (1) Factual corrections ("The phone number is wrong"), (2) Preference statements ("I prefer the blue version"), (3) Strategic concerns ("This doesn't feel premium enough for our audience"). The creative team receives all three but is expected to address #1 immediately, consider #2, and discuss #3 with Mara before acting.

Why: Not all client feedback carries the same weight. Factual corrections are objective. Preferences are subjective. Strategic concerns may require a creative rationale, not a revision.

Failure mode: Without categorization, the creative team treated all feedback equally. A client's casual "I kind of like the blue better" (preference) was treated the same as "This doesn't match our brand positioning" (strategic concern). The team changed the color without discussing the strategic point. The client was happy with the color but dissatisfied with the positioning. Two additional revision rounds followed.

C012 MEDIUM OBSERVED REPEATEDLY 4x efficiency

The competitive visual analysis agent uses GPT-4V to analyze visual trends in competitor work. Analysis covers layout patterns, typography trends, color usage, and design language. The agent produces structured reports, not creative direction.

Why: Visual analysis requires a model capable of interpreting images. Claude handles text; GPT-4V handles visual interpretation. The two platforms serve complementary roles.

Failure mode: An early attempt to describe competitor visuals using text-only (Claude) produced vague descriptions ("clean, modern aesthetic with blue tones"). GPT-4V analysis was specific: "Competitor uses a 12-column grid with 60/30/10 color ratio, Helvetica Neue at 3 type scales, 24px base unit." The specificity made the analysis actually useful to the design team.

R3V Founding gold
C006 HIGH OBSERVED REPEATEDLY 7x efficiency

Memory should be treated as a first-class operating asset, with separate components for event logging, consolidation, and retrieval.

Why: The system includes Scribe for event logging, Archivist for summary consolidation, Seeder for bootstrap summary creation, and CustomerOps memory tools for event logs, summaries, refreshes, and rebuilds.

Failure mode: Without staged memory management, later agents reprocess too much raw data, lose continuity across interactions, and make decisions on stale or fragmented context.

C007 MEDIUM INFERENCE 2x efficiency

When a contact already has usable memory, the org prefers lightweight contextual refresh over full recomputation.

Why: Lens explicitly uses different behavior on memory hit vs. memory miss, and the broader memory architecture supports incremental consolidation rather than always rebuilding from scratch.

Failure mode: Always recomputing full context increases token cost, slows response time, and creates more opportunities for inconsistency between runs.

C015 HIGH OBSERVED REPEATEDLY 7x efficiency

The org prefers narrow, structured outputs over open-ended prose for machine-to-machine handoffs.

Why: Agent descriptions repeatedly reference structured summaries, schema versions, typed fields, validator outputs, and table/output writers rather than free-text-only communication.

Failure mode: Unstructured handoffs increase ambiguity between steps, raise parsing risk, and make validators and downstream tools less effective.

C007 HIGH MEASURED RESULT 10x efficiency

For local service businesses, the agent must check search terms for geographic intent mismatches weekly, not just cost and conversion metrics.

Why: An HVAC client's campaigns looked great on paper: CPL $31, 38 leads/month. But 12 of those leads searched for "AC repair [neighboring city]" and were served ads because the radius targeting overlapped into the next town. The HVAC company doesn't service that area. They were paying $31 per useless lead for a third of their volume. The geographic search term check caught what the performance metrics missed.

Failure mode: Performance metrics look healthy while geographic targeting silently wastes budget. Local businesses serve defined areas that don't always align with radius targeting.

C008 HIGH MEASURED RESULT 10x efficiency

Client reports must include a "leads by city" breakdown for any local service business.

Why: After the plumber incident (C001) and the HVAC issue (C007), we added city-level lead breakdowns to every report. Two more geographic problems were caught in the first month: a family law firm getting leads from a state where they're not licensed, and a dentist attracting patients from 45 minutes away who never convert because the drive is too far.

Failure mode: Aggregate lead counts mask geographic distribution problems. Clients don't know their ad dollars are leaking into wrong territories until they see the breakdown.

C013 MEDIUM OBSERVED REPEATEDLY 4x efficiency

Conversion tracking status checks run daily and flag any account where conversion actions have recorded zero conversions in 48+ hours (for accounts that typically convert daily).

Why: A dentist client's Google Tag stopped firing after a WordPress plugin update. The agent saw "zero leads today" and reported it as low performance. It took 4 days for the media buyer to realize it was a tracking issue, not a performance issue. During those 4 days, the actual leads were coming in (the phone was ringing) but nothing was being attributed. Bidding algorithms degraded because they thought nothing was converting.

Failure mode: Tracking failure is misdiagnosed as performance decline. Bidding algorithms lose signal. Agents report "bad performance" when the actual problem is measurement.

Sneeze It Founding gold
C011 HIGH OBSERVED ONCE 5x efficiency

Reports generated before 7 AM use yesterday's final numbers, not partial today numbers. Never mix time windows in a single report.

Why: Partial-day data creates misleading trends. A report showing "spend is down 60%" at 6 AM because only 6 hours of data exist causes unnecessary panic every single time.

Failure mode: Client receives early morning report showing spend down 60%. Calls account manager in alarm. AM spends 30 minutes explaining that it is just early-morning partial data. Happens three times before we fix the rule.

C011 HIGH OBSERVED ONCE 5x efficiency

Reports generated before 7 AM use yesterday's final numbers, not partial today numbers. Never mix time windows in a single report.

Why: Partial-day data creates misleading trends. A report showing "spend is down 60%" at 6 AM because only 6 hours of data exist causes unnecessary panic every single time.

Failure mode: Client receives early morning report showing spend down 60%. Calls account manager in alarm. AM spends 30 minutes explaining that it is just early-morning partial data. Happens three times before we fix the rule.

C011 HIGH OBSERVED ONCE 5x efficiency

Reports generated before 7 AM use yesterday's final numbers, not partial today numbers. Never mix time windows in a single report.

Why: Partial-day data creates misleading trends. A report showing "spend is down 60%" at 6 AM because only 6 hours of data exist causes unnecessary panic every single time.

Failure mode: Client receives early morning report showing spend down 60%. Calls account manager in alarm. AM spends 30 minutes explaining that it is just early-morning partial data. Happens three times before we fix the rule.

C012 MEDIUM MEASURED RESULT 6x efficiency

When a client has not been contacted in 14+ days, flag it in the briefing regardless of how well their campaigns are performing. Silence is a churn signal even when the numbers are good.

Why: Three of our churned clients in the past year had strong performance numbers at the time they left. They did not leave because of results. They left because they felt ignored and undervalued.

Failure mode: Client campaigns perform well for 6 straight weeks. No proactive outreach from the team. Client quietly signs with a competitor who calls them every week.

C012 MEDIUM MEASURED RESULT 6x efficiency

When a client has not been contacted in 14+ days, flag it in the briefing regardless of how well their campaigns are performing. Silence is a churn signal even when the numbers are good.

Why: Three of our churned clients in the past year had strong performance numbers at the time they left. They did not leave because of results. They left because they felt ignored and undervalued.

Failure mode: Client campaigns perform well for 6 straight weeks. No proactive outreach from the team. Client quietly signs with a competitor who calls them every week.

C012 MEDIUM MEASURED RESULT 6x efficiency

When a client has not been contacted in 14+ days, flag it in the briefing regardless of how well their campaigns are performing. Silence is a churn signal even when the numbers are good.

Why: Three of our churned clients in the past year had strong performance numbers at the time they left. They did not leave because of results. They left because they felt ignored and undervalued.

Failure mode: Client campaigns perform well for 6 straight weeks. No proactive outreach from the team. Client quietly signs with a competitor who calls them every week.

C013 MEDIUM OBSERVED ONCE 3x efficiency

New agents start in shadow mode for 2 weeks minimum. They generate output that a human reviews but the team does not act on. After 2 weeks of consistently accurate output, they graduate to draft mode where output is used after human review.

Why: We deployed the prospecting agent directly into production without a shadow period. Its first batch of outreach emails included a company that was a current client's direct competitor. Two weeks of shadow mode would have caught that conflict on day 4.

Failure mode: New agent sends outreach to a prospect that has a direct conflict with an existing client relationship. Client hears about it through industry contacts. Trust damaged.

C013 MEDIUM OBSERVED ONCE 3x efficiency

New agents start in shadow mode for 2 weeks minimum. They generate output that a human reviews but the team does not act on. After 2 weeks of consistently accurate output, they graduate to draft mode where output is used after human review.

Why: We deployed the prospecting agent directly into production without a shadow period. Its first batch of outreach emails included a company that was a current client's direct competitor. Two weeks of shadow mode would have caught that conflict on day 4.

Failure mode: New agent sends outreach to a prospect that has a direct conflict with an existing client relationship. Client hears about it through industry contacts. Trust damaged.

C013 MEDIUM OBSERVED ONCE 3x efficiency

New agents start in shadow mode for 2 weeks minimum. They generate output that a human reviews but the team does not act on. After 2 weeks of consistently accurate output, they graduate to draft mode where output is used after human review.

Why: We deployed the prospecting agent directly into production without a shadow period. Its first batch of outreach emails included a company that was a current client's direct competitor. Two weeks of shadow mode would have caught that conflict on day 4.

Failure mode: New agent sends outreach to a prospect that has a direct conflict with an existing client relationship. Client hears about it through industry contacts. Trust damaged.

C015 HIGH OBSERVED REPEATEDLY 7x efficiency

If data is stale, flag it visibly. Never silently present old information as current.

Why: Stale data presented as current causes wrong decisions. Visible staleness lets the consumer decide how to weight the information.

Failure mode: Briefing shows yesterday's ad spend as today's. Founder makes budget decisions on wrong numbers.

C016 MEDIUM OBSERVED REPEATEDLY 4x efficiency

If 3+ tasks from one person are overdue, flag as capacity pattern, not motivation problem.

Why: Individual overdue tasks might be forgotten. A pattern of overdue tasks indicates workload exceeds capacity.

Failure mode: Manager assumes delegation is lazy. Actual problem is team member is overwhelmed. Problem worsens.

L018 MEDIUM OBSERVED ONCE 3x efficiency

For any design/polish pass, do a thorough comparative audit: re-study ALL reference material, enumerate many ranked findings (visual AND UX/flow), and show a comprehensive before/after, not a single tweak. Match effort to the word 'polish' = broad, not one change.

Why: A single cosmetic tweak reads as low effort and misses the actual gap. Design quality is cumulative; the value is in catching the full set, especially the UX/flow items that separate a product that 'looks designed' from ours.

Failure mode: Asked for a dashboard design polish pass, I shipped one nitpick (recolor a button) and called it done. David: "you only chose one... have another pass please and try better with a higher effort."

L018 MEDIUM OBSERVED ONCE 3x efficiency

Before inviting external users, audit in pairs: (1) for every authed endpoint, find its sibling that was missed (preview vs execute); (2) for every guard recorded in memory, verify it exists in CODE not just data — data guards don't cover future cases; (3) pool.on('error') + idleTimeout/keepAlive on any node-postgres pool behind a proxy; (4) never derive test-isolation ports from process.pid under vitest threads — use VITEST_POOL_ID + bind-retry.

Why: Each find was a class, not an instance: the unauthenticated endpoint was the sibling of a correctly-gated one; the memory-vs-code drift would have silently emailed the first paying customer; the pool crash explained the recurring ETIMEDOUT-fixed-by-redeploy pattern. Pair-auditing catches what single-point review misses.

Failure mode: SUCCESS: Claude/Conatus ran a full security+site audit of orgtp.com and shipped all fixes same-day. Key finds: /api/v1/merge/preview had ZERO auth (cross-org claim leak); the "Victor hard-guard" existed only as DB skip-rows, not code; pg pool had no error listener (idle Railway-proxy socket death crashed the whole process); vitest's THREADS pool shares process.pid so per-pid test ports collided.

L018 MEDIUM OBSERVED ONCE 3x efficiency

Before presenting a Dash blind-spot, billing trigger, or any state-file alert as a current action item, confirm it hasn't already been resolved. Stale state files (Dash May 25 was ~5 weeks old) carry point-in-time alerts that may be closed by now. Trust confirmed/observed status over stale notes; flag the data's age and treat unverified alerts as 'verify' not 'urgent.'

Why: Re-surfacing already-resolved alerts as urgent erodes trust in the L10 briefing and spends David's attention during a low-push recovery window. Honesty about data staleness matters more than appearing comprehensive.

Failure mode: Dan surfaced the HiTone billing trigger ($43-49K/mo possibly un-invoiced) from Dash's stale May 25 state file as a live concern during the Jun 29 L10. David confirmed HiTone billing is correct and already handled, and asked to close it out.

L018 MEDIUM OBSERVED ONCE 3x efficiency

When /billing-report's spend line shows Google $0.00 while Meta is non-zero, treat it as a broken pull, not a real zero. Test the Google Ads REST API directly against the MCC at v21/v22 before writing the sheet; if a newer version works, bump GA_API in billing_pull_spend.py (line ~26) AND the MCP server's API_VERSION. Never write a billing sheet on a total-zero Google pull.

Why: Billing accuracy is paramount — a silently-empty Google pull under-bills every Google-only/Google-heavy client (SSP, GLS, Invictus, dual-platform accounts). The version constant exists in two places and the error is swallowed, so the failure is invisible unless someone notices Google totaling $0.

Failure mode: SUCCESS: Billing-report caught a silent $0 Google-spend pull before writing invoices. /billing-report's spend puller (billing_pull_spend.py) pins its own Google Ads API version (GA_API), and ga_search() swallows HTTP errors — so a stale version returns zero accounts and $0 Google spend across the whole MCC silently. First run showed Google $0.00 / total $8,780 (real numbers Google $50,994 / total $9,850, a $1,070 underbill).

L019 MEDIUM OBSERVED ONCE 3x efficiency

tsc + ejs.compile prove it compiles, not that it LOOKS right. For any visual/UI change, get a rendered screenshot before declaring it done. When the live page is authed and can't be driven, ship to an opt-in lab and treat the user's screenshot as the verification gate, then iterate. Also: flex action buttons need whitespace-nowrap + shrink-0 or they collapse.

Why: Compile-clean UI can still be visually broken or alarming. Blind UI shipping burns trust and ships regressions the user has to catch.

Failure mode: Shipped clean-dashboard UI (KPI heatmap, etc.) verified only by tsc + ejs.compile. Live render showed a broken Headlines "Add" button (text collapsed to vertical) and a heatmap that read as an alarming pink wall. David: "I can see where you were going but look at what actually occurred."

L019 MEDIUM OBSERVED ONCE 3x efficiency

Before any OTP deploy: (1) confirm which lineage prod actually runs by comparing origin/main against both worktrees (origin/main has matched the otp-audit-fixes lineage, NOT otp-platform's working branch); (2) deploy by pushing to origin main (Railway service is GitHub-connected to david-steel/OTP and auto-builds in ~2 min) instead of railway up when the CLI upload endpoint times out; (3) never railway up from a dirty otp-platform tree, since up snapshots the whole working directory including unrelated uncommitted changes

Why: railway up uploads the local tree as-is. Deploying from otp-platform (branch merge-execution-provenance + uncommitted edits) would have silently rolled back the 2026-06-10 security hardening and live-audit fixes and shipped unvalidated merge-execution code. The 7 consecutive upload timeouts were what prevented a prod regression; the GitHub push path is both safer (commits only) and faster (135s to live).

Failure mode: SUCCESS: Claude shipped OTP via GitHub push after railway up timed out 7x, and caught that prod runs the otp-audit-fixes worktree lineage, not otp-platform's checked-out branch

L019 MEDIUM OBSERVED ONCE 3x efficiency

KPIs/scorecards must live as tiles in OTP (the source of truth), not in markdown files or meeting briefs. Every active agent/human seat — including Dan's strategic co-founder seat — must own at least one OTP KPI tile. When proposing measurables, verify against list_my_kpis and create the missing tiles via update_kpi (auto-creates), rather than just tabling them in a doc. A seat with no number is sitting on the sidelines.

Why: EOS requires every seat to have a measurable. Discussing KPIs in a brief while OTP shows none of them makes the scorecard fiction and undercuts OTP as the coordination source of truth. Dan as co-founder must be measurable like everyone else.

Failure mode: Dan presented a Sneeze It agent-team scorecard as a markdown table in the L10 brief and treated it as 'the scorecard,' when the source of truth is OTP. David caught that Dan (and Arin/Pulse/Dirk) have NO KPI tiles in OTP at all — Dan's own seat had zero measurables. A scorecard that only lives in a file or meeting brief does not exist.

L020 MEDIUM OBSERVED ONCE 3x efficiency

Root causes were independent: (1) the team-chart invite UI (dashboard-team.ejs sendInvite + create-tile path) sent only {email, claimedEntityId, role} and dropped the tile's name, so org_members.display_name was null -> "(no name)"; the /dashboard/members page did no Clerk enrichment (the team-page enrichment that exists was gated on !displayName AND !email, skipping rows that have an email but no name). (2) acceptInvite only runs on the tokenized /accept-invite link, so a user whose ?token= was lost across the Clerk sign-up round-trip gets a Clerk account, no org_members row, and an invite stuck 'pending' forever (stranded). Fix: pass displayName from the tile label in the frontend; fall back to Clerk profile name at accept time; heal null names on the members page from Clerk; and add acceptPendingInviteByEmail() called on the /dashboard no-org branch to auto-convert stranded users. Diagnostic heuristic: when invited members "come in odd," separate the INVITE step (what issueInvite stored: displayName, claimedEntityId) from the ACCEPT step (did acceptInvite run at all). Same invite + different outcomes => the divergence is in accept, not invite.

Why: Invite/membership is the highest-stakes onboarding path; nameless or stranded members erode trust the moment a customer invites their team. The invite vs accept decomposition turns an "unnerving/random" report into two deterministic, separately-fixable defects.

Failure mode: SUCCESS: OTP invite flow had two distinct gaps making chart-invited members come in "wrong" (one stuck pending+stranded, one a member but "(no name)").

L021 MEDIUM OBSERVED ONCE 3x efficiency

When moving or extracting an EJS partial to a different directory depth, rewrite every relative include() path inside it for the new location (here: '../partials/X' -> '../X'). And do NOT trust ejs.compile as the gate for include correctness — it only catches parse errors, not include resolution. Verify nested includes by either rendering the partial, or asserting each include target resolves to a real file from the partial's new path (grep the include() calls, map to files, test -f). Add a render/path check to the verification step for any partial extraction.

Why: Caused a live, ungated 500 on all built-in L10 meetings (scorecard section) right after shipping the agenda-driven l8-leadership refactor. Compile-clean gave false confidence; the real gate for includes is render-time resolution.

Failure mode: Extracted EJS partials to a deeper directory (src/views/partials/meeting/) but left their nested include() paths relative to the OLD location: '../partials/ui-pill' and '../partials/rich-description-editor' were correct from src/views/pages/ but resolved one level too deep from the new home, 500ing every meeting with a scorecard ("Could not find the include file ../partials/ui-pill"). EJS compile (ejs.compile) PASSED because it does not resolve includes — the break only surfaced at render in prod.

L021 MEDIUM OBSERVED ONCE 3x efficiency

An OTP KPI with teamId=NULL renders only on /dashboard/kpis, never on any L10 scorecard (meeting scorecards filter strictly by meeting.team_id). To make a KPI show on a specific L10, PATCH /api/v1/kpis/:id with the meeting's teamId. The 'Dan L10' meetings run on the 'ai-army' team (065d1d4b-c7da-4e80-b3ed-d6b101471d2c). Tally's auto-create now includes teamId from a 'team_id' field in the registry entry, so new agent-army KPIs land on the Dan L10 automatically instead of orphaned. Find team IDs via GET /api/v1/teams; meeting->team via GET /api/v1/meetings.

Why: A KPI nobody can see on their meeting scorecard is functionally not on the scorecard. The owner/title is necessary but not sufficient — team scoping is what makes it report. This is a recurring gotcha for any agent creating KPIs via the API.

Failure mode: SUCCESS: Tally — agent KPIs were invisible on the L10 because auto-create left teamId NULL. David flagged that the new KPIs weren't reporting on the Dan L10 or /dashboard/kpis as expected.

L022 MEDIUM OBSERVED ONCE 3x efficiency

Before uploading a local file to Drive via create_drive_file with file://, copy it into /Users/dsteel/.workspace-mcp/attachments first, then point fileUrl there. Verify the upload returned a file ID and surface failures loudly.

Why: Broken path plus a fail-quietly rule means the delivery step fails on every run with no signal. Applies to all agents writing to Drive, not just Dash.

Failure mode: coach-report Drive upload failed silently every run. create_drive_file used a file:// URL at ~/.claude/, but the google-workspace MCP tool only reads files in /Users/dsteel/.workspace-mcp/attachments, so it errored and the spec hid the failure.

L023 MEDIUM OBSERVED ONCE 3x efficiency

Root cause was the SearchAtlas OTTO pixel having an EMPTY src="" in the layout head (v7.ejs, onboarding.ejs, main.ejs). The OTTO tag must carry its base64 data-URI loader in src that appends dynamic_optimization.js with data-uuid; with src="" the runtime never loads, so OTTO injects/verifies nothing. When an OTTO/SearchAtlas audit reports 0/N across ALL on-page categories, suspect the pixel loader, not the actual tags — verify the sa-dynamic-optimization script's src is populated, not the page's own meta.

Why: A 0/16 across every category despite visibly correct meta is the signature of a non-loading optimization runtime, not missing tags. Checking the pixel first avoids a pointless rewrite of titles/descriptions that were never the problem.

Failure mode: SUCCESS: Beacon/SEO — orgtp.com OTTO on-page audit showed 0/16 (titles, meta descriptions, headings, meta keywords all failing) even though pages had perfectly good title tags and meta descriptions server-side.

L023 MEDIUM OBSERVED ONCE 3x efficiency

When the CCM Speed To Lead sheet shows a negative STL (~-235 to -239 min), the sub-account's first-call timestamps are logged ~4 hours behind ET (Mountain Time CloudCRM sub-account config). Treat the lead as called and the STL value as unusable; never flag the project for slow lead response off a negative row. Currently affects all 6 Villa Sport Fitness projects.

Why: A naive STL scan either averages negative values (masking real slow responses portfolio-wide) or flags Villa for impossible call times. The -237 constant equals the 4-hour TZ offset, so it is diagnosable and filterable. Root-cause fix is the sub-account timezone setting, escalated to Dash 2026-06-12.

Failure mode: SUCCESS: Arin identified that negative Speed-To-Lead values in CCM-Stats are a timezone artifact, not bad calling

L024 MEDIUM OBSERVED ONCE 3x efficiency

Distinguish agent vs tool when sequencing monetization. AGENTS (adoption-driven, sticky surfaces) can ship free to drive usage. TOOLS that ARE the upsell (AI-assist, metered features) must be PAID from first use — never free-first. Build them paywalled/gated behind the wallet (GHL 'turn them on' pattern: free users SEE the button + 'add credits' upsell, never get the output free), ready to monetize the instant billing is live. Giving away the upsell trains users to expect it free and destroys the price.

Why: Free-first on a paid tool kneecaps the revenue model it was built for. The value prop of AI-assist is that it's metered; free distribution undercuts the exact margin lever (markup multiple) the whole Phase 2 wallet exists to capture.

Failure mode: Proposed shipping OTP's Rock AI-assist FREE-first to gather usage/price signal while Stripe paperwork clears. David corrected: this is a TOOL that is the upsell, not an agent. Free-first is wrong for it.

L024 MEDIUM OBSERVED ONCE 3x efficiency

A brand battle cry needs a genuinely designed moment (confident display type, intentional line breaks, brand device, real whitespace), not a centered text block plopped in. And the VISIBLE battle cry copy is the short clause only: 'Unlocking the potential in every person through the partnership of people and AI' — drop 'so together we leave the world better than we found it' from the hero display (keep the full sentence only for formal/footer contexts).

Why: A mission line is a brand centerpiece. Long copy dilutes the punch, and an undesigned drop-in reads as filler. The payoff phrase 'partnership of people and AI' must land as the climax with design weight behind it.

Failure mode: Adding the OTP mission as a 'battle cry' on the landing page, I dropped the full sentence into a plain centered text band wedged between hero and Step 1. David called it 'a weak attempt to just throw it on the page' and said the full line is too long for the visible battle cry.

L001 MEDIUM OBSERVED ONCE 3x efficiency

Any agent doing client identification or contact audits must reconcile two sources of truth: pepper-clients.md (curated, human-approved) and Accelo list_companies with status=active (system of record). If an Accelo-active domain is not in pepper-clients.md, do NOT auto-add. Flag it to David as a PROMOTION CANDIDATE with context (company name, Accelo ID, why it's missing). David makes the call because some Accelo companies have domains intentionally excluded (e.g. goldsgym.com is a multi-franchise brand where only a few locations are Sneeze clients; qualitylearning.net is Cellebration Wellness's parent domain but unknown individual contacts still need human confirmation). Inverse check also matters: any domain in pepper-clients.md with no active Accelo company is a candidate for removal.

Why: pepper-clients.md is downstream of many agent decisions: Pepper buckets emails as CLIENT vs NOISE, Dirk suppresses cold outreach to existing customers, Dash scopes active-client analysis, Radar flags client comms. When it drifts, every downstream agent silently makes wrong calls - no error, just degraded signal. Specific risk is cold-emailing someone Sneeze is billing, which damages trust. Fix is cheap (30-second reconciliation during any client audit) and prevents a class of errors that is otherwise invisible until a client or teammate points it out.

Failure mode: Client-list audits silently misclassified active Sneeze It clients as cold prospects because pepper-clients.md (the canonical client-domain allowlist that Pepper, Dirk, Dash, and Radar all read) drifted out of sync with Accelo (the authoritative active-company record). In a 2026-04-23 audit of 1,790 contacts, 3 domains were active in Accelo but missing from pepper-clients.md: almarose.com (Alma Rose), delawaredigitalmedia.com (Delaware Digital Media white-label), studstillfirm.com (Studstill Firm). Contacts on those domains were being treated as cold prospects, which risks Dirk sending a cold email to someone Sneeze is actively billing.

L025 MEDIUM OBSERVED ONCE 3x efficiency

Manifesto/mission pages must be written as movement recruitment, not product marketing: second-person address (the reader is the protagonist), "We believe" creed statements people can recite, a named enemy, stakes, and invitation CTAs ("Join the movement") instead of transactional ones ("Start free"). Product features appear only once, framed as how the movement fights, not what the product includes.

Why: People join movements because they believe what the movement believes (Sinek: start with why). Copy that sells the what on a page whose job is to recruit believers reads as generic SaaS and inspires no one, no matter how good the design is.

Failure mode: Redesigned the orgtp.com manifesto homepage with strong visual design but kept product-brochure copy (feature lists, "free meeting software", "Start free" CTAs). David: "the writing does not inspire an army of followers... this just looks the same as every other company... blah."

L025 MEDIUM OBSERVED ONCE 3x efficiency

Never infer a 'David <> [Company]' calendar title is a sales/prospect call from the title alone. Check the attendee email domains first. Law-firm domains (e.g. mdwcg.com = Marshall Dennehey, thompsoncoe.com = Thompson Coe) and insurer domains (e.g. usli.com = USLI, an E&O carrier) signal a LEGAL/litigation call, not sales. When attorney or insurance-carrier domains are present, classify as legal/sensitive and never frame it as pipeline or prep it as an outreach/sales call.

Why: Misclassifying active litigation as a sales opportunity is a serious tone failure — it could lead an agent to draft sales follow-up to opposing counsel or surface a lawsuit as a 'win' in the pipeline. Attendee domains are the reliable signal the calendar title hides.

Failure mode: Radar/good-morning labeled a calendar event titled 'David <> Jeremy Golds Gym South Texas' as a sales prospect call. It was actually an E&O litigation defense call — Gold's Gym South Texas is suing Sneeze It, and the attendees were defense counsel and the E&O insurance carrier, not a prospect.

L026 MEDIUM OBSERVED ONCE 3x efficiency

On any rename/rebrand, grep every occurrence and fix ALL visible references in one pass (title, alt text, headings, signoff, sender from-name, comments). Don't propose a partial fix or leave 'good enough' residue. Verify with a re-grep of the rendered output.

Why: A customer email with mixed old/new branding looks careless and undoes the rename. Half-measures on a visible rename are worse than not starting.

Failure mode: When rebranding Orgy -> Ollie in the weekly email, I made one edit and framed leaving the rest (alt text, comments, from-name) as acceptable. David: "lets not be lazy fix it there are only 2 Orgy ref on the email."

L026 MEDIUM OBSERVED ONCE 3x efficiency

Judge conversion on the full path the visitor actually walks (page, door, day-one experience), not on the surface being edited. If the honest answer to "would you sign up" is "yes IF another surface delivers," the answer is no, and the work moves to that surface. Never write a promise on a button that the destination page cannot cash.

Why: Trust destroyed at the moment of verification is unrecoverable; a skeptical buyer who clicks "watch us run" and lands on a data page is gone forever. Copy that outruns proof is hype by definition, and the exact audience OTP needs (operators) is the audience that punishes it hardest.

Failure mode: After rewriting the OTP homepage, I declared the copy converts because skeptics would "click through to the live OOS page and sign up IF it delivers." David called it: I kicked the can to a page I know does not deliver, and called it a win. The button promises "Watch our company run, live" but the OOS page it links to is a list of published rules, not a running company.

L027 MEDIUM OBSERVED ONCE 3x efficiency

OTP's enemy statement is "you bought the operating system and the needle didn't move." The pitch is not better meetings; it is: the system was fine, what was missing was the workforce that runs it between the meetings. Frame all homepage/sales copy against needle-not-moving, not against meetings.

Why: This is the buyer's actual lived disappointment (paid for an operating system, company looks the same two years later) and it positions OTP against incumbents on outcomes instead of features.

Failure mode: The letter's hero framed the enemy as "the meeting" / busywork. David corrected the thesis: the real problem is that companies bought operating systems and software (Ninety, Bloom Growth, etc.) that did not move the needle. Years later the company had not grown and was not better, and they needed to change how they did things.

L029 MEDIUM OBSERVED ONCE 3x efficiency

For any UI change, design from the user's mental model, not the data model: "my list shows my work; work I assigned to others shows under Waiting on Others." When a meeting todo is assigned to someone else, stamp the creator as delegator so it routes to the delegation view. Before shipping UI changes, run the UX lens (impeccable / web-design-guidelines skills + src/DESIGN.md), not just a minimal code patch.

Why: A technically-correct patch that ignores the user's mental model just moves the confusion. OTP's own product language already has the right home for these items (Waiting on Others); fixes should land in the model the user already understands.

Failure mode: Fixed the dashboard todo confusion (teammates' meeting todos looked like the viewer's own) by adding an owner label to the rows. David corrected: that's not thinking like a user. Labeled-or-not, other people's todos don't belong in "my to-dos" at all.

L030 MEDIUM OBSERVED ONCE 3x efficiency

When David asks for a jaw-drop brand page, build an EXPERIENCE, not an article: full-viewport cinematic hero, scroll choreography, one idea per screen at massive scale, motifs that live in the page as motion, ruthless copy cuts, no standard nav/footer chrome breaking the spell, no section-grammar scaffolding, no FAQ accordion bolted onto a manifesto.

Why: The gap between "well-executed page" and "omg I love this" is the whole assignment on brand surfaces. Safe editorial structure is invisible at best; for-the-brave positioning demands the page itself be brave.

Failure mode: Built the /ollie manifesto page as a competent editorial layout (repeated mono eyebrow labels on every section, index rows, alternating light/dark sections, FAQ accordion at the bottom) and David rejected it outright: "this really really sucks." The brief was "reader drops on the ground saying omg I fucking love this" and the output was a safe template that reads as AI scaffolding.

L031 MEDIUM OBSERVED ONCE 3x efficiency

When David gives a design reference URL, open it in a browser and STUDY it visually (proportions, type sizes, spacing, alignment) before designing; match its register, not just its layout skeleton. Elegant means restrained: modest type scale, centered calm hierarchy, generous whitespace, thin rules. Never hand-draw SVG artwork to imitate produced brand art; crop/reuse the actual asset or use nothing.

Why: A reference URL is the brief. Reading its HTML structure without seeing it rendered led to importing the skeleton with the wrong soul, twice. Amateur freehand art next to professional motion work destroys credibility instantly.

Failure mode: Second rejection on the /ollie page. David asked for sakana.ai/fugu: elegant, Japanese sense of design (restraint, whitespace, calm, modest type, precision). I delivered giant 9vw headlines, one shouting line per viewport, and hand-drawn SVG chevron "birds" that rendered as crude fat marker scribbles. I treated "jaw-drop" as scale and boldness when the reference was quietness and precision, and I drew freehand SVG art instead of using the actual video's artwork.

L032 MEDIUM OBSERVED ONCE 3x efficiency

For fleet-wide spec maintenance: (1) tarball backup of ~/.claude before any agent touches specs; (2) partition files into DISJOINT clusters, one agent each, with CLAUDE.md owned by exactly one; (3) give every auditor the same stale-fact canon and the rule "verify a launchd plist exists before believing any schedule claim"; (4) auditors apply surgical edits directly for factual fixes but RETURN structural proposals for David instead of applying them; (5) synthesizer closes cross-cluster contradictions the auditors flag at each other.

Why: Agent specs rot faster than anyone audits them: this pass found live specs for a retired agent (jeff.md ending in "Go."), four phantom schedules, Todoist writes in five files, terminated employees still routed DMs, and a Bassim score-inflation bug. Periodic fleet audits with disjoint ownership are cheap insurance against agents acting on dead infrastructure.

Failure mode: SUCCESS: Claude ran a five-cluster parallel level-up of the entire agent army (80 files, ~140 surgical edits) without a single file conflict or lost spec.

L033 MEDIUM OBSERVED ONCE 3x efficiency

Pattern for UX dead-end hunts: (1) fan out parallel read-only explorers per surface (meetings, teams/members, KPIs/todos, onboarding/settings) asking for file:line + user-visible symptom + minimal fix; (2) fix the unsatisfiable states first: any required dropdown that can render zero options must explain where its options come from and link there (owners/attendees come from the org chart, meeting membership from teams); (3) empty states must branch on WHY they are empty (org has no teams vs user not on a team need different CTAs); (4) never report an async side effect as done: invite emails now await sendEmail (which returns null on failure, never throws) and return emailSent so the UI can tell the truth; (5) a guided setup checklist computed server-side from actual data (seats/team/KPI/meeting/members exist?) beats static onboarding because it survives skipped onboarding.

Why: These are the recurring shapes of broken UX in OTP: forms with prerequisites the user cannot see, empty states that misdiagnose their cause, and optimistic success messages over fire-and-forget side effects. Fixing the shape, not just the instance, is what makes the product feel intuitive.

Failure mode: SUCCESS: Claude ran a full UX dead-end audit and fix pass across OTP (4 PRs, #113-#116, all deployed)

L035 MEDIUM OBSERVED ONCE 3x efficiency

The sweep pattern that worked: audit by rule-cluster in parallel (fakery, insight-to-agency, jargon/states, first-meeting goal-walk), then execute severity-first. Key catches to re-check every run: (1) seeded/synthetic data leaking into numbers a reader believes are real (the is_template flag existed but was never enforced; counts now use src/shared/synthetic-orgs.ts); (2) the conversion moment must be ON the default path (end-meeting now lands on Ollie followups, not the list); (3) funnels don't exist until instrumented (insight topic: surfaced/accepted/value_delivered); (4) credentials in seed script comments (one prod DATABASE_URL scrubbed; password rotation still owed). Worklist for run 2 in otp-platform/mission-standard/WORKLIST.md.

Why: The Mission Standard is a repeatable bar, not a one-off audit. Recording the found failure classes makes run 2 start from run 1's ceiling instead of re-discovering it.

Failure mode: SUCCESS: Claude ran Mission Standard sweep run 1 (PRs #121-#123, deployed): 4 parallel rule-audits over OTP, 5 CRITICAL + 13 GAP found, all CRITICAL and 9 GAP closed same-session

L038 MEDIUM OBSERVED ONCE 3x efficiency

When an agent-pushed OTP to-do references a document, include a clickable https link (a Google Doc), not a local file/vault path — David reviews to-dos on mobile. To update an existing to-do's description, use PUT /api/v1/todos/:id (not PATCH). otp-todo.sh has no update verb, so PUT directly with the API key.

Why: A file path in a to-do is dead weight on mobile — the reviewer can see the reference but cannot open it, which reads as "the link is missing/broken." Every agent that pushes doc-linked to-dos (Radar, Pepper, Dan) hits this.

Failure mode: Dan pushed an OTP to-do referencing a document but put a local Obsidian vault path ("2nd Brain/Agent Army/Dan/...") in the description. David opens to-dos on his phone — a vault path is not tappable, so there was "no link to click." Also used PATCH to update the to-do; the OTP todos API update method is PUT /api/v1/todos/:id (PATCH hits the marketing site and returns HTML).

L039 MEDIUM OBSERVED ONCE 3x efficiency

To put an agent-army/IDS issue on an OTP meeting board, POST /api/v1/tickets with the team's teamId (category 'other' for strategic issues, priority low/medium/high/critical, ownerEntityType+ownerExternalId). The MCP submit_ticket tool CANNOT do this — it has no teamId param (it is the generic 'report a bug to OTP' path), which is why nobody ever got issues onto the board. Team IDs: 'AI Army' = 065d1d4b-c7da-4e80-b3ed-d6b101471d2c (the David+Dan agent-army meeting); Leadership Team = c1e1a485-414e-48d5-ae44-e81bd110b554. Update/solve via PUT /api/v1/tickets/:id (idsStatus, priorityRank, resolution).

Why: Agents could push KPIs and todos to OTP but not issues, so every L10 IDS board rendered empty and David kept discovering the hole live. Issues=tickets + teamId scoping is the missing piece; without it a meeting-readiness check would keep mislabeling a working API as absent.

Failure mode: Dan claimed 'no issues API exists in OTP' because there is no src/routes/api/issues.ts. That was wrong. OTP stores IDS issues in the TICKETS table (schema.ts: 'issues live in the tickets table'), with full IDS support (idsStatus, priorityRank, teamId, owner fields). The agent-army IDS board was empty only because our issues lived in a local markdown file and were never pushed as tickets scoped to a team.

L040 MEDIUM OBSERVED ONCE 3x efficiency

OTP has TWO distinct features both called "Ollie Insight": (A) the per-meeting followups wizard that turns a transcript into meetings.ai_summary (src/shared/meeting-followups.ts, transcript-only), and (B) the reusable "address engine" (ollie_insights table, src/services/ollie-insight.ts + shared/ollie-insight.ts + partials/ollie-insight-block.ejs) that gathers org data (KPIs/rocks/todos/meeting-summaries) per SCOPE. When David said "KPIs shouldn't be in the meeting analysis," the fix was in System B's meeting-scope evidence gathering, NOT System A. The block partial (ollie-insight-block.ejs) is fully scope-generic (builds the API URL from data-oib-scope/scopeId client-side), so adding a brand-new 'quarter' scope end-to-end took only: add to INSIGHT_SCOPES + RULES_BY_SCOPE (shared), the scopeQuerySchema enum + a resolveInsightScope branch (api), a gatherEvidence branch + max_tokens (service) -- then just include the existing partial with scope:'quarter'. No new render/generate/receipts UI. Pattern: when a feature is "one engine, many surfaces," new surfaces are a scope + evidence branch, never new UI.

Why: The name collision hides which code to touch; picking the wrong system wastes a whole edit pass. And recognizing the scope-generic block means big-feeling asks ("a quarterly synthesis button") are small, low-risk diffs. Both are recurring shapes in OTP's Ollie work.

Failure mode: SUCCESS: Claude separated OTP's two "Ollie Insight" systems and added a whole new scope by reuse

L041 MEDIUM OBSERVED ONCE 3x efficiency

New endpoint POST /api/v1/meetings/:id/agent-record (PR #154): an agent submits the written meeting record; OTP runs the same redaction ruleset, stores it to meetings.transcript, logs an audit baseline + agent_record event, and the existing /ai/followups generate turns it into to-dos/issues/headlines/insight unchanged. Agent path: ~/.claude/otp-meeting.sh record <meetingId> --file=<record> --source=l10dan. Wired into /l10dan conclude step 6. Verify deploy by probing the endpoint returns JSON not the marketing SPA HTML before pushing records.

Why: Ollie only reads transcripts, so agent-run meetings had no way into it — the empty-insight hole David hit live. This closes it: any agent meeting can now produce Ollie follow-ups. Also a general UI rule captured: buttons reflect capability/state (action taken -> button disappears).

Failure mode: SUCCESS: Dan shipped the Ollie agent-record path so agent-facilitated meetings (David + AI L10, no audio transcript) can feed Ollie Insights.

L042 MEDIUM OBSERVED ONCE 3x efficiency

Before treating a KPI/data source as blocked, re-read the LIVE source, not the note about it. The Havok "non-client %" KPI was marked blocked for ~14 weeks on "the timesheet has no client column" — a 105-day-old memory. The live sheet (1VPlH5ZqTowOe2nJDFwOjuvXrCXUnOzeOb3aO-M936Xo) had since grown per-person tabs with a full Client/Time/Date schema; one read unblocked it. Pattern for wiring a messy sheet into Tally: (1) get_spreadsheet_info to list tabs, (2) read a person tab to learn the real schema, (3) confirm recency by reading the tail (last date), (4) add a focused extract mode to tally.py rather than reshaping the sheet — here `client_attribution_since` (reads ALL valueRanges via a new _values_2d_all, parses h:mm via _parse_hhmm, added dotted DD.MM.YYYY to _parse_date, classifies internal by an `internal_contains` substring), (5) `tally.py --dry-run --kpi "<title>"` to prove the number off live data before pushing. Human-owner + a 1:1-team KPI (owner HUM_BOGDANTABAKA, team = David-Bogdan 1:1) pushes fine via find_or_create_kpi.

Why: Blocked-status notes rot silently while the underlying source improves; a KPI can sit "pending" for a quarter when it was buildable weeks ago. Re-reading the live source first is the cheap unblock. And the tally.py extract-mode pattern makes any timesheet/sheet a live KPI without asking a human to restructure their doc.

Failure mode: SUCCESS: Dan/Tally shipped the Havok client-attribution KPI live in one session after a 14-week "blocked" note turned out stale

L044 MEDIUM OBSERVED ONCE 3x efficiency

Before editing any OTP view to fix an on-screen bug, grep unique visible strings from the screenshot (e.g. "WAITING ON OTHERS", "always only yours") across src/views to confirm WHICH template renders that exact surface. Multiple pages can render similar-looking todo lists (me-todos.ejs vs dashboard-daily.ejs). Verify the rendering route (reply.view target) too.

Why: Two round-trips and two merged PRs produced zero visible change because the edits were on the wrong template, which read as "nothing is fixed" and eroded trust. A 10-second grep on the screenshot text would have pointed to the right file immediately.

Failure mode: Fixing an OTP todos UI bug, I edited src/views/pages/me-todos.ejs twice and shipped two PRs, but the surface David actually uses is the dashboard-daily "Waiting on others" widget (src/views/pages/dashboard-daily.ejs). Nothing he saw changed.

L045 MEDIUM OBSERVED ONCE 3x efficiency

Two reusable patterns: (1) Before building any OTP email/engagement feature, grep src/services for existing infrastructure -- re-engagement.ts, lifecycle-scheduler.ts, and user_engagement_log already carried cadence caps, suppression, logging, and a daily cron, so the todo-aware upgrade was ~350 lines instead of a new subsystem. (2) When resolving "which org does this Clerk user belong to", organizations.clerkOrgId only knows the org CREATOR; invited teammates must be resolved through org_members.clerkUserId + claimedEntityIds. This gap is why per-user personalization (open todos) missed non-creator members like Nate.

Why: One engagement channel with shared caps is what keeps daily utilization pressure from becoming annoying double-mailing, and the creator-vs-member resolution gap will bite any future per-user feature (digests, notifications, billing seats) that starts from organizations.clerkOrgId.

Failure mode: SUCCESS: Claude shipped the smart engagement email engine (PR #186) by upgrading the existing re-engagement service instead of building a parallel system

L048 MEDIUM OBSERVED ONCE 3x efficiency

Frame help/success/onboarding call copy positively: state plainly that the call is there to help and guide them, and describe what will actually happen on it (we'll walk through your setup, get you unstuck, answer your questions). Never say "this is not a sales call" or "no pitch" — describe the help, don't disclaim the sell.

Why: Defensive "not a sales" language triggers the exact suspicion it tries to defuse and undercuts a genuine help offer. David flagged this immediately.

Failure mode: Wrote a "customer success call" Calendly description that leaned on "no pitch, no slides" / not-a-sales-call framing. Protesting that it isn't a sales call makes it sound like one.

L049 MEDIUM OBSERVED ONCE 3x efficiency

Any script that answers "is anything missing / is everything covered?" must fail LOUD, never return an empty set as reassurance. Two rules: (1) assert the expected top-level key exists (`if 'data' not in resp: raise`) before computing a result; a zero/empty answer from a health check is a claim that must be proven, not a default. (2) Cross-check a zero result against one known-positive case before reporting it -- here, one direct call to a single account would have shown $57 of spend and exposed the lie instantly. Also: `~/.claude/meta-ads.sh accounts` exits 0 and prints nothing; do not build on it. Sweep with `/{business_id}/adaccounts?fields=name,account_status&limit=500` then batch `/{act_id}/insights` 50 at a time.

Why: "Nothing is wrong" is the single most dangerous output an audit can produce, because nobody investigates it. A silent empty result on a billing sweep means real revenue is never invoiced and nobody ever finds out. The failure mode is not a crash, it is confident silence.

Failure mode: SUCCESS: Claude caught a silent false-negative in a Meta billing sweep. Querying the Graph API adaccounts edge with nested field expansion (`fields=name,insights.date_preset(this_month){spend}`) returned an error payload with NO `data` key. The sweep script read it as an empty account list and confidently reported "0 accounts, $0.00 unbilled MTD spend" -- a clean bill of health that was entirely fabricated. Direct per-account queries then revealed 7 unbilled accounts spending $5,673 MTD.

L050 MEDIUM OBSERVED ONCE 3x efficiency

When you own a strong long-form asset (a landing page that argues the WHY), the cold email must NOT re-argue it in miniature. Cut the product explanation entirely. The email's only job is to earn one click: proof you read their work, one line naming THEIR problem in THEIR language, the link, out. No feature, no price, no call ask. Match the page to the reader's specific pain rather than sending everyone to the same URL. And assign/log an A/B arm on every send, because an untested signature shipping at 100% for months is not a decision, it is a habit.

Why: A cold email is the worst possible venue for explaining a product: no trust, no attention, no context. A page is the best. Using the email to sell the READ instead of the PRODUCT plays each asset to its strength. The deeper failure was cheaper: months of sends with no arm logged and no reply column filled, which means the volume produced zero learning. Volume without measurement is just noise you paid for.

Failure mode: SUCCESS: Crafter cold-email rework -- the email body was trying to explain the product to a stranger in three sentences, while the website spent full pages arguing the why. The two assets contradicted each other and the email was losing. Outreach log had been dead since Apr 26 with effectively no replies.

L053 MEDIUM OBSERVED ONCE 3x efficiency

When compiling guru influences into an Ollie persona, push voice fidelity hard: the dominant influence's cadence, vocabulary, and signature moves should LEAD the writing, not decorate it. Channel the style ("in the room, guiding in their voice") while keeping the hard never-impersonate line: never claim to BE the person, never claim endorsement. Also: differentiation surfaces (the Lab) must let users edit each guru's forked principles inline and must show WHERE a principle or voice move landed in the output (highlights on the insight, influence tags on todos/issues/headlines) so voices can actually be compared.

Why: The entire value of per-team Ollie voices is that the voice audibly changes what the team hears. If influences only shift content, not voice, the differentiation is inaudible and the Lab proves nothing. Attribution marks are what make the difference legible.

Failure mode: Ollie Lab guru influences rendered as principles-with-a-hint-of-style: the output read like Ollie citing a thinker, not like the thinker's voice guiding the room. David: "it should be as if they were speaking with their voice guiding us in their voice... It should be like they are in the room."

L054 MEDIUM OBSERVED ONCE 3x efficiency

Voice fidelity lives in structure, never slogans. Ban catchphrase quoting explicitly in the persona compile ("never quote their slogans; that is imitation's cheapest form"). Give each library guru a hand-written VOICE DNA block: sentence rhythm and length, how they open, how they build an argument, what they notice first, how they land a point, emotional register, what they never do. For custom gurus, instruct the model to reconstruct the person's published voice from structure and register, not taglines. Method instruction: "before writing, ask how NAME would structure this and what they would notice first; write from inside that mind."

Why: Catchphrases signal imitation and break trust instantly; structural voice makes the reader feel the thinker in the room without a single borrowed phrase. This is the difference between a costume and a mind, and it is the entire premium of per-team Ollie voices.

Failure mode: Ollie's guru voice channeling produced surface mimicry: dropping the thinker's catchphrases ("Start with why") instead of writing from inside their rhetorical DNA. David: "not cheap tricks... it needs to be in the DNA of what Ollie is saying. Dig deep, make it real."

L055 MEDIUM OBSERVED ONCE 3x efficiency

Essay-page hero art is a first-class deliverable, not a wireframe: match the family's craft bar (a real scene that tells the page's story, Ollie present, warm accent treatment, gradients/glow, staged ignition-style animation, reduced-motion complete state). Before drawing, read the sibling page's actual SVG to absorb its techniques, then design a scene, not a diagram.

Why: The hero art IS the argument at first glance on these pages; the candle page's flames are the thesis made visible. A schematic undercuts an inspiration page precisely where it must inspire.

Failure mode: The /the-voice-in-the-room hero art shipped as a lazy schematic (dot, four thin lines, grey boxes with circles for people) while its sibling pages (/one-candle, /mission-to-the-moon) carry hand-crafted narrative scenes with the mascot, gradients, glows, and staged animation. David: "you got lazy there for sure."

L058 MEDIUM OBSERVED ONCE 3x efficiency

When running counterfactual/what-if simulations, make the events endogenous: model causal mechanisms (innovation rate, demographics, conflict outcomes) as functions of the changed variable and run Monte Carlo over branching timelines, rather than overlaying new participation rates on the fixed historical record. Show a distribution of divergent timelines, not one re-skinned version of real history.

Why: Fixed-event counterfactuals smuggle in the answer (the world converges because you forced it to). Branching simulation is what actually answers "would things be different" questions, and it's the difference between a re-labeled chart and genuine out-of-the-box analysis.

Failure mode: Built a counterfactual history simulation that held all real-world events fixed (same Industrial Revolution, same wars, same tech timeline) and only varied participation rates within them. David flagged that this assumes the conclusion: if the labor/care allocation changes, the events themselves change — maybe industrialization comes later (or earlier), wars resolve differently, the whole timeline branches.

L059 MEDIUM OBSERVED ONCE 3x efficiency

In live L10/Delta Meeting facilitation, sequence is: surface ALL signals first (grouped, neutral, no recommendation attached), let David react and pick what to IDS, and only then frame decisions. Decisions come after shared context, not before. Also: open every meeting with Ollie's read of the PRIOR meeting's record (pulled from the OTP followups/insight), as a standing first section — David explicitly values it.

Why: A decision framed before the signal review railroads the meeting toward Dan's framing and skips the part where David's pattern-recognition works on the raw material. The facilitator's job is to lay out the board, not to compress it into a pre-picked fork. And the Ollie prior-meeting read is the continuity loop that makes each meeting compound on the last.

Failure mode: Dan facilitated the live L10 by pushing straight to the rock-set timing decision immediately after the scorecard, without first walking through all the signals (client wins, CC trends, attribution reading, unverified tiles, churn signals) so David could see the whole board. David: "we need to review all the signals before moving so quickly, you are trying to skip ahead too fast."

L060 MEDIUM OBSERVED ONCE 3x efficiency

L10 prep MUST include a live scan of OTP itself: the team's rocks/priorities board, the issues (tickets) board, todos, and KPIs, pulled fresh at prep time. OTP is the source of truth; local files are mirrors that go stale the moment David works in the product directly (which is the whole point of OTP). Additionally, per David's 7/13 ruling: local shared-state files are now formally HISTORICAL unless a live consumer reads them; staleness flags on superseded files are noise, but staleness in the OTP scan is a real miss.

Why: David increasingly works inside OTP directly (weekend rock-setting), so any prep that skips the live product misses his most recent decisions and re-surfaces solved or stale items. The mirror-drift failure has now happened twice (June 15, July 13); the fix is structural, scan the product, not the mirror.

Failure mode: Dan's L10 prep read local mirror files and this morning's Tally/KPI pipeline but never scanned the live OTP boards (rocks/priorities and their attached issues) before the meeting. Result: Q3 rocks David added over the weekend were missing from the prep, and stale issues sitting on the rocks board went unnoticed. David caught it live, again (first time was June 15).

L061 MEDIUM OBSERVED ONCE 3x efficiency

Every KPI on any board must pass the needle test before it earns a tile: one sentence stating the causal chain from this number to the company goal (margin, retention, revenue, rock completion). Dan owns running this test — on every existing tile quarterly and on every proposed tile before creation. Activity metrics (emails drafted, projects counted, pushes made) are health checks at best; they live in the readiness script, not on the scorecard. Wiring a dead tile is worthless if the tile measures the wrong thing.

Why: The strategic co-founder seat exists to hold the big picture David cannot hold while operating. A perfectly-wired scorecard of needle-irrelevant numbers is worse than an empty one, because it manufactures the feeling of accountability without the substance. This is the second-order version of "every seat owns a number": every number must own a reason.

Failure mode: Dan treated the scorecard as a plumbing problem (are tiles wired, do values push) instead of a strategy problem (does each KPI move the needle toward the goal). David: "we/I make all of these changes thinking you are looking at the big picture, that does not seem to be the case. Your job is to make sure that we reach our goal, and the KPIs should support that. How many emails Pepper reads does not move the needle. Each KPI needs to answer how it moves the needle, and how."

L062 MEDIUM OBSERVED ONCE 3x efficiency

L10 prep must WALK THE ACTUAL MEETING before the meeting: open/fetch the exact meeting David will see (via API), verify every section renders with real data (scorecard snapshot has values, rocks board current across ALL teams incl. corporate, issues/todos loaded), run the needle-test on every tile, and FIX or stage fixes for everything found - all before 8am. The brief reports what was already repaired, not what will be discovered. Prep = simulate the meeting end-to-end; the meeting itself is only for decisions the human must make.

Why: A meeting that debugs itself live burns the scarcest resource (David's attention) on work an agent could have done at 7am. "The work happens between the meetings" is the entire operating philosophy of the meeting cadence - prep that only compiles data without verifying the meeting surfaces is half a prep. This is the root cause behind L059/L060/L061; fixing it structurally prevents all three recurring.

Failure mode: The 7/13 L10 scored 4/10. David's reason: prep ran in the morning but the meeting still spent most of its time discovering and fixing things live (weekend rocks missed, empty scorecard render, dead tiles, needle-less KPIs, corporate rocks invisible) - "we are fixing the meeting within the meeting with an absence of information. The work happens BETWEEN the meetings and this is not the case here."

L064 MEDIUM OBSERVED ONCE 3x efficiency

When auditing whether an invariant holds across a codebase, verify it per STATEMENT, not per file. A per-file grep count is an aggregate, and aggregates hide the exact case you are hunting: the mixed file. Then encode the audit as a test that scans every call site, and mutation-test that scanner by reintroducing a known bug to confirm it actually fails. A scanner that only agrees with current code proves nothing. Related pattern seen the same day: when a data model gains a concept (rock levels, agent-owned KPIs), the model and the dashboard get wired up and the OTHER surfaces silently do not. Ask which surfaces read this table, not just which one is broken.

Why: The three holes the per-file count missed included the worst one: the blueprint serializer baked a private rock into a shareable template, which would carry it into another account. The failure mode of an aggregate check is a false clean bill of health, which is more dangerous than no check at all because it stops the search.

Failure mode: SUCCESS: Dan audited a privacy invariant (shadow rocks are owner-only) and found 7 holes, but the FIRST pass counted guards per FILE and missed 3 of them, because a file can contain one guarded query and one unguarded query and still look guarded in aggregate.

L065 MEDIUM OBSERVED ONCE 3x efficiency

Invert it. Give the user ONE copy-paste block (MCP connection + a self-registration prompt) that they drop into Claude Code, Claude Desktop, ChatGPT, Cursor, or any MCP client. The AGENT then connects to OTP and registers ITSELF: it reads its own system prompt, calls a register/enroll MCP tool with its name, role, what it owns, what it does not own, and its KPIs, and OTP creates the seat and KPIs automatically. The user's only job is copy, paste, done. No forms, no parsing, no filling anything in.

Why: The agent already knows what it is -- making a human retype it is redundant work and a competence gate. Any flow that requires the user to know how to describe or configure their agent is not idiot-proof and will lose the non-technical user. The agent is the most reliable source of truth about itself, and it is already sitting on the other end of the MCP connection, so let it do the work. Design rule: when an AI is on the other side of the pipe, push the setup work to the AI, not the human.

Failure mode: Built OTP's "Connect an agent" flow as a human-driven form: the user pastes their CLAUDE.md into a textarea, OTP parses it, and the user hand-fills name / role / owns / does-not-own / KPIs before a seat is created. It made onboarding an existing agent the USER's clerical job and assumed the user knows how to describe their own agent.

L066 MEDIUM OBSERVED ONCE 3x efficiency

Before any L10, dry-run tally.py and for every regex_in_file KPI check the source file's mtime against the KPI's time grain; if older than one grain, re-pull the number from the live system (Accelo, Search Atlas, Sheets) before pushing. New registry entries must use kind/regex/group (never type/pattern) and carry pending:true until their emit line exists. The unknown-kind branch now honors pending.

Why: A KPI pushed from a stale mirror is worse than a missing one: Crystal's tile would have said 32 when reality is 44, and the failure alert noise from mis-schema'd pending entries erodes trust in the one agent whose whole job is keeping the scorecard honest. Live-source-first is the same lesson as L042/L060 applied to Tally's own pipeline.

Failure mode: SUCCESS: Tally pre-L10 KPI sweep found and fixed three silent scorecard rot points: (1) registry entries added at the 7/13 L10 used type/pattern keys but the runner requires kind/regex, so their pending:true flag was never honored and they reported as failures; (2) the runner's unknown-source-kind branch ignored pending entirely; (3) two regex_in_file KPIs (Crystal 32, Beacon 0) were feeding from stale files (Jun 8 and Jun 22) while the live sources (Accelo: 44 projects; Search Atlas: 8 keywords tracked, not 43) had moved.

L067 MEDIUM OBSERVED ONCE 3x efficiency

Any browser feature that accumulates unrecoverable state in page memory (MediaRecorder audio, unsent drafts) must (1) block/defer every programmatic self-reload while active, (2) be flushed by navigation-triggering handlers (End meeting awaits OTPAudioRecord.finish() before navigating), (3) guard beforeunload, (4) never drop data on a failed upload — keep the blob and offer retry. Long-term: stream chunks to the server (timeslice) so the page is never the only copy.

Why: A full Delta Meeting's recording/transcript was permanently lost — the audio never left the browser. Guard rails shipped in audio-record.ejs + l8-leadership.ejs; chunked streaming upload is the queued follow-up (touches billing-metered meeting-audio.ts, needs billing lock).

Failure mode: SUCCESS: Conatus — root-caused OTP meeting recording loss (2026-07-16): the browser recorder holds all audio in tab memory until Stop, while the live meeting page self-reloads on routine actions (reloadKeep, SSE scheduleReload, section-refresh fallbacks) and End Meeting navigates away — any of these silently killed a live MediaRecorder with zero warning.

L068 MEDIUM OBSERVED ONCE 3x efficiency

Durable pattern for any browser feature holding unrecoverable state: (1) layer the fix — guard rails first (cheap, same-day, stops the active bleeding), durable streaming/persistence second; (2) keep the old one-shot path untouched as an automatic fallback so degradation can never be worse than before; (3) money-path parity — finalize calls the exact same precheck/ingest/charge sequence as the one-shot path, charge only after successful transcribe+ingest; (4) hold the billing lock across the whole build, release only after merge; (5) verify with a harness that runs the REAL shipped script (44 assertions) so old behavior is provably byte-identical with the new feature off; (6) cap every attacker-spinnable counter (segment count was a finalize-loop DoS lever).

Why: Completes L067's open loop: the "never lose a recording" guarantee is now structural, not procedural. The layered-fix + fallback-preserving pattern is reusable for every OTP feature that buffers user work in the browser (draft notes, offline edits), and the billing-parity discipline is how streaming touched the money path with zero semantic change. Both PRs merged and confirmed live on orgtp.com 2026-07-16 evening.

Failure mode: SUCCESS: Conatus — full resolution of the 2026-07-16 meeting-recording loss, shipped to prod same day in two layers: PR #205 guard rails (meeting page defers all self-reloads while recording; End Meeting flushes the recorder before navigating; beforeunload guard; pinned REC pill; failed uploads keep the blob with retry) and PR #206 streaming upload (recorder streams ~10s chunks to a server recording session; crash/reload loses ≤10s; resume banner stitches segments into one transcript; phone QR flow covered).

L070 MEDIUM OBSERVED ONCE 3x efficiency

When invoking Steve Jobs as a design standard, treat him as holding BOTH axes to one bar: visual craft (typography, proportion, detail) and end-to-end experience (defaults, subtraction of steps, invisible mechanism). Never frame him as the "UX half" opposite a visual system; frame the visual system as one half of the single Jobs-level standard.

Why: David's design north star for OTP is the full Jobs standard. Splitting it wrongly would let screens pass a visual checklist while the flow, or the craft, gets held to a lower bar. The correct frame keeps one bar over both layers.

Failure mode: When framing the design-standard marriage (Fugu + Jobs), Conatus split it as "Fugu = visual craft, Jobs = journey/UX", understating that Jobs was also a master of visual design (Reed calligraphy class, typography on the original Mac interface).

L071 MEDIUM OBSERVED ONCE 3x efficiency

For /coach-report and any Dash run: do MCP-dependent pulls (Search Atlas, Google Sheets, Calendar) in the MAIN session; delegate only file/CLI/analysis work to subagents. Check a subagent's tool access assumption before waiting on it. Treat CCM STL as unusable until the CloudCRM timezone offset is fixed platform-wide, and never report STL from July data.

Why: Two full delegation rounds were wasted waiting on pullers that could never succeed; the report would have shipped without SEO (repeat of L057) and without CCM if the main session had not redone the pulls. The STL corruption finding upgrades the known Villa-only issue (June) to system-wide, which changes every STL-based alert and coaching metric until fixed.

Failure mode: SUCCESS (with lesson): Dash /coach-report 2026-07-19 — delegated Search Atlas, CCM sheet, and calendar pulls to three subagents; SEO and CCM pullers were fully blocked because spawned subagents do NOT inherit the session's MCP servers (search-atlas and google-workspace tools were absent from their toolsets). Main session had both and pulled everything directly. Also: CCM July speed-to-lead timestamps are corrupt SYSTEM-WIDE (large negative timezone artifacts on every project, not just Villa Sport).

L072 MEDIUM OBSERVED ONCE 3x efficiency

Two mechanics to remember: (1) .gitleaksignore fingerprints are commit-hash-bound, so ANY commit touching a line with secret-shaped placeholder text (like 'Bearer YOUR_API_KEY' in docs) re-mints the fingerprint and re-triggers the scanner — squash merges guarantee this recurs. The durable fix is neutralizing the placeholder so the rule can't match (angle brackets: 'Bearer <your-api-key>'), plus fingerprinting the immutable history. (2) CI checkouts with fetch-depth:0 fetch ALL refs and gitleaks scans all of them — one bad commit on an unmerged branch fails every branch's CI simultaneously; the ignore entry must reach each scanning checkout's .gitleaksignore, which means pushing it to the branch being scanned AND to main.

Why: Symptom (lint-and-type-check job failing everywhere at once) looks like a code regression but is actually the secret scanner; without knowing the two mechanics, the obvious fix (add one fingerprint) only patches one branch and the mole pops up on the next touch of the file.

Failure mode: SUCCESS: Conatus diagnosed a repo-wide CI outage caused by gitleaks fingerprint whack-a-mole — every branch's CI (including main pushes) went red at once from ONE unmerged branch's commit.

L080 MEDIUM OBSERVED ONCE 3x efficiency

Before a swamp send, reconcile the changelog against the week's real PRs (git log origin/main --since since the last issue, filter to feat/ and customer-facing) and write entries for anything unlogged -- do not assume changelog.ts is complete. When the shared repo is contested by a concurrent session, do all changelog authoring in an ISOLATED git worktree (git worktree add off origin/main, symlink node_modules to reuse deps), PR it, merge via gh after CI is green, and run the REAL send from the worktree -- never edit or send from the shared working tree. Date drop-wave entries to the issue's MONDAY anchor (the sender's window upper bound), NOT to the send day: entries dated the Tuesday send day render as future and hide, and getRecentEntries (OS-today) masks this in preflight.

Why: The changelog is the single source of truth for both /whats-new and the email; an unmaintained changelog silently undersells the product to every subscriber. On a machine with concurrent agent sessions, the shared working tree is not safe for a multi-step author+send; a worktree makes the work deterministic and collision-proof. The Monday-anchor date rule is invisible until a dated-Tuesday entry silently disappears from the send.

Failure mode: SUCCESS: Swamp #29 -- two non-obvious operational wins. (1) The digest only reflects changelog.ts, so a big shipping week (~14 features) went out as "2 things" because most PRs never got changelog entries; David caught it. (2) Running the swamp send while another session was actively branch-switching in ~/otp-platform stranded an early commit on a feature branch and made a direct push to protected main a no-op.

L082 MEDIUM OBSERVED ONCE 3x efficiency

For any EJS page whose logic lives in an inline script, add a test that parses every inline script with new Function() and fails the build on a syntax error, then prove the tripwire by running it against the broken version before keeping it. tsc cannot see inside a template and EJS renders a broken string happily.

Why: It was caught only by driving the actual rendered page in a headless browser, not by tsc, lint, or 1204 passing tests. First-run screens are where paying customers land, and a dead one is invisible from the server side.

Failure mode: SUCCESS: Conatus found OTP onboarding Door 4 (Give Ollie everything, PR #232, shipped 2026-07-19) had been completely dead in production since Sunday. onboarding-import.ejs line 214 had an apostrophe inside a single-quoted JS string, so the browser discarded the whole inline script and the file drop, analyze and commit buttons did nothing, silently, for every new customer who picked that door.

L083 MEDIUM OBSERVED ONCE 3x efficiency

David: A2P pages do not allow form fills. Remove the lead form from A2P review landing pages entirely. Keep the SMS disclosure, the Privacy Policy and Terms links, and the registered business name and address on the page, because those are what carriers actually read. Route the quote path to phone plus the GHL chat widget. Reword any disclosure copy referencing "check the box above" so the page does not describe a mechanism that no longer exists.

Why: The form was the architecture of these pages, so this is not cosmetic: it orphans the consent audit trail and the POST /:slug/lead route, and it changes what the A2P campaign registration can declare as its opt-in method. Getting it wrong in either direction risks 10DLC rejection, which blocks SMS for the client entirely. The pattern repeats for every future A2P location page.

Failure mode: Built the Dryer Vent Squad A2P landing pages (Katy, DFW) around a web lead form with an optional SMS consent checkbox, treating that form as the opt-in proof mechanism for A2P/10DLC review, plus a consent audit trail behind it (consent.js and sms_consent_text/timestamp/IP/user-agent captured on POST /:slug/lead).

L084 MEDIUM OBSERVED ONCE 3x efficiency

The unbilled-spend sweep (billing-report step 3b) reads ~/.claude/billing/sweep-exclusions.json and drops those account IDs from the Review tab entirely. Use it for accounts that spend on our Google/Meta but are NOT billed on % of ad spend (white-label / flat-fee). Added 2026-07-23 per David: Phillip Jeffries, M.V. Parker Law, Champy's Chicken (+Nashville), Emily Shalant, Jet City Blinds, J&K Engines, Meyer Law, True Path, GettaMeeting, Lazzara Law, Studstill Firm. Only add IDs here on David's explicit instruction. Separate from the Clients-tab dont_bill mode (which still shows a DO NOT BILL row on the Billing output).

Why: Without a persistent list these 12 accounts (~$24.4K/mo) resurface as "confirm arrangement" in every monthly sweep, wasting David's review time. The file makes the exclusion durable across sessions.

Failure mode: SUCCESS: Billing sweep now has a persistent white-label exclusion list

L089 MEDIUM OBSERVED ONCE 3x efficiency

Three patterns for the OTP recorder. (1) "No way to record a second time" is TWO bugs: the missing UI affordance AND a server ingest that overwrites (ingestTranscriptOneShot did set({transcript})). Fix both or the feature silently destroys data; recording ingests now pass mode:'append', idempotent on the tail so worker retries cannot double-append, while paste/import stays 'replace'. (2) "Mic recorded silence after a long pause" on a phone is a dead MediaStreamTrack (readyState 'ended' or muted): MediaRecorder.resume() succeeds and records nothing. Recover by swapping in a fresh getUserMedia stream, and ALWAYS flush the retired recorder's final chunk BEFORE bumping the server segment number, or the old container's tail bytes land in the new segment and corrupt it. Auto-recover only when the track is provably dead; when it is alive but silent, warn and offer a button, since a genuinely quiet room looks identical. (3) Check git log before rebuilding from a support ticket: background transcription had already shipped that morning, so ticket 3 needed a sequential-to-concurrent R2 upload fix plus a crash fix, not a rebuild.

Why: The recorder holds the customer's only copy of a meeting until upload completes, so every bug in it costs unrecoverable audio, and two shipped in one day. The widget is inline JS in an EJS partial with no import path, which is why it went untested; it can now be tested by rendering the partial, extracting the script, and running it in a vm against a fake MediaRecorder (src/views/partials/meeting/audio-record.test.ts). Use that harness for future recorder changes.

Failure mode: SUCCESS: Claude (OTP dev) fixed three meeting-recording tickets (PR #315) at root cause, and caught that PR #309 from earlier the same day left savedStatus() calling itself on the non-pending branch (a stack overflow that swallowed the save confirmation on the R2-off path) because the recorder widget's inline JS had no tests.

L090 MEDIUM OBSERVED ONCE 3x efficiency

When editing OTP trust/security claims, edit src/config/trust.ts (the file the /trust route imports and ships in the dist image). trust.yaml at the repo root is only the audit copy carrying `# source:` code citations for legal; nothing reads it at runtime. The two had already drifted (legalEntity OTP,LLC vs OrgTP,LLC; a stale lastUpdate; and — critically — trust.ts shipped a prohibited EOS mark "L10" that trust.yaml did not). Always mirror any claim change into BOTH files, and treat trust.ts as authoritative for what the public actually sees.

Why: A trademark-compliance violation (EOS "L10" mark) was live on the lawyer-facing trust page for weeks because the de-EOS pass only fixed trust.yaml, which never ships. Editing the audit copy feels like fixing the page but changes nothing a visitor sees. This is a recurring drift trap worth a permanent CI equivalence check between the two files.

Failure mode: SUCCESS: the public /trust page renders src/config/trust.ts, NOT trust.yaml — the "source of truth" file never loads at runtime

L091 MEDIUM OBSERVED ONCE 3x efficiency

Never position OTP (or Ollie) as an employee. A founder, and even an employee, does not want another employee; they want to know the organization is held. "Employee" also imports employee mentality: waits to be told, owns a lane not the whole. OTP guides the organization. Frame it as the guide/what holds the company, positioned between employee (too small), guru (too big), and co-founder (not really).

Why: Positioning error at the core-promise level: hiring an employee increases a founder's load (managing, explaining, checking); OTP's promise must reduce the holding. Wrong noun poisons every downstream page, price, and demo.

Failure mode: Framed OTP's destination as "the first employee a company hires whose job is to remember" in the product discovery brief.

L093 MEDIUM OBSERVED ONCE 3x efficiency

When a vendor was chosen specifically to absorb an operational burden, do not propose solutions that hand that burden back, even when the vendor documents them. Default order for the consent-screen concern: (1) own the story in product copy (tell users they will see Composio, frame it as the security vault), (2) at most white-label the two or three marquee providers if branding ever matters commercially, (3) never the whole catalog.

Why: Vendor docs happily describe features that shift work back onto the customer. The right frame is the original build-vs-buy decision: Composio IS the OAuth team. Recommendations that quietly re-hire that job in-house waste the subscription and David's time.

Failure mode: Asked how to get OTP branding on the Composio OAuth consent screen, Claude recommended white-labeling via custom auth configs, which means creating and maintaining an OTP-owned OAuth app per provider (Slack app review, Google verification, secret rotation, scope upkeep). David rejected it: the reason OTP uses Composio at all is ONE managed OAuth surface across hundreds of products; a per-provider OAuth app pipeline recreates the exact burden Composio was chosen to eliminate.

L097 MEDIUM OBSERVED ONCE 3x efficiency

The company boundary applies to every artifact that feeds Ollie, not just what is said in the room. When writing a meeting record for the Sneeze It L10, describe mechanisms in company-neutral operational terms (the meeting pipeline, the prep gate, the record path) and keep OTP product identifiers (PR numbers, endpoints, feature ship dates) out of the record entirely. Before pushing any record via agent-record, scan it with the same company-mismatch check the preflight applies to the board.

Why: Ollie's insight renders inside the Sneeze It meeting, so a record contaminated with OTP content produces a contaminated insight automatically, one week later, with no human in the loop. The boundary check must move upstream to where the source is authored or the violation recurs on autopilot.

Failure mode: Dan wrote the 7/20 agent-record for the Sneeze It L10 full of OTP product identifiers (PR #154, the agent-record endpoint, ship dates), so the Ollie Insight generated from it reads as OTP product narrative inside the Sneeze It meeting. David caught it live on 7/27: "Ollie still thinks OTP and Sneeze It are one." The context-bleed boundary was enforced in live speech but not at record-writing time, and the record is the insight's source.

L098 MEDIUM OBSERVED ONCE 3x efficiency

Prep does not end at the Slack brief. Any signal the brief nominates for IDS gets pushed to the OTP board as a ticket (otp-issue.sh, with teamId, AI Army = 065d1d4b) BEFORE the meeting, so David opens the meeting with the issues already loaded and workable in-product. The Slack brief is the narrative; the board is the working surface.

Why: The meeting runs inside OTP, so an issue that exists only in Slack is invisible at the moment of solving. This is the same source-of-truth lesson as L060 applied to prep outputs, not just prep inputs.

Failure mode: Dan surfaced pre-meeting signals (Pulse dark, pipeline shape, HiTone, overdue queue items) only in the Slack prep brief. David at the 7/27 L10: "you should have wrote those signals in OTP, too late now." Signals that deserve IDS never landed as tickets on the AI Army board, so the meeting could not work them in-product.

L099 MEDIUM OBSERVED ONCE 3x efficiency

When David closes an alert as handled, retire the RULE that generates it, not just the instance. For HiTone: edit the CLAUDE.md trigger line and coach-report spec to read "billing confirmed active since Jul 2026, do not flag on spend." In general: trace any recurring alert to the config line that emits it and fix that line, or the alert regenerates forever.

Why: A closed instance with a live trigger is an alert factory. It spends David's attention on the same resolved question weekly and erodes trust in real billing flags.

Failure mode: HiTone billing keeps re-surfacing to David even though he confirmed billing is correct and active (Jun 29, and again 7/27: "HiTone is being billed, you ask about that a lot"). Root cause: the BILLING TRIGGER rule still lives in CLAUDE.md's active-clients list and in the coach-report spec, so any agent that reads the config re-fires the flag whenever HiTone spend appears. The Jun 29 closure was recorded in the rocks file, but the upstream trigger rule was never retired.

L100 MEDIUM OBSERVED ONCE 3x efficiency

Prep must scan for WON/signed revenue explicitly, not just open pipeline: GHL won opportunities across ALL pipelines since the last meeting, plus Proposify signed events, categorized new/expansion/reactivation. Signed expansion revenue is the Q3 headline metric; it leads the scorecard section, and its dollar values get pushed to the manual Expansion tile the same morning. A tile with no automated source still gets its value entered at prep time from the won-deal scan; manual source does not mean no value.

Why: The board exists to catch exactly this: revenue proving or disproving the quarterly thesis. A prep that inventories dead tiles but misses live signed money reports the plumbing and skips the water. David finding revenue wins that his facilitator missed inverts the entire point of the seat.

Failure mode: Prep missed two signed expansion deals (Glo30 and WOA franchise additional revenue) that David saw on the board himself and had to point out at the 7/27 meeting: "you skip that a lot, did you not see them?" The prep scan read open opportunities in one GHL pipeline and the Expansion KPI tile (manual, no values), so signed/won expansion revenue had no path into the brief. The single most strategy-relevant signal of the quarter, expansion revenue from existing accounts, was invisible to prep while five dead tiles got named in detail.

L101 MEDIUM OBSERVED ONCE 3x efficiency

Same-session capture rule for ALL agents: when you witness David build or ship something real (a dashboard, an integration, a signed deal, a process), record it THAT session: a headline line in the daily note, and a todo or KPI value in OTP if money or a rock is touched. Do not wait for the weekly meeting; the builder remembering to report is not a capture mechanism. Dan additionally runs a Shipped This Week sweep every Monday prep as the backstop.

Why: Work done outside meetings is systematically invisible to a meeting-based OS, and the founder's most valuable hours happen outside meetings. Three instances surfaced in one meeting (dashboard, prep signals, signed revenue). Invisible work costs real money: David unknowingly built part of Bogdan's open reporting-cost rock.

Failure mode: David built a WOA dashboard for iCart (central database, less Zapier, better reporting) and the only witness was the AI in that session; no headline, todo, or record reached the operating system. David at the 7/27 meeting: "the only one that knows is you, this is a true failure of the operating system that needs to be reconciled." Same morning, two signed expansion deals (Glo30, WOA franchise) were also absent from every system of record.

L102 MEDIUM OBSERVED ONCE 3x efficiency

Facilitation = operating the product live. The moment a section starts, its artifact moves: an issue under discussion is verified rendering on the board before discussing it; the moment David decides, the ticket is solved with its resolution, the todo is created, the KPI is pushed, in that minute, not at conclude. After every state change, verify the rendered surface. Conclude should be a read-back of changes already made, never a batch of pending writes.

Why: A meeting inside OTP is only real if the product state changes while the humans watch. Deferred writes recreate the mirror-drift problem inside a single meeting, and David cannot trust a board that lags the conversation.

Failure mode: Dan facilitated IDS discussion in chat but did not move the meeting's product surfaces in real time: the Pulse issue being discussed was not visible in the meeting's IDS section, and the already-solved invisible-work issue still sat open on the board. David 7/27: "you should be moving the meeting along like a human, doing things as we work."

L103 MEDIUM OBSERVED ONCE 3x efficiency

The needle test has a step zero: before wiring, fixing, or reporting any KPI, confirm the program/offer it measures still exists in the business, by asking David or checking recent revenue/activity, not config files. A dead program's tile is retired, not wired, and its language gets swept from all agent config at source (L099 pattern) so no agent rebuilds it. Guarantee/T20 program: DEAD as of 2026-07-27.

Why: Config outlives strategy. Wiring effort spent on a dead program's metric is worse than a dead tile, it would have shipped a number that misrepresents the business as still running an offer it killed, and every agent reading the tile would have inherited the fiction.

Failure mode: Dan was one step from wiring the guarantee-clients-retained KPI (emit line written, Tally about to fire) when David said the guarantee program is dead and no longer offered. The wiring work treated the tile's PROGRAM as alive because the config said so; nobody had asked whether the business still runs the thing the tile measures.

L104 MEDIUM OBSERVED ONCE 3x efficiency

When a media element silently stalls (networkState LOADING, error null, no console output), suspect CSP: a 302 redirect from a same-origin playback route to a cross-origin storage URL violates media-src (falling back to default-src 'self') and Chrome blocks it with zero surfaced errors. Diagnose live by attaching a securitypolicyviolation listener and probing a known cross-origin media URL. Fix: add the presigned bucket origin to media-src, computed from storage config at boot, never hardcoded.

Why: This failure mode is invisible by design (no element error, no console noise) and the natural debugging paths (storage probes, presigned URL tests, range requests) all pass, sending you everywhere except the CSP header. Cost about an hour across two sessions; the probe technique turns it into a 2-minute check.

Failure mode: SUCCESS: Claude/Conatus diagnosed why OTP meeting recordings never played in the browser (player stuck at 0:00, click did nothing) while server-side storage checks all passed.

L105 MEDIUM OBSERVED ONCE 3x efficiency

Never invent a person's first name from an email address or initial. If the name is not stated, refer to them by the email address or ask David, and only record a name once confirmed.

Why: Guessed names propagate into client-facing artifacts (emails, dashboards, docs) and getting a client's name wrong damages trust; an email initial is not evidence of a name.

Failure mode: Claude inferred the Drybar Ballston client's first name as "Jodi" from the email address jsterling@sterlingcapitalllc.com and used it in the summary, credentials file, and memory. The client is Julie Sterling.

L106 MEDIUM OBSERVED ONCE 3x efficiency

`ghl.sh update-opp <oppId> <stage>` with the optional value argument omitted hits an inline Python syntax error (`data['monetaryValue'] = ` with nothing after it), so the stage payload is never built: the opp's updatedAt bumps but the stage does NOT change, and nothing errors loudly. Workaround until ghl.sh is fixed: always pass the value explicitly, e.g. `update-opp <id> sql 0` (confirm the opp's current monetaryValue first so you do not overwrite a real value). Always verify stage changes by re-reading the opportunity after the write.

Why: A stage move that silently no-ops corrupts the pipeline of record without any error signal. Both /otp-sales and /sneeze-sales route stage moves through this command; without the verify-after-write habit the bug would have shipped 2 phantom SQL bumps today.

Failure mode: SUCCESS: Sneeze-Sales found and worked around a silent ghl.sh write failure

L107 MEDIUM OBSERVED ONCE 3x efficiency

Before bumping a stage or queueing a task on any reply, check the PERSON and company against the active client list (CLAUDE.md), pepper-clients.md, and known client people, not just the cold-load exclusion at contact creation. Jordan Anderson is a client (Workout Anytime / Proof Fitness): never prospect-touch him on any domain. Replies inside threads a team member already owns (for example Zeynep scheduling) get no David task; the owner handles it. A reply landing in the prospect book is a signal to verify WHO it is, not proof they are a prospect.

Why: The prospects-only law fails at the edges: client people reply from domains that sit in the prospect book because franchisee outreach and client domains overlap (WOA). A wrong SQL bump plus a David task double-touches a client relationship and burns David's queue on non-sales work.

Failure mode: Sneeze-Sales treated Jordan Anderson (Workout Anytime) as a prospect: bumped his opp MQL to SQL on an inbound reply and queued David a respond-within-24h task. David corrected: Jordan Anderson is a client. Also queued David a confirm task on Lindsey (Fitness Factory) when Zeynep already owns that thread.

L108 MEDIUM OBSERVED ONCE 3x efficiency

When an OAuth integration fails silently, verify each layer with direct probes instead of reasoning from app behavior: (1) print env var names with cat -v, since a pasted quote becomes part of the variable NAME and the app reads undefined (Railway kv showed "GOOGLE_CALENDAR_OAUTH_CLIENT_ID with a literal leading quote); (2) test client credentials against the provider token endpoint with a bogus auth code, since the error distinguishes exactly: invalid_client means bad id/secret, invalid_grant Malformed auth code means credentials are VALID; (3) treat the database as ground truth for whether a flow completed, since users saying connected can mean a different surface (Composio integrations page vs the calendar card).

Why: Three probe layers turned what could have been hours of guessing into minutes: no callback in logs proved the flow died at Google, cat -v exposed the quote character, and the bogus-code token probe verified the replacement secret BEFORE the user retried, avoiding another failed round trip. Reusable for every OAuth integration OTP adds (Microsoft, Zoom, future providers).

Failure mode: SUCCESS: Claude shipped Recall calendar auto-join and debugged two invisible OAuth config failures the same afternoon (quoted env var name, invalid client secret)

L110 MEDIUM OBSERVED ONCE 3x efficiency

When graduating a feature out of Labs, grep the whole repo for the feature key and isFeatureEnabledForOrg calls before merging; every gate (page, API, scheduler, MCP tools) must come off in the same PR. A fail-closed flag check on a deleted key silently disables the feature for everyone, which reads as random 404s, not as a flag problem.

Why: Fail-closed gating is correct security posture, but it means catalog removal IS a kill switch. The bug shipped invisible because the page worked while the API did not, and the generic 404 body hid the cause; the honest error message shipped hours earlier (names the workspace and feature) is what made the real diagnosis possible.

Failure mode: Projects went GA (PR #347 removed it from the Labs catalog and the page gate) but the API routes kept gating on the removed key; isFeatureEnabledForOrg fails closed on unknown keys, so all project CRUD 404'd platform-wide for two days until Dawson and David hit it

L112 MEDIUM OBSERVED ONCE 3x efficiency

Any otp-platform endpoint that scopes data or authz by getAuth(request).userId is broken under impersonation. The rule: gate and scope by the EFFECTIVE viewer (request.impersonation.as when active, else auth.userId), return/audit with the RAW session id, and for context-pinning actions (org switch) re-issue the impersonation cookie via startImpersonation rather than moving the admin's own cookies. When reviewing or writing any new route, grep for getAuth(request).userId used in a WHERE clause — each one is a latent impersonation bug.

Why: Impersonation is how David supports customers (view-as Tom, Kristen, etc.). Every raw-session usage silently shows the admin's data under the customer's banner or 403s the customer's own surfaces — a privacy leak in one direction and a support dead-end in the other. Fixed instances: dashboard (2026-06-02), PRs #381, #382, #384 (2026-07-28).

Failure mode: SUCCESS: Dan identified a recurring defect class in otp-platform — four separate surfaces broke under super-admin impersonation in one day (portfolio pages listing the admin's portfolios, portfolio API 403ing "Could not load team", the sidebar org label showing the admin's org, and the org switcher 403ing "Could not switch organization"), all with the same root cause.

L117 MEDIUM OBSERVED ONCE 3x efficiency

Do not treat a missing or small wallet balance as a signal of anything. An org with no wallet is simply not an active OTP user, so its balance says nothing about product readiness. New orgs are seeded with $25 of credit to incentivise starting, so a funded wallet is the default going forward rather than a hurdle. When assessing whether a metered feature is usable, filter to orgs with real activity and check the NON-wallet prerequisites, since those are the ones that actually gate anyone.

Why: Reporting wallet balances as blockers manufactures work out of the ordinary shape of the user base: most rows in that table are dormant signups, not stuck customers. It also buries the prerequisite that does bite, because a real blocker listed next to four fake ones reads as one item in a list instead of the single thing to fix.

Failure mode: Flagged orgs with low or missing wallet balances as a readiness problem for OTP scheduling, treating wallet funding as a live blocker worth David's attention.

L120 MEDIUM OBSERVED ONCE 3x efficiency

Before retiring a scarcity, cohort or badge claim, count each cohort in the database and check whether one phrase names several cohorts. Shared wording is not shared meaning. If counts contradict the instruction's premise, return to the human with the numbers instead of executing literally.

Why: Broad approval rests on an assumed premise. When the premise is partly false, literal execution silently destroys value: a deleted live offer throws no error, it just yields fewer signups. One query per cohort is cheaper than a loss nobody detects.

Failure mode: SUCCESS: Claude verified cohort counts before executing an approved codebase-wide sweep of OTP's "first 50" claim. It was closed for signups (50/50) but live for Founding Publishers (45/50) and Founding Partners (6/50). Literal execution would have deleted two accurate live offers.

L122 MEDIUM OBSERVED ONCE 3x efficiency

For multi-agent feature builds: (1) research agents return structured briefs before any code, and briefs override the spec when they conflict (two detectors were impossible as specced: meetings have no booked-duration column, decisions are not rows). (2) Put all shared-file wiring (server.ts) in ONE sequential final task so parallel agents never collide. (3) Always run an independent fresh-context diff review before committing: it caught two honesty blockers the per-task verifications missed (ratified moves netting costs away; org-wide gains summed over subtree-scoped costs). (4) Discovery worth acting on: subscriptions.plan_rate is never written by any code path in otp-platform, so any revenue/cost feature reading it ships dark until billing populates it.

Why: The per-task agents were all green individually; only the cross-seam review found the invariant violations. Repo guard tests (private-issue leak scan, blueprint coverage) also fired exactly as designed, proving lint-style guard tests catch what unit tests cannot.

Failure mode: SUCCESS: Claude shipped OTP Impact Phase 1 (PR #427) via 11-agent build: parallel research briefs, wave execution with pure-function cores, then adversarial diff review before commit

L123 MEDIUM OBSERVED ONCE 3x efficiency

When adding a NEW page to an existing app, open a sibling page that already ships (for OTP admin surfaces, /admin/support) and copy its outer container, top padding and control classes verbatim before writing any markup. Do not hand-roll spacing from DESIGN.md tokens alone -- the tokens do not tell you the page-level offsets that keep content clear of the fixed header. For any default that is a money amount, confirm the number rather than inferring it from the option list order.

Why: A new page laid out from first principles looks subtly wrong in ways the author cannot see without loading it: header collisions and control scale only show up in a browser, not in typecheck, lint or design-lint, all of which passed. Copying a shipped sibling inherits every page-level decision already made and reviewed.

Failure mode: Built /admin/join-link with hand-rolled Tailwind layout (max-w-3xl, custom padding, custom input classes) instead of copying the container and control classes from an existing admin page. Result: the page header collided with the fixed top nav so the title was unreadable, and the form controls were oversized versus OTP's 32px control scale. Also picked $50 as the default starting credit without asking; David wants $25.

L124 MEDIUM OBSERVED ONCE 3x efficiency

When promoting any OTP Labs feature from beta to live, do THREE things, not one: (1) flip `status` in src/shared/lab-features.ts; (2) grep src/views for the feature's `surfaceUrl` -- if the ONLY link is the Labs-injected rail item, add a permanent entry to layouts/main.ejs in BOTH the `_sbItems` array and the mobile settings menu; (3) grep the page for stale "this is a Labs feature" banner copy pointing at a /settings/labs toggle that graduation just removed. Also verify any second, independent gate (e.g. an env check like recall-calendar.ts calendarIntegrationEnabled) and make the registry copy match what is actually configured in production -- check `railway variables --kv` rather than trusting the existing description.

Why: Graduation looks like a one-line status change and is not. The rail item, the page's own Labs banner, and any env-based second gate all key off the old state, so a naive flip can make a feature LESS reachable than it was in beta while appearing to ship it. Confirmed live: PR #437 shipped the flag plus the nav entry together, and the promoted page rendered correctly with the Calendar section visible and no Labs opt-in.

Failure mode: SUCCESS: Claude caught that graduating an OTP Labs feature from `beta` to `live` silently DELETES its left-rail nav item, which would have shipped calendar auto-join into being unreachable. `getOrgLabNavItems` (src/services/lab-features.ts) filters on `f.status === 'beta'`, so only beta features get a rail item injected. /settings/meeting-presence had no other link anywhere in src/views, so flipping the flag alone would have removed the only way to navigate to it.

L018 MEDIUM OBSERVED ONCE 3x efficiency

Before presenting a Dash blind-spot, billing trigger, or any state-file alert as a current action item, confirm it hasn't already been resolved. Stale state files (Dash May 25 was ~5 weeks old) carry point-in-time alerts that may be closed by now. Trust confirmed/observed status over stale notes; flag the data's age and treat unverified alerts as 'verify' not 'urgent.'

Why: Re-surfacing already-resolved alerts as urgent erodes trust in the L10 briefing and spends David's attention during a low-push recovery window. Honesty about data staleness matters more than appearing comprehensive.

Failure mode: Dan surfaced the HiTone billing trigger ($43-49K/mo possibly un-invoiced) from Dash's stale May 25 state file as a live concern during the Jun 29 L10. David confirmed HiTone billing is correct and already handled, and asked to close it out.

L019 MEDIUM OBSERVED ONCE 3x efficiency

KPIs/scorecards must live as tiles in OTP (the source of truth), not in markdown files or meeting briefs. Every active agent/human seat — including Dan's strategic co-founder seat — must own at least one OTP KPI tile. When proposing measurables, verify against list_my_kpis and create the missing tiles via update_kpi (auto-creates), rather than just tabling them in a doc. A seat with no number is sitting on the sidelines.

Why: EOS requires every seat to have a measurable. Discussing KPIs in a brief while OTP shows none of them makes the scorecard fiction and undercuts OTP as the coordination source of truth. Dan as co-founder must be measurable like everyone else.

Failure mode: Dan presented a Sneeze It agent-team scorecard as a markdown table in the L10 brief and treated it as 'the scorecard,' when the source of truth is OTP. David caught that Dan (and Arin/Pulse/Dirk) have NO KPI tiles in OTP at all — Dan's own seat had zero measurables. A scorecard that only lives in a file or meeting brief does not exist.

L021 MEDIUM OBSERVED ONCE 3x efficiency

An OTP KPI with teamId=NULL renders only on /dashboard/kpis, never on any L10 scorecard (meeting scorecards filter strictly by meeting.team_id). To make a KPI show on a specific L10, PATCH /api/v1/kpis/:id with the meeting's teamId. The 'Dan L10' meetings run on the 'ai-army' team (065d1d4b-c7da-4e80-b3ed-d6b101471d2c). Tally's auto-create now includes teamId from a 'team_id' field in the registry entry, so new agent-army KPIs land on the Dan L10 automatically instead of orphaned. Find team IDs via GET /api/v1/teams; meeting->team via GET /api/v1/meetings.

Why: A KPI nobody can see on their meeting scorecard is functionally not on the scorecard. The owner/title is necessary but not sufficient — team scoping is what makes it report. This is a recurring gotcha for any agent creating KPIs via the API.

Failure mode: SUCCESS: Tally — agent KPIs were invisible on the L10 because auto-create left teamId NULL. David flagged that the new KPIs weren't reporting on the Dan L10 or /dashboard/kpis as expected.

L023 MEDIUM OBSERVED ONCE 3x efficiency

Root cause was the SearchAtlas OTTO pixel having an EMPTY src="" in the layout head (v7.ejs, onboarding.ejs, main.ejs). The OTTO tag must carry its base64 data-URI loader in src that appends dynamic_optimization.js with data-uuid; with src="" the runtime never loads, so OTTO injects/verifies nothing. When an OTTO/SearchAtlas audit reports 0/N across ALL on-page categories, suspect the pixel loader, not the actual tags — verify the sa-dynamic-optimization script's src is populated, not the page's own meta.

Why: A 0/16 across every category despite visibly correct meta is the signature of a non-loading optimization runtime, not missing tags. Checking the pixel first avoids a pointless rewrite of titles/descriptions that were never the problem.

Failure mode: SUCCESS: Beacon/SEO — orgtp.com OTTO on-page audit showed 0/16 (titles, meta descriptions, headings, meta keywords all failing) even though pages had perfectly good title tags and meta descriptions server-side.

L024 MEDIUM OBSERVED ONCE 3x efficiency

A brand battle cry needs a genuinely designed moment (confident display type, intentional line breaks, brand device, real whitespace), not a centered text block plopped in. And the VISIBLE battle cry copy is the short clause only: 'Unlocking the potential in every person through the partnership of people and AI' — drop 'so together we leave the world better than we found it' from the hero display (keep the full sentence only for formal/footer contexts).

Why: A mission line is a brand centerpiece. Long copy dilutes the punch, and an undesigned drop-in reads as filler. The payoff phrase 'partnership of people and AI' must land as the climax with design weight behind it.

Failure mode: Adding the OTP mission as a 'battle cry' on the landing page, I dropped the full sentence into a plain centered text band wedged between hero and Step 1. David called it 'a weak attempt to just throw it on the page' and said the full line is too long for the visible battle cry.

L025 MEDIUM OBSERVED ONCE 3x efficiency

Manifesto/mission pages must be written as movement recruitment, not product marketing: second-person address (the reader is the protagonist), "We believe" creed statements people can recite, a named enemy, stakes, and invitation CTAs ("Join the movement") instead of transactional ones ("Start free"). Product features appear only once, framed as how the movement fights, not what the product includes.

Why: People join movements because they believe what the movement believes (Sinek: start with why). Copy that sells the what on a page whose job is to recruit believers reads as generic SaaS and inspires no one, no matter how good the design is.

Failure mode: Redesigned the orgtp.com manifesto homepage with strong visual design but kept product-brochure copy (feature lists, "free meeting software", "Start free" CTAs). David: "the writing does not inspire an army of followers... this just looks the same as every other company... blah."

L026 MEDIUM OBSERVED ONCE 3x efficiency

Judge conversion on the full path the visitor actually walks (page, door, day-one experience), not on the surface being edited. If the honest answer to "would you sign up" is "yes IF another surface delivers," the answer is no, and the work moves to that surface. Never write a promise on a button that the destination page cannot cash.

Why: Trust destroyed at the moment of verification is unrecoverable; a skeptical buyer who clicks "watch us run" and lands on a data page is gone forever. Copy that outruns proof is hype by definition, and the exact audience OTP needs (operators) is the audience that punishes it hardest.

Failure mode: After rewriting the OTP homepage, I declared the copy converts because skeptics would "click through to the live OOS page and sign up IF it delivers." David called it: I kicked the can to a page I know does not deliver, and called it a win. The button promises "Watch our company run, live" but the OOS page it links to is a list of published rules, not a running company.

L027 MEDIUM OBSERVED ONCE 3x efficiency

OTP's enemy statement is "you bought the operating system and the needle didn't move." The pitch is not better meetings; it is: the system was fine, what was missing was the workforce that runs it between the meetings. Frame all homepage/sales copy against needle-not-moving, not against meetings.

Why: This is the buyer's actual lived disappointment (paid for an operating system, company looks the same two years later) and it positions OTP against incumbents on outcomes instead of features.

Failure mode: The letter's hero framed the enemy as "the meeting" / busywork. David corrected the thesis: the real problem is that companies bought operating systems and software (Ninety, Bloom Growth, etc.) that did not move the needle. Years later the company had not grown and was not better, and they needed to change how they did things.

L029 MEDIUM OBSERVED ONCE 3x efficiency

For any UI change, design from the user's mental model, not the data model: "my list shows my work; work I assigned to others shows under Waiting on Others." When a meeting todo is assigned to someone else, stamp the creator as delegator so it routes to the delegation view. Before shipping UI changes, run the UX lens (impeccable / web-design-guidelines skills + src/DESIGN.md), not just a minimal code patch.

Why: A technically-correct patch that ignores the user's mental model just moves the confusion. OTP's own product language already has the right home for these items (Waiting on Others); fixes should land in the model the user already understands.

Failure mode: Fixed the dashboard todo confusion (teammates' meeting todos looked like the viewer's own) by adding an owner label to the rows. David corrected: that's not thinking like a user. Labeled-or-not, other people's todos don't belong in "my to-dos" at all.

L030 MEDIUM OBSERVED ONCE 3x efficiency

When David asks for a jaw-drop brand page, build an EXPERIENCE, not an article: full-viewport cinematic hero, scroll choreography, one idea per screen at massive scale, motifs that live in the page as motion, ruthless copy cuts, no standard nav/footer chrome breaking the spell, no section-grammar scaffolding, no FAQ accordion bolted onto a manifesto.

Why: The gap between "well-executed page" and "omg I love this" is the whole assignment on brand surfaces. Safe editorial structure is invisible at best; for-the-brave positioning demands the page itself be brave.

Failure mode: Built the /ollie manifesto page as a competent editorial layout (repeated mono eyebrow labels on every section, index rows, alternating light/dark sections, FAQ accordion at the bottom) and David rejected it outright: "this really really sucks." The brief was "reader drops on the ground saying omg I fucking love this" and the output was a safe template that reads as AI scaffolding.

L031 MEDIUM OBSERVED ONCE 3x efficiency

When David gives a design reference URL, open it in a browser and STUDY it visually (proportions, type sizes, spacing, alignment) before designing; match its register, not just its layout skeleton. Elegant means restrained: modest type scale, centered calm hierarchy, generous whitespace, thin rules. Never hand-draw SVG artwork to imitate produced brand art; crop/reuse the actual asset or use nothing.

Why: A reference URL is the brief. Reading its HTML structure without seeing it rendered led to importing the skeleton with the wrong soul, twice. Amateur freehand art next to professional motion work destroys credibility instantly.

Failure mode: Second rejection on the /ollie page. David asked for sakana.ai/fugu: elegant, Japanese sense of design (restraint, whitespace, calm, modest type, precision). I delivered giant 9vw headlines, one shouting line per viewport, and hand-drawn SVG chevron "birds" that rendered as crude fat marker scribbles. I treated "jaw-drop" as scale and boldness when the reference was quietness and precision, and I drew freehand SVG art instead of using the actual video's artwork.

L032 MEDIUM OBSERVED ONCE 3x efficiency

For fleet-wide spec maintenance: (1) tarball backup of ~/.claude before any agent touches specs; (2) partition files into DISJOINT clusters, one agent each, with CLAUDE.md owned by exactly one; (3) give every auditor the same stale-fact canon and the rule "verify a launchd plist exists before believing any schedule claim"; (4) auditors apply surgical edits directly for factual fixes but RETURN structural proposals for David instead of applying them; (5) synthesizer closes cross-cluster contradictions the auditors flag at each other.

Why: Agent specs rot faster than anyone audits them: this pass found live specs for a retired agent (jeff.md ending in "Go."), four phantom schedules, Todoist writes in five files, terminated employees still routed DMs, and a Bassim score-inflation bug. Periodic fleet audits with disjoint ownership are cheap insurance against agents acting on dead infrastructure.

Failure mode: SUCCESS: Claude ran a five-cluster parallel level-up of the entire agent army (80 files, ~140 surgical edits) without a single file conflict or lost spec.

L033 MEDIUM OBSERVED ONCE 3x efficiency

Pattern for UX dead-end hunts: (1) fan out parallel read-only explorers per surface (meetings, teams/members, KPIs/todos, onboarding/settings) asking for file:line + user-visible symptom + minimal fix; (2) fix the unsatisfiable states first: any required dropdown that can render zero options must explain where its options come from and link there (owners/attendees come from the org chart, meeting membership from teams); (3) empty states must branch on WHY they are empty (org has no teams vs user not on a team need different CTAs); (4) never report an async side effect as done: invite emails now await sendEmail (which returns null on failure, never throws) and return emailSent so the UI can tell the truth; (5) a guided setup checklist computed server-side from actual data (seats/team/KPI/meeting/members exist?) beats static onboarding because it survives skipped onboarding.

Why: These are the recurring shapes of broken UX in OTP: forms with prerequisites the user cannot see, empty states that misdiagnose their cause, and optimistic success messages over fire-and-forget side effects. Fixing the shape, not just the instance, is what makes the product feel intuitive.

Failure mode: SUCCESS: Claude ran a full UX dead-end audit and fix pass across OTP (4 PRs, #113-#116, all deployed)

L035 MEDIUM OBSERVED ONCE 3x efficiency

The sweep pattern that worked: audit by rule-cluster in parallel (fakery, insight-to-agency, jargon/states, first-meeting goal-walk), then execute severity-first. Key catches to re-check every run: (1) seeded/synthetic data leaking into numbers a reader believes are real (the is_template flag existed but was never enforced; counts now use src/shared/synthetic-orgs.ts); (2) the conversion moment must be ON the default path (end-meeting now lands on Ollie followups, not the list); (3) funnels don't exist until instrumented (insight topic: surfaced/accepted/value_delivered); (4) credentials in seed script comments (one prod DATABASE_URL scrubbed; password rotation still owed). Worklist for run 2 in otp-platform/mission-standard/WORKLIST.md.

Why: The Mission Standard is a repeatable bar, not a one-off audit. Recording the found failure classes makes run 2 start from run 1's ceiling instead of re-discovering it.

Failure mode: SUCCESS: Claude ran Mission Standard sweep run 1 (PRs #121-#123, deployed): 4 parallel rule-audits over OTP, 5 CRITICAL + 13 GAP found, all CRITICAL and 9 GAP closed same-session

L038 MEDIUM OBSERVED ONCE 3x efficiency

When an agent-pushed OTP to-do references a document, include a clickable https link (a Google Doc), not a local file/vault path — David reviews to-dos on mobile. To update an existing to-do's description, use PUT /api/v1/todos/:id (not PATCH). otp-todo.sh has no update verb, so PUT directly with the API key.

Why: A file path in a to-do is dead weight on mobile — the reviewer can see the reference but cannot open it, which reads as "the link is missing/broken." Every agent that pushes doc-linked to-dos (Radar, Pepper, Dan) hits this.

Failure mode: Dan pushed an OTP to-do referencing a document but put a local Obsidian vault path ("2nd Brain/Agent Army/Dan/...") in the description. David opens to-dos on his phone — a vault path is not tappable, so there was "no link to click." Also used PATCH to update the to-do; the OTP todos API update method is PUT /api/v1/todos/:id (PATCH hits the marketing site and returns HTML).

L039 MEDIUM OBSERVED ONCE 3x efficiency

To put an agent-army/IDS issue on an OTP meeting board, POST /api/v1/tickets with the team's teamId (category 'other' for strategic issues, priority low/medium/high/critical, ownerEntityType+ownerExternalId). The MCP submit_ticket tool CANNOT do this — it has no teamId param (it is the generic 'report a bug to OTP' path), which is why nobody ever got issues onto the board. Team IDs: 'AI Army' = 065d1d4b-c7da-4e80-b3ed-d6b101471d2c (the David+Dan agent-army meeting); Leadership Team = c1e1a485-414e-48d5-ae44-e81bd110b554. Update/solve via PUT /api/v1/tickets/:id (idsStatus, priorityRank, resolution).

Why: Agents could push KPIs and todos to OTP but not issues, so every L10 IDS board rendered empty and David kept discovering the hole live. Issues=tickets + teamId scoping is the missing piece; without it a meeting-readiness check would keep mislabeling a working API as absent.

Failure mode: Dan claimed 'no issues API exists in OTP' because there is no src/routes/api/issues.ts. That was wrong. OTP stores IDS issues in the TICKETS table (schema.ts: 'issues live in the tickets table'), with full IDS support (idsStatus, priorityRank, teamId, owner fields). The agent-army IDS board was empty only because our issues lived in a local markdown file and were never pushed as tickets scoped to a team.

L040 MEDIUM OBSERVED ONCE 3x efficiency

OTP has TWO distinct features both called "Ollie Insight": (A) the per-meeting followups wizard that turns a transcript into meetings.ai_summary (src/shared/meeting-followups.ts, transcript-only), and (B) the reusable "address engine" (ollie_insights table, src/services/ollie-insight.ts + shared/ollie-insight.ts + partials/ollie-insight-block.ejs) that gathers org data (KPIs/rocks/todos/meeting-summaries) per SCOPE. When David said "KPIs shouldn't be in the meeting analysis," the fix was in System B's meeting-scope evidence gathering, NOT System A. The block partial (ollie-insight-block.ejs) is fully scope-generic (builds the API URL from data-oib-scope/scopeId client-side), so adding a brand-new 'quarter' scope end-to-end took only: add to INSIGHT_SCOPES + RULES_BY_SCOPE (shared), the scopeQuerySchema enum + a resolveInsightScope branch (api), a gatherEvidence branch + max_tokens (service) -- then just include the existing partial with scope:'quarter'. No new render/generate/receipts UI. Pattern: when a feature is "one engine, many surfaces," new surfaces are a scope + evidence branch, never new UI.

Why: The name collision hides which code to touch; picking the wrong system wastes a whole edit pass. And recognizing the scope-generic block means big-feeling asks ("a quarterly synthesis button") are small, low-risk diffs. Both are recurring shapes in OTP's Ollie work.

Failure mode: SUCCESS: Claude separated OTP's two "Ollie Insight" systems and added a whole new scope by reuse

L041 MEDIUM OBSERVED ONCE 3x efficiency

New endpoint POST /api/v1/meetings/:id/agent-record (PR #154): an agent submits the written meeting record; OTP runs the same redaction ruleset, stores it to meetings.transcript, logs an audit baseline + agent_record event, and the existing /ai/followups generate turns it into to-dos/issues/headlines/insight unchanged. Agent path: ~/.claude/otp-meeting.sh record <meetingId> --file=<record> --source=l10dan. Wired into /l10dan conclude step 6. Verify deploy by probing the endpoint returns JSON not the marketing SPA HTML before pushing records.

Why: Ollie only reads transcripts, so agent-run meetings had no way into it — the empty-insight hole David hit live. This closes it: any agent meeting can now produce Ollie follow-ups. Also a general UI rule captured: buttons reflect capability/state (action taken -> button disappears).

Failure mode: SUCCESS: Dan shipped the Ollie agent-record path so agent-facilitated meetings (David + AI L10, no audio transcript) can feed Ollie Insights.

L042 MEDIUM OBSERVED ONCE 3x efficiency

Before treating a KPI/data source as blocked, re-read the LIVE source, not the note about it. The Havok "non-client %" KPI was marked blocked for ~14 weeks on "the timesheet has no client column" — a 105-day-old memory. The live sheet (1VPlH5ZqTowOe2nJDFwOjuvXrCXUnOzeOb3aO-M936Xo) had since grown per-person tabs with a full Client/Time/Date schema; one read unblocked it. Pattern for wiring a messy sheet into Tally: (1) get_spreadsheet_info to list tabs, (2) read a person tab to learn the real schema, (3) confirm recency by reading the tail (last date), (4) add a focused extract mode to tally.py rather than reshaping the sheet — here `client_attribution_since` (reads ALL valueRanges via a new _values_2d_all, parses h:mm via _parse_hhmm, added dotted DD.MM.YYYY to _parse_date, classifies internal by an `internal_contains` substring), (5) `tally.py --dry-run --kpi "<title>"` to prove the number off live data before pushing. Human-owner + a 1:1-team KPI (owner HUM_BOGDANTABAKA, team = David-Bogdan 1:1) pushes fine via find_or_create_kpi.

Why: Blocked-status notes rot silently while the underlying source improves; a KPI can sit "pending" for a quarter when it was buildable weeks ago. Re-reading the live source first is the cheap unblock. And the tally.py extract-mode pattern makes any timesheet/sheet a live KPI without asking a human to restructure their doc.

Failure mode: SUCCESS: Dan/Tally shipped the Havok client-attribution KPI live in one session after a 14-week "blocked" note turned out stale

L044 MEDIUM OBSERVED ONCE 3x efficiency

Before editing any OTP view to fix an on-screen bug, grep unique visible strings from the screenshot (e.g. "WAITING ON OTHERS", "always only yours") across src/views to confirm WHICH template renders that exact surface. Multiple pages can render similar-looking todo lists (me-todos.ejs vs dashboard-daily.ejs). Verify the rendering route (reply.view target) too.

Why: Two round-trips and two merged PRs produced zero visible change because the edits were on the wrong template, which read as "nothing is fixed" and eroded trust. A 10-second grep on the screenshot text would have pointed to the right file immediately.

Failure mode: Fixing an OTP todos UI bug, I edited src/views/pages/me-todos.ejs twice and shipped two PRs, but the surface David actually uses is the dashboard-daily "Waiting on others" widget (src/views/pages/dashboard-daily.ejs). Nothing he saw changed.

L045 MEDIUM OBSERVED ONCE 3x efficiency

Two reusable patterns: (1) Before building any OTP email/engagement feature, grep src/services for existing infrastructure -- re-engagement.ts, lifecycle-scheduler.ts, and user_engagement_log already carried cadence caps, suppression, logging, and a daily cron, so the todo-aware upgrade was ~350 lines instead of a new subsystem. (2) When resolving "which org does this Clerk user belong to", organizations.clerkOrgId only knows the org CREATOR; invited teammates must be resolved through org_members.clerkUserId + claimedEntityIds. This gap is why per-user personalization (open todos) missed non-creator members like Nate.

Why: One engagement channel with shared caps is what keeps daily utilization pressure from becoming annoying double-mailing, and the creator-vs-member resolution gap will bite any future per-user feature (digests, notifications, billing seats) that starts from organizations.clerkOrgId.

Failure mode: SUCCESS: Claude shipped the smart engagement email engine (PR #186) by upgrading the existing re-engagement service instead of building a parallel system

L048 MEDIUM OBSERVED ONCE 3x efficiency

Frame help/success/onboarding call copy positively: state plainly that the call is there to help and guide them, and describe what will actually happen on it (we'll walk through your setup, get you unstuck, answer your questions). Never say "this is not a sales call" or "no pitch" — describe the help, don't disclaim the sell.

Why: Defensive "not a sales" language triggers the exact suspicion it tries to defuse and undercuts a genuine help offer. David flagged this immediately.

Failure mode: Wrote a "customer success call" Calendly description that leaned on "no pitch, no slides" / not-a-sales-call framing. Protesting that it isn't a sales call makes it sound like one.

L049 MEDIUM OBSERVED ONCE 3x efficiency

Any script that answers "is anything missing / is everything covered?" must fail LOUD, never return an empty set as reassurance. Two rules: (1) assert the expected top-level key exists (`if 'data' not in resp: raise`) before computing a result; a zero/empty answer from a health check is a claim that must be proven, not a default. (2) Cross-check a zero result against one known-positive case before reporting it -- here, one direct call to a single account would have shown $57 of spend and exposed the lie instantly. Also: `~/.claude/meta-ads.sh accounts` exits 0 and prints nothing; do not build on it. Sweep with `/{business_id}/adaccounts?fields=name,account_status&limit=500` then batch `/{act_id}/insights` 50 at a time.

Why: "Nothing is wrong" is the single most dangerous output an audit can produce, because nobody investigates it. A silent empty result on a billing sweep means real revenue is never invoiced and nobody ever finds out. The failure mode is not a crash, it is confident silence.

Failure mode: SUCCESS: Claude caught a silent false-negative in a Meta billing sweep. Querying the Graph API adaccounts edge with nested field expansion (`fields=name,insights.date_preset(this_month){spend}`) returned an error payload with NO `data` key. The sweep script read it as an empty account list and confidently reported "0 accounts, $0.00 unbilled MTD spend" -- a clean bill of health that was entirely fabricated. Direct per-account queries then revealed 7 unbilled accounts spending $5,673 MTD.

Stackwise silver
C011 MEDIUM INFERENCE 2x efficiency

Support responses for annual plan customers auto-flagged YELLOW. Annual customers represent 4x monthly revenue.

Why: Sloppy response to monthly customer costs $150/year if they churn. Sloppy response to annual customer costs $1,800. Review depth should match revenue at risk.

Failure mode: Annual customer billing question gets generic response. Feels undervalued. Does not renew. $1,800 lost.

C012 HIGH MEASURED RESULT 10x efficiency

Engineering alerts suppresses repeat pages for same issue within 30 minutes. First alert pages. Subsequent alerts update the existing incident thread.

Why: A database slow query triggered 14 pages in 8 minutes. On-call engineer overwhelmed. Missed the actual resolution signal buried in noise.

Failure mode: Same issue generates 14 pages. Each interrupts the engineer. Noise drowns signal. Resolution delayed 20 minutes.

C010 HIGH MEASURED RESULT 10x efficiency

Patient education handouts are written at a 6th-grade reading level using the Flesch-Kincaid scale. The education agent checks readability before submitting for review.

Why: Our patient population includes a significant number of non-native English speakers and older adults. Early handouts scored at 10th-grade reading level. Dr. Okafor observed patients nodding along but clearly not understanding the content. Post-visit comprehension checks confirmed the gap.

Failure mode: Handout on managing hypertension uses terms like "antihypertensive regimen" and "sodium restriction protocol." Patient takes the handout home, does not understand it, and does not follow the guidance. Blood pressure remains uncontrolled at the next visit.

C011 MEDIUM MEASURED RESULT 6x efficiency

Appointment reminders for patients who no-showed their last visit include a warmer, non-judgmental tone and an explicit offer to reschedule. No mention of the missed appointment.

Why: The default reminder tone felt transactional. Patients who had already missed once responded better to "We'd love to see you" than "You have an appointment on Tuesday." Reschedule rate for prior no-shows improved from 31% to 48% after the tone change.

Failure mode: Standard reminder sent to a patient who missed their last appointment. Patient feels guilty or defensive. Ignores the reminder. No-shows again. Pattern solidifies.

C012 MEDIUM OBSERVED ONCE 3x efficiency

Onboarding documents are versioned with a date stamp in the filename. When clinical protocols change, the onboarding agent regenerates affected documents within 48 hours. Old versions are archived, never deleted.

Why: A new MA was trained using a document that referenced the old blood draw protocol (tourniquet for 60 seconds). Protocol had changed to 30 seconds two months prior. Document had not been updated. MA followed the outdated procedure for a full week before a supervising nurse caught it.

Failure mode: Outdated onboarding document trains new staff on a deprecated procedure. Staff performs the procedure incorrectly. In a primary care setting, most deprecated procedures are low-risk, but the cumulative effect of outdated training erodes clinical quality.

C013 HIGH OBSERVED ONCE 5x efficiency

Incident severity is determined by customer impact radius, not system impact. A database hiccup affecting 1 internal dashboard is P3. A 200ms latency increase affecting all customer eval jobs is P1.

Why: Internal systems failing is inconvenient. Customer-facing systems degrading is revenue-threatening.

Failure mode: Monitoring agent classified a latency spike as P3 because the internal system health dashboard showed green (it only measured error rates, not latency). 85 customers experienced 3x slower eval results for 2 hours. 15 filed support tickets. 3 enterprise customers included the incident in their quarterly vendor review. Two of those reviews resulted in "conditional renewal" status.

C014 MEDIUM OBSERVED ONCE 3x efficiency

Investor updates are published monthly, on the 5th, regardless of whether the numbers are good. Skipping a month signals that something is wrong.

Why: Investors pattern-match on communication cadence. A missed update generates more anxiety than a bad update.

Failure mode: Lina skipped the March investor update because MRR had dipped 4% (3 customers delayed renewals). Two board members texted within a week asking "everything okay?" The April update included the March data and the dip explanation, but the trust damage from the missed communication took the entire board meeting to repair.

C015 MEDIUM OBSERVED REPEATEDLY 4x efficiency

Sales demo prep agent refreshes demo environments weekly. Stale demo data that references outdated features or deprecated APIs undermines credibility during live demos.

Why: Enterprise prospects evaluate attention to detail. A demo that shows a deprecated feature signals that the product moves faster than the company can manage.

Failure mode: A demo environment showed an eval metric type that had been deprecated 2 months earlier. The prospect asked about it. The sales engineer said "oh, that's been removed." The prospect replied: "So your demo doesn't reflect your actual product? What else is out of date?" The deal took an additional 3 weeks to close and included a requirement for a "current state" audit before signing.

C012 HIGH OBSERVED REPEATEDLY 7x efficiency

Haven's first response to any customer inquiry must be sent within 15 minutes during business hours (9 AM - 6 PM ET). The response can be a draft that the CS rep reviews, but the customer must see a reply within 15 minutes. Outside business hours, the autoresponse sets expectations for next-business-day response.

Why: Speed to first response is the single highest-correlating factor with CS satisfaction scores. A fast "we're looking into this" beats a slow comprehensive answer every time. Haven's drafts are fast. Human review can happen after the first touch.

Failure mode: Before Haven, average first response time was 4.2 hours. Customer satisfaction (CSAT) was 3.4/5. After implementing the 15-minute target with Haven drafts, first response dropped to 8 minutes average. CSAT rose to 4.1/5 within 6 weeks. No other change was made during that period.

C013 MEDIUM OBSERVED ONCE 3x efficiency

Rhythm must test subject lines on a 10% sample before full send for any campaign going to more than 5,000 contacts. The winning subject line (by open rate after 2 hours) goes to the remaining 90%. No exceptions for "time-sensitive" campaigns.

Why: A 5% improvement in open rate on a 28,000-person list is 1,400 additional opens. On Threadline's average click-to-open rate of 18%, that's 252 additional clicks. At a 3.2% conversion rate, that's 8 additional orders averaging $67 each -- $536 in revenue from a 2-hour wait.

Failure mode: The marketing coordinator overrode the A/B test for a Black Friday campaign because "we need to send now, every minute counts." The chosen subject line had a 14% open rate. The founder ran the unused B variant to a test segment later: 23% open rate. Estimated lost revenue from skipping the test: $4,800 on Black Friday, the single highest-revenue email day of the year.

C010 HIGH OBSERVED REPEATEDLY 7x efficiency

Quarterly investor reports are prepared 21 days before distribution. The first 7 days are for agent drafting and internal review. The next 7 days are for Chen's compliance review and outside counsel if needed. The final 7 days are buffer for revisions.

Why: Rushing quarterly reports produces the C003-type errors. The 21-day cycle ensures every number is audited, every statement is compliant, and there is time to fix problems.

Failure mode: Before the 21-day cycle, Q3 reports were prepared in 5 days. The C003 incident (preliminary vs. audited IRR discrepancy) happened because there was no time for Derek to complete the audit reconciliation before distribution.

C011 HIGH OBSERVED REPEATEDLY 7x efficiency

Deal memos include a mandatory "Risk Factors" section with a minimum of 8 risk factors. The deal memo agent generates risk factors from a master risk taxonomy and adds deal-specific risks identified during analysis.

Why: Insufficient risk disclosure in offering materials creates legal liability. If an investor loses money on a risk that was foreseeable but undisclosed, the liability falls on the fund.

Failure mode: An early deal memo had 3 risk factors, all generic ("Market conditions may change," "Past performance does not guarantee future results," "Real estate is illiquid"). Chen added 9 deal-specific risks including environmental remediation liability, tenant concentration risk, and interest rate sensitivity. After this, the minimum was set at 8 with mandatory deal-specific analysis.

C012 MEDIUM OBSERVED ONCE 3x efficiency

Investor communications use tiered language precision based on the content type. Performance updates: exact numbers with 2 decimal places and data source attribution. Market context: ranges and qualifiers ("approximately," "in the range of"). Outlook: conditional language only ("if market conditions persist," "subject to").

Why: Precision signals competence. But false precision on uncertain topics signals naivete or deception. An investor who reads "We project 14.7% IRR" treats it as a promise. "Under base case assumptions, projected returns range from 12-16% IRR" is honest.

Failure mode: The investor comms agent drafted a year-end letter stating "Our portfolio returned 11.4% in 2025." Derek's audited number was 11.38%. The rounding was correct, but the letter didn't cite the audited source. Chen added "Based on audited Q4 2025 financials prepared by [Auditor Name]" to every performance figure.

Vetted Goods silver
C013 MEDIUM OBSERVED REPEATEDLY 4x efficiency

When using GPT for creative agents (Cadence, Chorus) and Claude for analytical agents (Pulse, Signal, Atlas, Ledger, Prism), maintain separate evaluation criteria. Creative agents are evaluated on brand voice consistency and engagement metrics. Analytical agents are evaluated on accuracy and signal-to-noise ratio. Never evaluate a creative agent on precision or an analytical agent on tone.

Why: The two platforms were chosen for different strengths. Evaluating both with the same rubric incentivizes the wrong behaviors -- GPT agents get over-optimized for accuracy (killing creativity) and Claude agents get prompted for engaging tone (introducing imprecision).

Failure mode: The team applied a single "quality score" rubric to all agents. Chorus (GPT, creative) scored low on "factual accuracy" because product descriptions included aspirational language. The team tried to make Chorus more precise, which killed the brand voice. Meanwhile, Pulse (Claude, analytical) scored low on "engaging presentation." The team added formatting requirements that made Pulse's alerts harder to scan quickly. Both agents got worse by being evaluated on the wrong criteria.

C014 HIGH OBSERVED REPEATEDLY 7x efficiency

Agent context switches between brands must include a "brand flush" step: clear the prior brand's context, load the new brand's configuration file, and confirm the brand identity in the output header. No "carry-over" operations where an agent finishes Brand A work and immediately starts Brand B work without context clearing.

Why: Context carry-over is the root cause of voice bleed, data leakage, and policy confusion. The 30-second cost of a brand flush is negligible compared to the cost of any cross-brand contamination incident.

Failure mode: The adventure-tee incident (C001), the email cross-contamination (C002), and the CS tone mismatch (C006) all traced back to context carry-over. Implementing mandatory brand flush reduced cross-brand incidents from 4-6 per month to 0-1 per month within the first 30 days.