Generated 2026-06-22T02:13:06.464Z · window 30d (2026-05-23T02:13:06.464Z → 2026-06-22T02:13:06.464Z) · schema v1.1
Gate 1 OK, Gate 2 7/7 MUST + 0/3 SHOULD
| Criterion | Status | Current | Target |
|---|---|---|---|
| 1. Extraction works | PASS | 81 live organic (0 extract-stored events, 81 total organic entries) | >= 5 organic entries |
| 2. Dedup / hygiene works | PASS | OK | 0 exact duplicates, 0 placeholder entries |
| 3. Interception fires | PASS | 1013 this week | >= 10/week |
| 5. Non-blocking | PASS | OK (timeout enforced) | 100% < 3s |
| 7. Evolution works | PASS | 1 T0 + 6 probationary T1 + 2 mature T1 | >= 1 principle |
| 8. Memory shrinks | PASS | Yes — evolution reduced entries | Entries decrease after evolution |
| 9. Novel coverage | PASS | 1 principle(s) with novel-case evidence (1/1 holdout matches) | >= 1 principle matches unseen case |
| Criterion | Status | Current | Target |
|---|---|---|---|
| 4. Interception accurate | FAIL | 0% precision (0/2 classified surfaced hints, 1/1013 surfaced total) — insufficient sample (need >= 20), pending | >= 70% of classified surfaced hints are relevant (need >= 20 classified) |
| 6. Error recurrence drops | FAIL | Insufficient data (need 2+ weeks of mistake telemetry) | >= 30% reduction |
| 10. Cost stable | FAIL | Insufficient data (need 2+ weeks of cost telemetry) | No material increase (>10%) |
| Q1 mistake avoidance | PENDING |
| Q2 novel coverage | PENDING |
| Q3 memory shrinks | PENDING |
| Q4 auto-narrow | PENDING |
Precision over events since the pre-surface relevance gate went live (2026-06-19T06:59:22.531Z). Sidesteps the 7-day rolling window so a fix is visible in ~1–2 days.
| Precision | 78.9% |
| Relevant | 30 |
| Irrelevant | 8 |
| Classified | 38 / 30 min |
| Gate drops | 23566 (2669 events) |
| Reason | Count |
|---|---|
| wrong_task | 6 |
| wrong_language | 2 |
| Tier | Count | % |
|---|---|---|
| T0 — New (never surfaced) | 145 | 32.6% |
| T1 — Bootstrap (surfaced 1-3x) | 70 | 15.7% |
| T2 — Active (has follows) | 42 | 9.4% |
| T3 — Dying (surfaced >3x, 0 follows) | 188 | 42.2% |
| Type | Count | % |
|---|---|---|
| runtime | 164 | 36.9% |
| other | 112 | 25.2% |
| user-correction | 77 | 17.3% |
| test | 36 | 8.1% |
| log | 33 | 7.4% |
| review | 22 | 4.9% |
| workflow | 1 | 0.2% |
| Metric | Coverage | % |
|---|---|---|
| project_slug | 387/445 | 87.0% |
| structured conditions | 269/445 | 60.4% |
| lang (not "all") | 279/445 | 62.7% |
| judgment | 445/445 | 100.0% |
| Collection | Total | T0 | T1 | T2 | T3 | Top type |
|---|---|---|---|---|---|---|
| principles | 1 | 0 | 0 | 1 | 0 | workflow (1) |
| behavioral | 54 | 2 | 4 | 12 | 36 | user-correction (41) |
| selfqa | 390 | 143 | 66 | 29 | 152 | runtime (155) |
| Band | Surfaced | Followed | Ignored | Noise | Precision |
|---|---|---|---|---|---|
| 0.50-0.65 | 11 | 4 | 0 | 0 | 100.0% |
| 0.65-0.70 | 4 | 1 | 0 | 0 | 100.0% |
| 0.70-0.75 | 4 | 3 | 1 | 0 | 100.0% |
| 0.75-0.80 | 180 | 1 | 1 | 0 | 100.0% |
| 0.85-1.00 | 30 | 4 | 0 | 0 | 100.0% |
| Collection | Surfaced | Followed | Ignored | Noise | Precision |
|---|---|---|---|---|---|
| experience-behavioral | 229 | 3 | 1 | 0 | 100.0% |
| experience-principles | 30 | 4 | 0 | 0 | 100.0% |
| experience-selfqa | 16 | 11 | 1 | 0 | 100.0% |
| static-rules | 0 | 0 | 0 | 4 | 0.0% |
| Framework | Surfaced | Followed | Ignored | Noise | Precision |
|---|---|---|---|---|---|
| unknown | 230 | 9 | 2 | 4 | 73.3% |
| any | 44 | 9 | 0 | 0 | 100.0% |
| dotnet | 1 | 0 | 0 | 0 | — |
| Runtime | Surfaced | Followed | Ignored | Noise | Precision |
|---|---|---|---|---|---|
| api | 265 | 5 | 2 | 0 | 100.0% |
| codex-windows | 10 | 10 | 0 | 0 | 100.0% |
| claude-code | 0 | 0 | 0 | 0 | — |
| unknown | 0 | 3 | 0 | 4 | 42.9% |
| Reason | Count |
|---|---|
| wrong_task | 3 |
| wrong_language | 1 |
| Window | Surfaced | Followed | Ignored | Noise | No response |
|---|---|---|---|---|---|
| 7d | 275 | 18 | 2 | 4 | 251 |
| 30d | 275 | 18 | 2 | 4 | 251 |
| ID | Coll | Tier | Conf | Surf | Ignore% | Framework | Last noise reasons | Principle |
|---|---|---|---|---|---|---|---|---|
| 688832b7 | selfqa | T2 | 50.0% | 15 | 100.0% | dotnet | — | Attempting to load and inspect types from a .NET assembly wi… |
| 44cda635 | selfqa | T2 | 50.0% | 12 | 100.0% | any | — | Using Windows-specific shell commands in a Unix-like shell e… |
| abf24d06 | selfqa | T2 | 50.0% | 8 | 100.0% | react | — | When modifying keyboard event handling in a React TypeScript… |
| cd156f32 | selfqa | T2 | 50.0% | 8 | 100.0% | any | — | Running a command-line tool (hyperfine) via Bash in a mixed … |
| 8fb23450 | selfqa | T2 | 65.0% | 6 | 100.0% | react | — | When modifying process exit handling in a TypeScript CLI app… |
| 764ed106 | selfqa | T2 | 54.0% | 14 | 92.9% | any | — | Attempting to run a TypeScript type check using a direct com… |
| a841d268 | selfqa | T2 | 54.0% | 11 | 90.9% | any | — | Attempting to execute PowerShell commands in a Bash shell en… |
| 5195f824 | selfqa | T2 | 65.0% | 5 | 80.0% | any | — | When running a remote grep via SSH to check if a patch has b… |
| ccc17478 | behavioral | T1 | 60.0% | 5 | 20.0% | any | — | When reviewing gap analysis for a Next.js frontend project w… |
| 6bb6f392 | behavioral | T1 | 68.0% | 8 | 0.0% | any | — | Running git commands via Bash on a Windows filesystem mounte… |
| de45ba11 | behavioral | T1 | 68.0% | 7 | 0.0% | any | — | Running a Python script that imports a missing module |
| a1acbfed | selfqa | T2 | 68.0% | 5 | 0.0% | any | — | Copying and executing a Python script in a Docker container … |
Per-session hint trace. Click a session row to expand its hints, click a hint to see the tool action(s) that triggered it.