86% lower cost.
85% recall.
| Method | Estimated cost | Wakes | Recall |
|---|---|---|---|
| Wake on every event | $39.49 | 334 | 100% |
| Keyword alert | $6.15 | 52 | 25% |
| Heartbeat (30 min) | $40.06 | 240 | 100% |
| jevable | $5.44 | 46 | 85% |
Method & full results
We replayed 334 events across five streams, treating each stream as one day. Costs use measured Claude Code wakes plus Jev calls. The heartbeat is assumed to catch everything; these are replay estimates.
Results by stream
| case | events (matter) | Wake on every event | Keyword alert | Check every 30 minutes | jevable |
|---|---|---|---|---|---|
| status | 35 (30) | 35 / 0 / 5 | 2 / 28 / 0 | 48 / 0 / 18 | 26 / 4 / 0 |
| hn-problems | 91 (7) | 91 / 0 / 84 | 10 / 5 / 8 | 48 / 0 / 41 | 5 / 2 / 0 |
| regressions | 97 (11) | 97 / 0 / 86 | 22 / 4 / 15 | 48 / 0 / 37 | 12 / 1 / 2 |
| releases | 11 (1) | 11 / 0 / 10 | 1 / 0 / 0 | 48 / 0 / 47 | 1 / 0 / 0 |
| review-asks | 100 (3) | 100 / 0 / 97 | 17 / 2 / 16 | 48 / 0 / 46 | 2 / 1 / 0 |
Per case: agent turns / missed / needless wakes.
Misses and unnecessary wakes
jevable missed 8 relevant events and triggered 2 unnecessary wakes.
7 misses scored between 0.45 and the 0.7 threshold. 1 were filtered out before reaching Jev.
- missedstatusService disruption on Claude services0.48
- missedstatusService disruption on Claude services0.48
- missedstatusElevated errors across many models0.52
- missedstatusElevated errors across many models0.52
- missedhn-problemsA few months ago Claude nuked one of our databases. We keep backups, so, "no big deal", but since then we just put "Never delete files, use the the …not asked
- missedhn-problems> In Claude Code, all plan mode does is add a little reminder to every user message I was quite surprised when I learned this (when Claude Code ed…0.65
- needlessregressions[BUG] 2.1.281: scrub sandbox still fails to create ancestor .claude stub when the ancestor is owned by another uid (as root, --cap-drop ALL)0.83
- needlessregressions[BUG] 2.1.281: CLAUDE_CODE_SUBPROCESS_ENV_SCRUB sandbox denies $HOME/actions-runner, so every Bash call fails on a default self-hosted GitHub Action…0.73
- missedregressions[GitHub integration]0.68
- missedreview-asksditto. tests are trivial and not providing useful coverage0.61
Method and limits.
Events came from status pages, Hacker News, GitHub issues, releases and PR reviews. A separate agent labelled them without seeing the rules or scores.
This is a small, model-labelled sample. Quiet streams make scheduled heartbeats expensive. Savings and recall will vary with your workload.
Single-event wakes averaged $0.12 across 8 runs ($0.10–$0.16). Batched wakes averaged $0.17 across 3 runs. Jev input costs $0.042 per million tokens.
Method
- Five cases, each a live source and a rule from
cases/; events fetched after the rules were last changed, never used to write any rule. - Labels by a separate agent that saw only each case's goal — not the rules, not any score. Events it was unsure about (12 of 346) are left out.
- Keyword regexes were written down before any results were seen.
- Wake cost measured with headless Claude Code (claude-opus-5-5[1m]): $0.13 to act on an event, $0.10 to decide it does not matter.
- Jev answers are real (jev-1.13.0 via TypeSafe).
Limits
- Measured wakes start a fresh session. Costs in a long-running session may differ as context grows.
- The labeller is a model, not a person; a few calls are arguable (listed above as needless wakes).
- 52 events that matter is a small sample; one case (review comments) had only 3.
- The heartbeat looks bad because these streams are small: it pays 48 checks a day however quiet the day is.
- jevable misses events. Near-threshold scores are the next thing to fix: collect them for a periodic look instead of dropping them.