Release 0.237
0.237.13: box-maintain.sh, and a bounded chunk-windows backfill
scripts/box-maintain.sh <box> <task> [args]runs a longpnpm maintaintask in a throwaway sibling ofmantle_web(same image and network, the web env through a pipe, its own memory limit,--rm, a mode-600 log file). It refuses a second run on the box.--status,--logs,--follow,--stop.pnpm maintainnow always ends with one line: finished, FAILED with the exit code, or killed by a signal. docs/maintenance-runner.md, update-prod.md.- The chunk-windows backfill held a page of 500 chunks as JS arrays and
one big JSON parameter, and was killed at
--parallel=16. It now reads chunks as text, copies a one-window chunk in SQL, and writes the others in batches of about 100 windows. Measured peak memory at--parallel=16: 1,047 MB down to 486 MB.
0.237.11: one shared connection pool for every provider call
Node 26.5’s built-in fetch sends POSTs one at a time on a warm HTTP/2
session, so N parallel provider calls took N request times. providerFetch
(packages/voice/src/adapters/provider-fetch.ts) is the built-in fetch with
one shared undici Agent (HTTP/1.1 keep-alive, 32 connections per origin).
Every voice adapter, the OpenRouter client and the decisions judge use it;
32 parallel POSTs went from 9.9 s to 0.24 s. The tailnet proxy loads undici
the same way (its bare require threw under ESM). Embed calls back off on a
429 (2, 4, 8, 16 s). The windows backfill takes --parallel=N.
docs/provider-http.md (new).
0.237.10: passage windows, a deeper judge pool, a parallel judge
- Passage windows (opt in): each chunk also gets about 800-character
sentence windows with their own vectors, and passage search adds a window
arm that returns the window’s chunk, so the prompt budget does not change.
Migration 0229 (
0229_chunk_windows):embedding_config.chunk_windows(default false) andcontent_chunk_windows(no text, HNSW halfvec, RLS follows the node).pnpm maintain chunk-windows(dry run,--apply,--off,--clear) andeval:route --windows. On the library test corpus, paraphrased questions R@10 rose from 40% to 63%. docs/embeddings.md. passage_scoring.poolgoes up to 200 (was 100); with windows on the pool doubles.- The judge fan-out runs side by side (it ran one request after another under the built-in fetch). Pool 50: p50 1.39 s to 0.91 s.
- docs/recall-eval.md, docs/decisions.md.
0.237.7: a keyword-found passage skips the cosine cutoff
The 0.65 cosine cutoff threw away literal matches that embed poorly (a code,
a reference, a coined word) even when the keyword arm ranked them first.
Under KEYWORD_PASSAGE_RULE = 'exempt' a passage with a keyword-arm rank is
not held to the cutoff; its place and the chunk_limit cut are unchanged.
The trace says exempt:keyword. Gated with eval:route.
docs/recall-eval.md.
0.237.6: the context decision trace, and eval:route
- Decision trace v1: every turn’s
load_contextsnapshot carriestrace(ContextTracein @mantle/client-types): per stage the candidates in, kept, dropped and milliseconds; per candidate the block, key, the stage and reason code, which arm found it with its ranks, the distances and the judge score. Thesearch_chunksstep output carries the same trace. Observation only: the prompt is unchanged. Ids and codes, no text, 150 rows at most. docs/observability.md. pnpm -C server/web eval:routeruns a typed case set through named rulesets and reports R@1, R@10, MRR, latency and cost per question type, with a paired gate against the reference. Manual only; prints its cost.- Fix:
recall_eval’schunksline measured the vector arm alone; it is now the hybrid path agents use (chunksVectorkeeps the old number). docs/recall-eval.md.
0.237.4: an optional deeper pool for passage_scoring
uses.passage_scoring.pool (per brain; unset keeps the old pool) sets how
many passages to fetch and score, up to 100, fanned out in requests of 25.
With a pool set, auto-context scores before its budget cut even when
context_pruning is on. eval:recall gains passage-scored. On the
library test set, a pool of 50 lifted search_chunks R@10 from 43% to 53%
at about three times the judge cost. docs/decisions.md.
0.237.3: the keyword arm speaks only on rare literals
On a large single-topic corpus the hybrid passage search scored below
vector only: a question’s frame words outvoted its one rare word. The k-th
rarest term now weighs idf * 0.5^k, question-frame words are dropped like
chat filler, and the passage keyword arm returns rows only when they hold a
rare term (gateRareTerms). Passage search p50 went from 150 ms to 14 ms.
Node search is unchanged.
0.237.1: per-agent thinking effort
- Migration 0228 (
0228_agent_thinking_effort):agents.thinking_effort, nullable. NULL inherits the person’s profile setting, as before;offnever reasons; a tier is that effort whatever the profile says. - One rule (
resolveAgentThinkingin content-core) for every turn path: web, Telegram, sim and resumed runs, team and client turns, delegated agents, heartbeats and run workers. - Read and set through
GET/POST/PATCH /api/agents, Agent Studio, andagent_set_thinking_effort(new; no self-change, an agent asking waits at /pending). Shipped agents stay on inherit. docs/thinking.md.
0.237.0: passage-level recall eval, and the measured capacity policy
eval:recallgains passage retrievers scored on the exact chunk and on the document (passage,passage-vector,passage-keyword). Cases may nameexpectChunksand a group;--retrieverspicks a subset.corpusCapacity(the dashboard dial andbrain_capacity) also returnsretrieval: the passage recall@10 and MRR of the newestrecall_evalrun, or null. Optional onBrainCapacity.- Passage-vector policy: watch at 100k and split at 250k (was 50k and 100k), from a measured scale curve: recall@10 falls about 6 points per doubling, with no cliff. docs/recall-eval.md, “Scale curve”.