Skip to content

Release 0.237

0.237.13: box-maintain.sh, and a bounded chunk-windows backfill

  • scripts/box-maintain.sh <box> <task> [args] runs a long pnpm maintain task in a throwaway sibling of mantle_web (same image and network, the web env through a pipe, its own memory limit, --rm, a mode-600 log file). It refuses a second run on the box. --status, --logs, --follow, --stop. pnpm maintain now always ends with one line: finished, FAILED with the exit code, or killed by a signal. docs/maintenance-runner.md, update-prod.md.
  • The chunk-windows backfill held a page of 500 chunks as JS arrays and one big JSON parameter, and was killed at --parallel=16. It now reads chunks as text, copies a one-window chunk in SQL, and writes the others in batches of about 100 windows. Measured peak memory at --parallel=16: 1,047 MB down to 486 MB.

0.237.11: one shared connection pool for every provider call

Node 26.5’s built-in fetch sends POSTs one at a time on a warm HTTP/2 session, so N parallel provider calls took N request times. providerFetch (packages/voice/src/adapters/provider-fetch.ts) is the built-in fetch with one shared undici Agent (HTTP/1.1 keep-alive, 32 connections per origin). Every voice adapter, the OpenRouter client and the decisions judge use it; 32 parallel POSTs went from 9.9 s to 0.24 s. The tailnet proxy loads undici the same way (its bare require threw under ESM). Embed calls back off on a 429 (2, 4, 8, 16 s). The windows backfill takes --parallel=N. docs/provider-http.md (new).

0.237.10: passage windows, a deeper judge pool, a parallel judge

  • Passage windows (opt in): each chunk also gets about 800-character sentence windows with their own vectors, and passage search adds a window arm that returns the window’s chunk, so the prompt budget does not change. Migration 0229 (0229_chunk_windows): embedding_config.chunk_windows (default false) and content_chunk_windows (no text, HNSW halfvec, RLS follows the node). pnpm maintain chunk-windows (dry run, --apply, --off, --clear) and eval:route --windows. On the library test corpus, paraphrased questions R@10 rose from 40% to 63%. docs/embeddings.md.
  • passage_scoring.pool goes up to 200 (was 100); with windows on the pool doubles.
  • The judge fan-out runs side by side (it ran one request after another under the built-in fetch). Pool 50: p50 1.39 s to 0.91 s.
  • docs/recall-eval.md, docs/decisions.md.

0.237.7: a keyword-found passage skips the cosine cutoff

The 0.65 cosine cutoff threw away literal matches that embed poorly (a code, a reference, a coined word) even when the keyword arm ranked them first. Under KEYWORD_PASSAGE_RULE = 'exempt' a passage with a keyword-arm rank is not held to the cutoff; its place and the chunk_limit cut are unchanged. The trace says exempt:keyword. Gated with eval:route. docs/recall-eval.md.

0.237.6: the context decision trace, and eval:route

  • Decision trace v1: every turn’s load_context snapshot carries trace (ContextTrace in @mantle/client-types): per stage the candidates in, kept, dropped and milliseconds; per candidate the block, key, the stage and reason code, which arm found it with its ranks, the distances and the judge score. The search_chunks step output carries the same trace. Observation only: the prompt is unchanged. Ids and codes, no text, 150 rows at most. docs/observability.md.
  • pnpm -C server/web eval:route runs a typed case set through named rulesets and reports R@1, R@10, MRR, latency and cost per question type, with a paired gate against the reference. Manual only; prints its cost.
  • Fix: recall_eval’s chunks line measured the vector arm alone; it is now the hybrid path agents use (chunksVector keeps the old number). docs/recall-eval.md.

0.237.4: an optional deeper pool for passage_scoring

uses.passage_scoring.pool (per brain; unset keeps the old pool) sets how many passages to fetch and score, up to 100, fanned out in requests of 25. With a pool set, auto-context scores before its budget cut even when context_pruning is on. eval:recall gains passage-scored. On the library test set, a pool of 50 lifted search_chunks R@10 from 43% to 53% at about three times the judge cost. docs/decisions.md.

0.237.3: the keyword arm speaks only on rare literals

On a large single-topic corpus the hybrid passage search scored below vector only: a question’s frame words outvoted its one rare word. The k-th rarest term now weighs idf * 0.5^k, question-frame words are dropped like chat filler, and the passage keyword arm returns rows only when they hold a rare term (gateRareTerms). Passage search p50 went from 150 ms to 14 ms. Node search is unchanged.

0.237.1: per-agent thinking effort

  • Migration 0228 (0228_agent_thinking_effort): agents.thinking_effort, nullable. NULL inherits the person’s profile setting, as before; off never reasons; a tier is that effort whatever the profile says.
  • One rule (resolveAgentThinking in content-core) for every turn path: web, Telegram, sim and resumed runs, team and client turns, delegated agents, heartbeats and run workers.
  • Read and set through GET/POST/PATCH /api/agents, Agent Studio, and agent_set_thinking_effort (new; no self-change, an agent asking waits at /pending). Shipped agents stay on inherit. docs/thinking.md.

0.237.0: passage-level recall eval, and the measured capacity policy

  • eval:recall gains passage retrievers scored on the exact chunk and on the document (passage, passage-vector, passage-keyword). Cases may name expectChunks and a group; --retrievers picks a subset.
  • corpusCapacity (the dashboard dial and brain_capacity) also returns retrieval: the passage recall@10 and MRR of the newest recall_eval run, or null. Optional on BrainCapacity.
  • Passage-vector policy: watch at 100k and split at 250k (was 50k and 100k), from a measured scale curve: recall@10 falls about 6 points per doubling, with no cliff. docs/recall-eval.md, “Scale curve”.