The DB soft Gherkin mutation re-ran the whole feature for every mutant; with a fresh LevelDB world per example, a 48-mutation feature exceeded 900s. The runner-worker now diffs the mutated IR against the base feature at <work>/base/feature.json (derived from the mutation path, since job.work_dir may be mutation-specific) and runs only the changed scenario/example. Safe because each scenario uses an isolated world. The same feature now completes in ~74s. By architect.
13 KiB
Architect Process Notes
ROLE-SCOPED — ARCHITECT ONLY. DO NOT FOLLOW.
This file is the architect role's private working notes. It records process exceptions, tooling behavior, and observations specific to how the architect runs its workflow. It is not shared guidance and is not intended for the specifier, coder, or refactorer roles. If you are not the architect, ignore this file entirely — do not treat anything here as a directive, convention, or requirement for your own role. Your role's instructions come only from your own role prompt and the constitution.
Durable notes on process exceptions, tooling behavior, and recurring
observations discovered while running the architect workflow. These are
process-level notes (how the tools behave, what to expect, what to watch for),
distinct from per-task verification results, which live in
docs/reviews/<task>-summary.md.
Tooling behavior / runtime
-
Pure modules can legitimately report 0 mutation sites.
mutate4javascriptonly targets arithmetic, comparison, equality, boolean, logical, and0<->1constant sites. A module built from!guards, ternaries, and template literals (e.g.block-explorer.js,like-result.js) scans asTotal mutation sites: 0withKilled: 0, Survived: 0. Run--scanto confirm the zero is structural and not a skipped/under-selected run before treating it as a pass. -
Run
dry4javascriptscoped to the changed files/dirs, not the whole client. A broadsrc test acceptancerun reports hundreds of pre-existing duplicate blocks (475 in the client on 2026-09-16) — mostlyacceptance/lib/handlers.jsstep-handler boilerplate and repeated older-suite test setup — which buries the one or two task-local candidates. Scope the run to the changed production files, tests, and adapters; the broad run is only a noise floor, consistent with prior reviews. -
Bare
dry4javascriptruns the full test suite (~1m50s). It is a DRY analysis that invokes tests, so running it with no arguments is a slow full-suite run, not a fast readiness probe. Never use it as a startup smoke check. Usedry4javascript --help(fast, ~0.6s) to confirm the binary is present and runnable.architect-startup.shand the startup cheat-sheet must both use--helpfor this check. -
architect-startup.shnow runs in ~5s. After thedry4javascript --helpfix, the only remaining cost is thegit fetchon the four tool repos (~3.7s). That fetch is a network call and can hang in sandboxed environments; if a hang is ever observed, add atimeoutto the fetch loop. The smoke test (mutate4javascript usage, dry4javascript --help, gherkin-parser --help) is the fast readiness probe; the full startup script is optional confirmation. -
Mutation runs dominate wall-clock time. Each
mutate4javascript <file>invocation runs the full test suite as a baseline (coverage refresh) before running mutations, then runs mutations in parallel with--max-workers 8. Because the baseline re-runs the whole suite, mutating N files costs roughly N full-suite runs. Plan for this: batch the affected files, run them sequentially, and use--max-workers 8to keep the mutation phase fast. The DRY and soft-Gherkin-mutation steps are comparatively quick. -
A mutated synchronous infinite loop can orphan a mocha process and wedge later runs. On 2026-09-17 a
mutate4javascriptworker onhelpers.js(stripLeadingEmptyPushes,1 -> 0) timed out but left ac8/mochaprocess at ~100% CPU for 45+ minutes; the next mutation baseline then hung behind it. The indexer and DBnpm testmocha has no--timeout, so a synchronous loop in mutated code cannot be interrupted by mocha — only the tool's--timeout-factor(default 10x baseline) applies. Defenses: run each mutation file under a shell-leveltimeout, redirect output to a file and grep theMutation Reportsection, and after any timeoutkill -9orphanedmocha/mutate4javascriptprocesses before the next file. Tool-written manifest updates in the source are expected; re-checkgit statusto confirm no mutant source remains applied. -
DB acceptance was the dominant verification cost; the test phase is now pooled. Before 2026-09-17,
npm run acceptancein psf-memo-db ran its 15 generated test files strictly sequentially (~430–790s; 145 scenarios, each opening/closing 21 LevelDB stores at ~850ms close apiece).acceptance.jsnow runs the generated files with bounded concurrency (defaultmin(4, files); override withACCEPTANCE_CONCURRENCY,=1restores the old sequential behavior). Measured 96s wall for the full DB suite. Generation still runs sequentially before the pooled test phase, so the constitution's "generation then tests" ordering is preserved. Pooling is safe because each generated file creates its own uniquely namedtmp/acceptance/level-*world. -
mutate4javascriptcopies the whole project into each worker, includingtmp/. The worker copy skips only.git,node_modules, andtarget. A staletmp/acceptance(LevelDB dirs from prior acceptance runs) can be ~1.5G, so with 8 workers the copy alone is ~12G of file I/O and the run appears to hang (process inDstate, no mutation progress lines). Before any mutation run,rm -rf <component>/tmp/acceptance target/mutation-workersto keep the worker copies tiny. This cut a notifications-query mutation run from 20+ minutes to ~2 minutes. -
Which directories bloat and how to keep them clean. The two directories that grow without bound are
<component>/tmp/acceptance(LevelDB dirs from acceptance runs, up to ~1.5G) and<component>/target/mutation-workers(per-run worker copies, up to ~14G). Both are gitignored build artifacts. Clean them before any mutation run:rm -rf psf-memo-db/tmp/acceptance psf-memo-db/target/mutation-workers \ psf-memo-client/tmp/acceptance psf-memo-client/target/mutation-workersarchitect-startup.shnow checks these and reports them as[FAIL]when they exceed a size threshold, so a bloated dir is caught before it slows a mutation run. -
memo-db.js(client HTTP adapter) is excluded from mutation testing. It uses ESM + a directory import (../config) that is only resolvable via react-scripts/webpack, so it cannot be loaded under plainnode --test. Its read behavior is exercised end-to-end via the DB acceptance tests. This is a standing precedent (also applied to the search task); do not attempt to force mutation coverage on it. -
Soft Gherkin acceptance mutation survivors are usually genuine equivalents. For read-only features, single-character case mutations of example values (addresses, text, txids) survive because each example value is used consistently on both the setup and assertion sides of its scenario. These are intrinsic equivalents, not implementation gaps; document them in the review summary and do not chase them.
-
DRY reports pre-existing pattern-boilerplate. The layered conventions (follow/mute/poll controllers, route-registration
index.js, memo-follow/ memo-mute services) produce score-1.00 duplicates that prior reviews left as-is. A shared controller base would be a broad cross-module refactor beyond any single handoff. Only reduce duplication that is local to the task at hand. -
mutate4javascriptdefault differential run can under-select after a function-set change. With a manifest present, the default run is differential (effectiveSinceLastRun) and selected only 1 changed site forpost-query.jsafter a review addedlistChildTxidsand removedbuildLikeCountMap;--scanalso reportedChanged mutation sites: 1. Re-running the same file with--mutate-allran all 37 sites (all killed). When a file adds or removes functions, run--mutate-allfor that file (or confirmSelected mutation sitesequalsTotal mutation sites) so new mutations are not silently skipped. -
Capture the
gherkin-mutatorreport with--jsonredirected to a file. The text report (write-text-report!, which usesprint) did not appear in the captured output before the tool'sSystem/exit; the JSON report (--json) did. Redirect stdout to a file and use--jsonfor a reliable, parseable record of killed/survived mutations. -
gherkin-mutatorwrites an emptyscenariosmanifest when every scenario has an intrinsic survivor.new-manifestrecords only scenarios withSurvived = 0andErrors = 0. When all mutations are single-character case changes of example values used consistently on both the setup and assertion sides, every scenario survives and the committed manifest is"scenarios":[]with no# mutation-stamp. This is the tool's expected output (those scenarios are intentionally re-mutated next run), not a partial write; commit it as-is and document the equivalents in the review summary. -
gherkin-mutatorcan run fromtmp/apswith absolute component paths. The Babashka task must run wherebb.edndefines it (tmp/aps), but the runner worker resolves the job'sfeature_json/work_dir/generated_dirpaths as given. Runcd tmp/aps && bb gherkin-mutator --feature <abs>/<component>/specs/x.feature --work-dir <abs>/<component>/build/acceptance-mutation --runner-worker "node <abs>/<component>/acceptance/lib/runner-worker.js" --level soft --workers 8 --status-interval 15s --json > report.json. The runner command is split on whitespace, so it must be a barenode <path>with no spaces in the path. The tool writes its manifest (and, when clean, a# mutation-stamp) into the feature file; commit that tool-written change. -
DB soft Gherkin mutation was impractically slow until the runner-worker narrowed its scope. Each mutant re-ran the whole feature (12 examples), and DB acceptance creates a fresh LevelDB world per example, so a 48-mutation feature took >900s (only 32 mutations reached after 15 min).
runner-worker.jsnow diffs the mutated IR against<work>/base/feature.json— derived from the mutation path<work>/mutations/<id>/feature.json, because the mutator'sjob.work_dirmay be the mutation-specific directory — and runs only the changed scenario/example. This is safe because every scenario uses an isolated world. The same 48-mutation feature now completes in ~74s. The indexer and client runner-workers still run the full feature; apply the same pattern if their soft runs get slow. -
mutate-file.shworks from every component dir, including the DB. It defaultsMUTATE4JS_BINto the client's installednode_modules/mutate4javascript, and the tool'snpm testbaseline runs in the current component, so the DB mutation runs use the DB suite without a second tool install.
Workflow observations
-
**
architect-startup.shcheckstmp/apsrelative to the worktree, butensure-aps.shresolves the single canonical APS checkout to the git common root'stmp/aps(the main checkout). In a role worktree the two disagree, so the startup script reports[FAIL] tmp/aps missingeven though the tools run. Local workaround: symlink the worktree's ignoredtmp/apsto the canonical checkout,ln -sfn <common-root>/tmp/aps tmp/aps. Proper fix (for a task that owns tooling): havearchitect-startup.shcaptureAPS_DIR="$(ensure-aps.sh --update)"and check/run against$APS_DIRinstead of the relative path. -
Review summaries must be force-added.
docs/is in the root.gitignore, sogit add -Asilently skipsdocs/reviews/<task>-summary.md. The role requires the summary to be committed with the byline in the same commit as the review changes, so usegit add -f docs/reviews/<task>-summary.md(orgit add -f docs/process-notes.md) before committing. This has silently dropped 8 of 13 summaries in the past; verify withgit ls-files docs/reviews/after committing. -
ready_for_next.sh/done_with_current.share the source of truth for queued work.done_with_current.shprintsNO_TASKwhen the queue is empty; stop waiting for work in that case. -
Handoff
commitfield must be exactly 10 hex chars.swarm_handoff.shrejects shorter abbreviations; usegit rev-parse --short=10 HEAD. -
Run per-component verification for every component a task touches before handing off (client, db, indexer), per the monorepo rules.
-
verify.shrecords one component per invocation, so multi-component tasks need multiple records. A task touching the client and the DB cannot use a single<task>-verification.json. The first recent two-component task (txid-wire-encoding) used the canonical<task>-verification.jsonfor the primary component and<task>-db-verification.jsonfor the second; both carry the same reviewgit_sha. State the mapping explicitly in the review summary, since the specifier's brief names only the canonical file.