Upgrades the default path to an observation-driven controller: HTTP model providers propose constrained candidates, the operator locks tasks and baselines, execution yields concrete counterevidence and revision obligations, and evidence authorizes representation changes.
Philosophy
Models propose and revise, experiment services execute, and independent evidence controls confirm; exploratory feedback stays separate from final confirmation.
Architecture & control flow
- 01
Load versioned frontier and task contracts and reserve budgets before calls
- 02
Validate strict candidates against parent, observation, source and fixed-metric references
- 03
Execute candidates and fixed baselines with equal resources, separating failures, insufficient measurement, counterevidence and support
- 04
Feed observations, repair contracts and difference memory to later rounds; measured counterevidence authorizes bounded reframing
- 05
Select a leader on the fixed primary suite, run local ablations and irreversibly freeze
- 06
Run one-shot fresh-holdout confirmation separately, with independent recomputation and optional program reexecution
Architecture outline derived from this version’s control flow.
What this version changes
Connects real model protocols, observations, revisions, reframing and separate final-confirmation authority in the default v6 control path.
Inputs & outputs
- Inputs
A topic, model configuration, operator-defined tasks/baselines/resources and literature; holdout data is supplied separately after freezing.
- Outputs
Candidates and observations, repairs and difference memory, exploratory archives, freeze contracts, confirmation statistics, event ledgers and readable reports.
Implementation & evidence scope
Attached results use a handwritten generator and synthetic tasks, validating data flow and rules rather than real-model scientific gains. External models and Docker were not exercised in the attached record; missing tasks produce IDEAS_ONLY.
Code & bundled material
The introduction draws on bundled notes, changelogs and central code. Software tests, synthetic diagnostics and scientific effectiveness use different evidence standards.
Source references
pyproject.toml· 1–5README.md· 1–7docs/v68/ARCHITECTURE.md· 3–36harness/discovery_v68/pipeline.py· 140–190harness/discovery_v68/pipeline.py· 195–240harness/discovery_v68/confirmation.py· 56–105docs/v68/VALIDATION.md· 9–28docs/v68/VALIDATION.md· 36–44harness/tests_v68/test_workflow.py· 16–36docs/v68/ARCHITECTURE.md· 68–88