Can a language model discover useful predictors and make useful
probabilistic predictions about outcomes that occurred after its documented knowledge cutoff
without a model-visible outcome? A convincing affirmative test also requires proof that every
input existed at forecast time. In a locally precommitted retrospective replay of 2026 Federal
Judicial Center civil-case records, gpt-5.6-sol proposed restricted feature plans; a
fixed logistic regression evaluated them on successive temporal blocks; and a fresh-context
model scored a later lockbox from direct-record or feature-packet views. The feature plan was not
admitted and was worse than raw fields on Lockbox. Both judge arms improved Brier loss over a
fixed historical prior, but feature packets did not improve on direct records and calibration did
not improve. The packets excluded explicit outcome fields, yet came from final-vintage records
with fields that may have changed after the forecast date. The result therefore shows signal in
retained outcome-field-excluded packets, not verified foresight. Earlier coded-data analyses are
exploratory because researchers specified their features. The corrected SCDB/Martin–Quinn
effect is small and mixed; an ECtHR run has 14.77% syntactic grounding failures and chance-level
AUC. Neither clean ex-ante utility nor adaptive feature-engineering utility is established.
Can an LLM improve a prediction system using only information demonstrably available before the outcome?
Retrospective success is easy to manufacture accidentally. Outcome clues can enter through document provenance; feature selection can overfit a reused validation set; recurring entities can cross nominal folds; and an LLM judge can favor artifacts produced by its own model family. Here, adaptive domain intelligence has a deliberately narrow meaning: a precommitted system proposes auditable representations of domain records and must improve out-of-sample probabilistic prediction under a frozen evaluation policy.
A proposal has no demonstrated utility if it fails held-out admission or its lockbox interval includes no improvement. A judge has no demonstrated utility unless it improves Brier loss over a fixed prior on every assigned case, with abstentions scored as prior forecasts. This makes null and negative outcomes part of the result rather than exceptions to it.
The completed replay satisfies explicit outcome-field exclusion at the packet boundary, but not the stronger point-in-time input requirement. Its result is therefore evidence about one fixed set of retained artifacts, not a completed demonstration of adaptive domain intelligence.
outcome-field-excluded final-vintage fields
│
├──► Sol feature engineer ─► 3 typed plans ─► Development selection
│ └► at most 2 children
│
├──► fixed logistic models ─► Admission gate ─► Lockbox classical score
│
└──► paired direct / feature packets ─► fresh Sol judge ─► committed predictions
│
└► one scoreThe source is a final-vintage FJC FY2021–2026 civil Integrated Database extract. Eligible
cases were filed by the model cutoff, terminated in a designated later window, and had a final
disposition outside a registered exclusion set. The Lockbox window contained 14,473 otherwise
valid terminations; that rule removed 1,977 and left 12,496 eligible cases. Sampling is
deterministic and unstratified only within this outcome-defined population; membership was
not knowable at cutoff. The target is DISP = 13, labeled “dismissed:
settled.” It is a scorable database code, but undercovers latent settlement and says nothing
about case merit.
| Cohort | Temporal definition | N | Role |
|---|---|---|---|
| Historical landmark | Pending 31 Dec. 2025; terminated 1 Jan.–16 Feb. 2026 | 5,000 | Fit comparators and estimate the reference prior |
| Development | Terminated 17–28 Feb. 2026 | 160 | Select roots and one bounded refinement round |
| Admission | Terminated 1–15 Mar. 2026 | 160 | One held-out admission decision |
| Lockbox | Terminated 16–31 Mar. 2026 | 160 | One final feature and judge evaluation |
Model packets contain circuit, opaque district, nature-of-suit bucket, jurisdiction basis, class-action allegation, jury demand, pro-se and in-forma-pauperis status, removal status, demanded amount or band, filing year and month, and age at cutoff. Except for recomputed age, these values were read from the later termination extract; they were not reconstructed from archived pending records. Party names, docket numbers, source row numbers, termination dates, disposition codes, judgment fields, procedural-progress codes, and later congestion measures are excluded. Opaque identifiers, no-tool execution, and trace audits reduce join-back risk.
Official FJC guidance says quarterly records can overwrite changed information and specifically identifies pro-se, in-forma-pauperis, and class-action fields as mutable during a case. The packet boundary therefore controls explicit outcome fields, not temporal vintage. The staging controller also loaded private answers in the same process as public records and used Historical answers to compute the fixed prior. Lockbox answers were not supplied to feature transformations, prompts, or provider calls, but label separation was behavioral rather than process-isolated.
The engineer proposes exactly three root plans in a deterministic DSL and may refine the best root with at most two children after seeing aggregate Development metrics only. Every plan uses the same fixed pipeline: dictionary vectorization, sparse standardization, and logistic regression with no hyperparameter or model-family search. Models refit on Historical + Development for one Admission test. Admission requires positive Brier improvement over the allowlisted raw-field baseline with a district-cluster 95% interval wholly above zero; otherwise the classical artifact falls back to the raw-field baseline. A plan replaces the 14 raw fields rather than augmenting them, so this tests a selected compressed representation, not the incremental value of appending engineered features.
Each lockbox case is then judged twice in separate fresh contexts: once from deterministic
field=value lines and once from the frozen submitted feature packet with exact source
lines. The judge receives the Historical strict-code prevalence as a prior but no Development or
Admission scores. Missing and abstained forecasts are assigned that prior. All reported paired
intervals use 10,000 district-cluster bootstrap draws. A feature-engineering or judge claim is
positive only when its prespecified improvement interval lies wholly above zero.
Before the final prediction set was committed or scored, Protocol Amendment 2 aligned the provider's strict schema, the already registered 40-case batches, and the prompt's exact quote-to-source contract. Two provider-rejected calls produced no forecasts; two one-case, outcome-blind smoke forecasts were retained but excluded from the final commitment and score. No packet, model, outcome, estimand, or decision rule changed.
The cells below were populated mechanically after all 320 forecasts were committed and the lockbox was scored once. Prompts, plans, models, estimands, and stopping rules did not change after prediction commitment.
| Quantity | Committed value |
|---|---|
| Historical strict-code prevalence | 0.2164 |
| Selected submitted feature plan | F2 |
| Development selection summary | F2 selected (Brier: F0 0.1204; F1 0.1374; F2 0.1097; F2A 0.1335; F2B 0.1235; F3 0.1605) |
| Admission ΔBrier over raw fields | -0.0080 |
| Admission district-cluster 95% CI | [-0.0263, +0.0097] |
| Admission decision | not admitted; classical fallback to F0 |
| Lockbox admitted-classical Brier | 0.1400 |
| Lockbox classical ΔBrier over raw fields | +0.0000 (F0 fallback) |
| Lockbox classical district-cluster 95% CI | not applicable (F0 fallback) |
| Direct-judge coverage | 100.0% |
| Direct-judge intention-to-forecast Brier | 0.1614 |
| Direct-judge ΔBrier over reference prior | +0.0272 |
| Direct-judge district-cluster 95% CI | [+0.0106, +0.0439] |
| Feature-judge coverage | 100.0% |
| Feature-judge intention-to-forecast Brier | 0.1624 |
| Feature-judge ΔBrier over reference prior | +0.0262 |
| Feature-judge district-cluster 95% CI | [+0.0120, +0.0400] |
| Feature-versus-direct paired ΔBrier | -0.0010 |
| Feature-versus-direct district-cluster 95% CI | [-0.0067, +0.0046] |
| Syntactic quote-grounding rate | 1227/1227 = 100.0% syntactic |
| Primary verdict | feature-engineering utility not demonstrated; both arm-vs-prior improvements observed; feature-over-direct gain not demonstrated; ex-ante utility not established |
F2 looked best on Development but reversed on Admission: Brier 0.1441 versus
0.1361 for raw-field F0, an improvement of −0.0080 (95% CI −0.0263 to
+0.0097). It was not admitted; because the interval crosses zero, the tri-state evidence policy
treats the result as inconclusive rather than a universal refutation. The submitted plan was again worse on Lockbox (Brier 0.1560 versus
0.1400; improvement −0.0160, CI −0.0395 to +0.0067).
In this retained sample, the direct judge achieved Brier 0.1614 and improved +0.0272 over the historical prior (95% CI +0.0106 to +0.0439). The feature-packet judge also beat the prior (+0.0262; CI +0.0120 to +0.0400), but was 0.0010 Brier worse than direct records; that paired interval included zero. Both arms had 100% coverage. The raw-field logistic model remained stronger at Brier 0.1400.
All 1,227 cited spans passed exact-match grounding, which establishes syntax rather than semantic entailment. Eight 40-case batches completed without a format retry or tool event in an 18.1-minute two-worker wall span. That records one fixed batched run; it does not establish stable or sustainable performance across provider revisions, batchings, or repeated executions.
A Claude model was registered as a cross-family replication, but its installed CLI reported the weekly quota exhausted at freeze. No Claude lockbox predictions were committed. The replication remains unrun; it will not be replaced by a post hoc model or presented as corroboration.
| Evidence class | Examples | Permitted inference |
|---|---|---|
| Primary outcome-withheld replay | FJC final-vintage feature and judge arms | Conditional proper-score comparison for the retained outcome-field-excluded packets; not a clean ex-ante forecast or feature-engineering success |
| Exploratory construct check | SCDB/Martin–Quinn, Carp-Manning, Songer, earlier coded FJC analyses | Whether researcher-coded variables behave broadly as expected under the stated split; not evidence that an LLM discovered them |
| Mechanism or diagnostic check | ECtHR extraction, planted faults, identity and split tests | Whether a component exposes a specified failure; not real-world predictive validity |
This hierarchy corrects a category error in earlier drafts. The SCDB, Carp-Manning, Songer, and legacy FJC schemas were hand-coded. They can test construct definitions and statistical plumbing, but they cannot show that an LLM is an effective feature engineer. Recovering a familiar association retrospectively is also not equivalent to forecasting a future outcome without access to it.
An earlier analysis assigned each Supreme Court term an ideology summary constructed from that same term's votes, contaminating the predictor with its target period. The repaired feature gives term t only the court median for term t − 1. Evaluation uses five complete, forward temporal folds across 9,092 cases and 79 terms; no term straddles training and evaluation.
| Temporal fold | ΔAUC |
|---|---|
| 1 | +0.0045 |
| 2 | −0.0014 |
| 3 | +0.0135 |
| 4 | +0.0340 |
| 5 | −0.0101 |
| Unweighted mean | +0.0081 |
Mean fold AUC rises from 0.5164 to 0.5245, but two folds are negative and the largest positive fold materially influences the mean. The appropriate verdict is small, mixed, and non-robust. Martin–Quinn scores are also retrospective dynamic estimates rather than archived release vintages. This remains an exploratory, researcher-coded construct check.
The ECtHR session experiment sampled 160 records and produced usable artifacts for 150. Across 1,178 committed feature values, 174 had missing or non-verbatim evidence: a 14.77% syntactic grounding-failure rate. Downstream AUC is 0.498, essentially chance; a rule comparator is likewise near chance at 0.490.
These measurements replace earlier claims of zero hallucination and successful grounded extraction. Exact substring matching cannot distinguish fabrication from formatting or evidence omission, and a passing quote may still be semantically irrelevant. An independent semantic audit is pending, so no faithfulness rate is claimed. The run also predates the repaired section-level provenance firewall and is a negative diagnostic, not a clean prospective forecast.
Earlier Carp-Manning, Songer, and FJC analyses remain useful for expressing familiar judicial-party, panel, repeat-player, and duration constructs. Their features were specified by researchers, however, and several designs retain limitations: substantial judge recurrence in Carp partitions, non-lead panel-member overlap in Songer, and a narrow target with older row-level inference in legacy FJC work. They are not used to validate LLM feature engineering, judge calibration, or a general “known-answer” claim.
Metis separates authority among a model-facing label-excluded feature engineer, deterministic extractor or transformer, fixed statistical estimator, fresh-context outcome judge, and frozen evaluator. Textual features declare their permitted document section; missing sections fail closed, and every committed value needs a verbatim span. Exact spans establish only syntactic presence. Experiment identity binds ordered case content, labels, grouping and time arrays, target, seed, evaluator, policy, backend, and model so a resumed run cannot silently inherit memory from a different task.
Adaptive work persists as a hypothesis tree with three evidence states: accepted, rejected, and inconclusive. Only genuine refutations become suppressive memory; uncertainty is retained as an audit receipt rather than rewritten as failure.
| Source mechanism | Status in Metis |
|---|---|
| RUC-NLPIR/Microsoft Research, Toward Generalist Autonomous Research via Hypothesis-Tree Refinement (arXiv:2606.11926) | Primary inspiration for persistent branching, evidence-conditioned expansion, and insight propagation; Metis adds tri-state memory and held-out admission |
| RUC long-lived coordinator, isolated executors, specialists, and serving runtime | Not reproduced as that runtime |
| AMD, Arbor: Tree Search as a Cognition Layer for Autonomous Agents (arXiv:2606.12563) | Related cognition-layer design; its profiler, dynamic specialists, critic, knowledge-base scoring, and GPU-serving stack are not ported |
| Typed legal policies, full experiment identity, section provenance, split validation, cluster-aware inference, and outcome-field-excluded judge packets | Metis extensions |
See the RUC project page and repository. The FJC replay tests Metis's bounded procedure; Arbor's existence is not evidence that this replay works.
DISP = 13 undercovers latent
settlement, and the replay conditions on both termination by 31 March 2026 and a realized final
disposition outside the exclusion set. Neither membership condition was known at cutoff. The run
does not estimate settlement propensity among all pending cases or support individual legal advice.Court records describe people and organizations in consequential disputes. This is a methods experiment, not a system for adjudication, bail, sentencing, immigration enforcement, case valuation, or legal representation. Any later application would require a task-specific legal basis, representative data, measurement validation, independent bias and security review, human authority, appeal and monitoring procedures, and evidence of benefit over non-automated practice.
This page reports an ongoing methods study. The completed replay records proper-score signal in outcome-field-excluded final-vintage packets, but does not establish clean ex-ante prediction or adaptive feature-engineering utility. Comments welcome — see the portfolio for contact.