Adaptive Domain Intelligence: A Falsifiable Protocol for LLM Feature Engineering and Outcome Judging

Dave Liu · working paper · 24 August 2026 · PDF

Working paper; outcome-withheld retrospective replay complete. In one fixed Sol run, both packet conditions had lower Brier loss than a frozen historical prior. The source was a final-vintage termination extract, however, and cohort membership depended on later termination and final disposition. This is conditional proper-score signal, not verified ex-ante foresight. The engineered feature plan was not admitted and did not improve the paired judge. The local freeze has no external timestamp, the cross-family replication is unrun, and the private artifact set has not been independently reproduced.

Abstract

Can a language model discover useful predictors and make useful probabilistic predictions about outcomes that occurred after its documented knowledge cutoff without a model-visible outcome? A convincing affirmative test also requires proof that every input existed at forecast time. In a locally precommitted retrospective replay of 2026 Federal Judicial Center civil-case records, gpt-5.6-sol proposed restricted feature plans; a fixed logistic regression evaluated them on successive temporal blocks; and a fresh-context model scored a later lockbox from direct-record or feature-packet views. The feature plan was not admitted and was worse than raw fields on Lockbox. Both judge arms improved Brier loss over a fixed historical prior, but feature packets did not improve on direct records and calibration did not improve. The packets excluded explicit outcome fields, yet came from final-vintage records with fields that may have changed after the forecast date. The result therefore shows signal in retained outcome-field-excluded packets, not verified foresight. Earlier coded-data analyses are exploratory because researchers specified their features. The corrected SCDB/Martin–Quinn effect is small and mixed; an ECtHR run has 14.77% syntactic grounding failures and chance-level AUC. Neither clean ex-ante utility nor adaptive feature-engineering utility is established.

1. The falsifiable question

Can an LLM improve a prediction system using only information demonstrably available before the outcome?

Retrospective success is easy to manufacture accidentally. Outcome clues can enter through document provenance; feature selection can overfit a reused validation set; recurring entities can cross nominal folds; and an LLM judge can favor artifacts produced by its own model family. Here, adaptive domain intelligence has a deliberately narrow meaning: a precommitted system proposes auditable representations of domain records and must improve out-of-sample probabilistic prediction under a frozen evaluation policy.

A proposal has no demonstrated utility if it fails held-out admission or its lockbox interval includes no improvement. A judge has no demonstrated utility unless it improves Brier loss over a fixed prior on every assigned case, with abstentions scored as prior forecasts. This makes null and negative outcomes part of the result rather than exceptions to it.

The completed replay satisfies explicit outcome-field exclusion at the packet boundary, but not the stronger point-in-time input requirement. Its result is therefore evidence about one fixed set of retained artifacts, not a completed demonstration of adaptive domain intelligence.

2. Outcome-withheld FJC final-vintage replay

outcome-field-excluded final-vintage fields
        │
        ├──► Sol feature engineer ─► 3 typed plans ─► Development selection
        │                                           └► at most 2 children
        │
        ├──► fixed logistic models ─► Admission gate ─► Lockbox classical score
        │
        └──► paired direct / feature packets ─► fresh Sol judge ─► committed predictions
                                                                  │
                                                                  └► one score

2.1 Cohorts and target

The source is a final-vintage FJC FY2021–2026 civil Integrated Database extract. Eligible cases were filed by the model cutoff, terminated in a designated later window, and had a final disposition outside a registered exclusion set. The Lockbox window contained 14,473 otherwise valid terminations; that rule removed 1,977 and left 12,496 eligible cases. Sampling is deterministic and unstratified only within this outcome-defined population; membership was not knowable at cutoff. The target is DISP = 13, labeled “dismissed: settled.” It is a scorable database code, but undercovers latent settlement and says nothing about case merit.

Replay cohorts. Successive temporal blocks prevent Development feedback from reaching Admission or Lockbox.
CohortTemporal definitionNRole
Historical landmarkPending 31 Dec. 2025; terminated 1 Jan.–16 Feb. 20265,000Fit comparators and estimate the reference prior
DevelopmentTerminated 17–28 Feb. 2026160Select roots and one bounded refinement round
AdmissionTerminated 1–15 Mar. 2026160One held-out admission decision
LockboxTerminated 16–31 Mar. 2026160One final feature and judge evaluation

2.2 Information firewall

Model packets contain circuit, opaque district, nature-of-suit bucket, jurisdiction basis, class-action allegation, jury demand, pro-se and in-forma-pauperis status, removal status, demanded amount or band, filing year and month, and age at cutoff. Except for recomputed age, these values were read from the later termination extract; they were not reconstructed from archived pending records. Party names, docket numbers, source row numbers, termination dates, disposition codes, judgment fields, procedural-progress codes, and later congestion measures are excluded. Opaque identifiers, no-tool execution, and trace audits reduce join-back risk.

Official FJC guidance says quarterly records can overwrite changed information and specifically identifies pro-se, in-forma-pauperis, and class-action fields as mutable during a case. The packet boundary therefore controls explicit outcome fields, not temporal vintage. The staging controller also loaded private answers in the same process as public records and used Historical answers to compute the fixed prior. Lockbox answers were not supplied to feature transformations, prompts, or provider calls, but label separation was behavioral rather than process-isolated.

2.3 Feature tree, judge, and decision rules

The engineer proposes exactly three root plans in a deterministic DSL and may refine the best root with at most two children after seeing aggregate Development metrics only. Every plan uses the same fixed pipeline: dictionary vectorization, sparse standardization, and logistic regression with no hyperparameter or model-family search. Models refit on Historical + Development for one Admission test. Admission requires positive Brier improvement over the allowlisted raw-field baseline with a district-cluster 95% interval wholly above zero; otherwise the classical artifact falls back to the raw-field baseline. A plan replaces the 14 raw fields rather than augmenting them, so this tests a selected compressed representation, not the incremental value of appending engineered features.

Each lockbox case is then judged twice in separate fresh contexts: once from deterministic field=value lines and once from the frozen submitted feature packet with exact source lines. The judge receives the Historical strict-code prevalence as a prior but no Development or Admission scores. Missing and abstained forecasts are assigned that prior. All reported paired intervals use 10,000 district-cluster bootstrap draws. A feature-engineering or judge claim is positive only when its prespecified improvement interval lies wholly above zero.

Before the final prediction set was committed or scored, Protocol Amendment 2 aligned the provider's strict schema, the already registered 40-case batches, and the prompt's exact quote-to-source contract. Two provider-rejected calls produced no forecasts; two one-case, outcome-blind smoke forecasts were retained but excluded from the final commitment and score. No packet, model, outcome, estimand, or decision rule changed.

3. Committed results: signal, not ex-ante proof

The cells below were populated mechanically after all 320 forecasts were committed and the lockbox was scored once. Prompts, plans, models, estimands, and stopping rules did not change after prediction commitment.

FJC replay. Brier improvements are oriented so positive values are better.
QuantityCommitted value
Historical strict-code prevalence0.2164
Selected submitted feature planF2
Development selection summaryF2 selected (Brier: F0 0.1204; F1 0.1374; F2 0.1097; F2A 0.1335; F2B 0.1235; F3 0.1605)
Admission ΔBrier over raw fields-0.0080
Admission district-cluster 95% CI[-0.0263, +0.0097]
Admission decisionnot admitted; classical fallback to F0
Lockbox admitted-classical Brier0.1400
Lockbox classical ΔBrier over raw fields+0.0000 (F0 fallback)
Lockbox classical district-cluster 95% CInot applicable (F0 fallback)
Direct-judge coverage100.0%
Direct-judge intention-to-forecast Brier0.1614
Direct-judge ΔBrier over reference prior+0.0272
Direct-judge district-cluster 95% CI[+0.0106, +0.0439]
Feature-judge coverage100.0%
Feature-judge intention-to-forecast Brier0.1624
Feature-judge ΔBrier over reference prior+0.0262
Feature-judge district-cluster 95% CI[+0.0120, +0.0400]
Feature-versus-direct paired ΔBrier-0.0010
Feature-versus-direct district-cluster 95% CI[-0.0067, +0.0046]
Syntactic quote-grounding rate1227/1227 = 100.0% syntactic
Primary verdictfeature-engineering utility not demonstrated; both arm-vs-prior improvements observed; feature-over-direct gain not demonstrated; ex-ante utility not established

F2 looked best on Development but reversed on Admission: Brier 0.1441 versus 0.1361 for raw-field F0, an improvement of −0.0080 (95% CI −0.0263 to +0.0097). It was not admitted; because the interval crosses zero, the tri-state evidence policy treats the result as inconclusive rather than a universal refutation. The submitted plan was again worse on Lockbox (Brier 0.1560 versus 0.1400; improvement −0.0160, CI −0.0395 to +0.0067).

In this retained sample, the direct judge achieved Brier 0.1614 and improved +0.0272 over the historical prior (95% CI +0.0106 to +0.0439). The feature-packet judge also beat the prior (+0.0262; CI +0.0120 to +0.0400), but was 0.0010 Brier worse than direct records; that paired interval included zero. Both arms had 100% coverage. The raw-field logistic model remained stronger at Brier 0.1400.

All 1,227 cited spans passed exact-match grounding, which establishes syntax rather than semantic entailment. Eight 40-case batches completed without a format retry or tool event in an 18.1-minute two-worker wall span. That records one fixed batched run; it does not establish stable or sustainable performance across provider revisions, batchings, or repeated executions.

A Claude model was registered as a cross-family replication, but its installed CLI reported the weekly quota exhausted at freeze. No Claude lockbox predictions were committed. The replication remains unrun; it will not be replaced by a post hoc model or presented as corroboration.

4. Evidence hierarchy

Different artifacts support different inferences.
Evidence classExamplesPermitted inference
Primary outcome-withheld replayFJC final-vintage feature and judge armsConditional proper-score comparison for the retained outcome-field-excluded packets; not a clean ex-ante forecast or feature-engineering success
Exploratory construct checkSCDB/Martin–Quinn, Carp-Manning, Songer, earlier coded FJC analysesWhether researcher-coded variables behave broadly as expected under the stated split; not evidence that an LLM discovered them
Mechanism or diagnostic checkECtHR extraction, planted faults, identity and split testsWhether a component exposes a specified failure; not real-world predictive validity

This hierarchy corrects a category error in earlier drafts. The SCDB, Carp-Manning, Songer, and legacy FJC schemas were hand-coded. They can test construct definitions and statistical plumbing, but they cannot show that an LLM is an effective feature engineer. Recovering a familiar association retrospectively is also not equivalent to forecasting a future outcome without access to it.

5. Corrected and negative evidence

5.1 SCDB/Martin–Quinn: small, mixed, non-robust

An earlier analysis assigned each Supreme Court term an ideology summary constructed from that same term's votes, contaminating the predictor with its target period. The repaired feature gives term t only the court median for term t − 1. Evaluation uses five complete, forward temporal folds across 9,092 cases and 79 terms; no term straddles training and evaluation.

Corrected prior-term ideology effect. ΔAUC from adding the lagged court summary.
Temporal foldΔAUC
1+0.0045
2−0.0014
3+0.0135
4+0.0340
5−0.0101
Unweighted mean+0.0081

Mean fold AUC rises from 0.5164 to 0.5245, but two folds are negative and the largest positive fold materially influences the mean. The appropriate verdict is small, mixed, and non-robust. Martin–Quinn scores are also retrospective dynamic estimates rather than archived release vintages. This remains an exploratory, researcher-coded construct check.

5.2 ECtHR: syntactic failures and chance prediction

The ECtHR session experiment sampled 160 records and produced usable artifacts for 150. Across 1,178 committed feature values, 174 had missing or non-verbatim evidence: a 14.77% syntactic grounding-failure rate. Downstream AUC is 0.498, essentially chance; a rule comparator is likewise near chance at 0.490.

These measurements replace earlier claims of zero hallucination and successful grounded extraction. Exact substring matching cannot distinguish fabrication from formatting or evidence omission, and a passing quote may still be semantically irrelevant. An independent semantic audit is pending, so no faithfulness rate is claimed. The run also predates the repaired section-level provenance firewall and is a negative diagnostic, not a clean prospective forecast.

5.3 Other coded replications

Earlier Carp-Manning, Songer, and FJC analyses remain useful for expressing familiar judicial-party, panel, repeat-player, and duration constructs. Their features were specified by researchers, however, and several designs retain limitations: substantial judge recurrence in Carp partitions, non-lead panel-member overlap in Songer, and a narrow target with older row-level inference in legacy FJC work. They are not used to validate LLM feature engineering, judge calibration, or a general “known-answer” claim.

6. Architecture and Arbor lineage

Metis separates authority among a model-facing label-excluded feature engineer, deterministic extractor or transformer, fixed statistical estimator, fresh-context outcome judge, and frozen evaluator. Textual features declare their permitted document section; missing sections fail closed, and every committed value needs a verbatim span. Exact spans establish only syntactic presence. Experiment identity binds ordered case content, labels, grouping and time arrays, target, seed, evaluator, policy, backend, and model so a resumed run cannot silently inherit memory from a different task.

Adaptive work persists as a hypothesis tree with three evidence states: accepted, rejected, and inconclusive. Only genuine refutations become suppressive memory; uncertainty is retained as an audit receipt rather than rewritten as failure.

Two different Arbor projects. The name alone is ambiguous.
Source mechanismStatus in Metis
RUC-NLPIR/Microsoft Research, Toward Generalist Autonomous Research via Hypothesis-Tree Refinement (arXiv:2606.11926)Primary inspiration for persistent branching, evidence-conditioned expansion, and insight propagation; Metis adds tri-state memory and held-out admission
RUC long-lived coordinator, isolated executors, specialists, and serving runtimeNot reproduced as that runtime
AMD, Arbor: Tree Search as a Cognition Layer for Autonomous Agents (arXiv:2606.12563)Related cognition-layer design; its profiler, dynamic specialists, critic, knowledge-base scoring, and GPU-serving stack are not ported
Typed legal policies, full experiment identity, section provenance, split validation, cluster-aware inference, and outcome-field-excluded judge packetsMetis extensions

See the RUC project page and repository. The FJC replay tests Metis's bounded procedure; Arbor's existence is not evidence that this replay works.

7. Interpretation, limitations, and ethics

Next decisive test

Prospectively commit an unresolved cohort before outcomes exist; define eligibility without final-disposition information; snapshot every model-visible field at the forecast date; isolate labels in a separate scoring process; and precommit independent, repeated judge runs, including a cross-family replication. Score once after the forecast horizon. That design would test adaptive utility rather than only conditional signal in retained artifacts.

Court records describe people and organizations in consequential disputes. This is a methods experiment, not a system for adjudication, bail, sentencing, immigration enforcement, case valuation, or legal representation. Any later application would require a task-specific legal basis, representative data, measurement validation, independent bias and security review, human authority, appeal and monitoring procedures, and evidence of benefit over non-automated practice.

Selected references

  1. Jin et al. (2026). Toward Generalist Autonomous Research via Hypothesis-Tree Refinement. arXiv:2606.11926.
  2. Prakriya et al. (2026). Arbor: Tree Search as a Cognition Layer for Autonomous Agents. arXiv:2606.12563.
  3. McInerney et al. (2023). CHiLL: Zero-shot Custom Interpretable Feature Extraction from Clinical Notes. Findings of ACL.
  4. Han, Yoon, Arik & Pfister (2024). FeatLLM: LLMs Can Automatically Engineer Features for Few-Shot Tabular Learning. ICML.
  5. Huang et al. (2023/2024). Large Language Models Cannot Self-Correct Reasoning Yet. ICLR.
  6. Kamoi et al. (2024). When Can LLMs Actually Correct Their Own Mistakes? TACL 12.
  7. Panickssery, Bowman & Feng (2024). LLM Evaluators Recognize and Favor Their Own Generations. NeurIPS.
  8. Aletras et al. (2016). Predicting Judicial Decisions of the European Court of Human Rights. PeerJ Computer Science 2:e93.
  9. Medvedeva, Wieling & Vols (2023). Rethinking the Field of Automatic Prediction of Court Decisions. Artificial Intelligence and Law 31.
  10. Martin & Quinn (2002). Dynamic Ideal Point Estimation for the U.S. Supreme Court. Political Analysis 10(2).

This page reports an ongoing methods study. The completed replay records proper-score signal in outcome-field-excluded final-vintage packets, but does not establish clean ex-ante prediction or adaptive feature-engineering utility. Comments welcome — see the portfolio for contact.