51.9% of applicable items were fully reported
Each of the 133 eligible papers was assessed against 32 reporting items derived from Princeton's REFORMS. Of 4,150 applicable item judgments, 2,153 were fully reported, 1,028 were partly reported, and 969 were absent or unclear. The secondary score gives half credit to partial reporting.
The 51.9% figure is an item-level reporting rate; papers were not assigned pass/fail results. Every eligible paper fully reported at least 2 items.
133 candidates met the full scoring criteria
The candidate frame includes papers whose titles or abstracts use explicit ML language. Screening then asks whether model performance or output is used as evidence for a scientific claim, and whether a complete original paper is available.
Eligibility outcome
| Outcome | Papers | Share |
|---|---|---|
| Eligible full ML-based-science papers | 133 | 29.2% |
| Not ML-based science | 317 | 69.5% |
| ML-based science, but no complete paper body | 4 | 0.9% |
| Unresolved | 2 | 0.4% |
Scientific-use classification
| Class | Papers | Share |
|---|---|---|
| ML-based science | 137 | 30.0% |
| ML mentioned, but not used as evidence | 103 | 22.6% |
| ML methods research | 86 | 18.9% |
| Predictive analytics | 83 | 18.2% |
| Non-ML quantitative analysis | 45 | 9.9% |
| Unresolved | 2 | 0.4% |
The 137 ML-based-science candidates comprise 133 complete original papers and 4 records without a complete paper body, which could not be scored against the full reporting standard.
Publication form in the candidate set
Publication form is recorded separately from scientific use. This keeps reviews, editorials, responses, and conference abstracts visible as sample outcomes without treating them as complete empirical papers.
| Publication form | Papers | Share |
|---|---|---|
| Original Research | 345 | 75.7% |
| Review | 47 | 10.3% |
| Editorial Opinion | 17 | 3.7% |
| Systematic Review | 15 | 3.3% |
| Other | 10 | 2.2% |
| Conference Abstract | 9 | 2.0% |
| Commentary Response | 5 | 1.1% |
| Bibliometric Review | 4 | 0.9% |
| Meta Analysis | 2 | 0.4% |
| Protocol | 1 | 0.2% |
| Uncertain | 1 | 0.2% |
Pre/post comparison
The currently eligible papers include 63 before REFORMS and 70 after it. Full reporting was 52.0% before and 51.8% after, a change of -0.2 percentage points. The half-credit score changed by -0.6 points. These figures remain descriptive: the comparison contains 133 papers, and human validation has not started.
| Period | Papers | Applicable items | Present | Partial | Absent or unclear | Full reporting | Half-credit score |
|---|---|---|---|---|---|---|---|
| Before REFORMS | 63 | 1,965 | 1,022 | 493 | 450 | 52.0% | 64.6% |
| After REFORMS | 70 | 2,185 | 1,131 | 535 | 519 | 51.8% | 64.0% |
Paper-level reporting varied substantially
Every eligible paper reported at least some REFORMS-aligned information. The table treats each paper equally; the preceding comparison treats each applicable item equally.
| Period | Papers | Mean paper score | Median | Range | Papers with zero present items |
|---|---|---|---|---|---|
| Before REFORMS | 63 | 52.0% | 53.1% | 19.4% to 93.8% | 0 |
| After REFORMS | 70 | 51.7% | 53.2% | 6.2% to 90.3% | 0 |
Run instructions and reproduction workflows were rarely reported
Across the 133 papers, run instructions and end-to-end reproduction workflows were almost never supplied. Code versions, missing-data frequencies, uncertainty estimates, external validity, exclusions, and safeguards against leakage were also often absent or incomplete.
| Item | Reporting requirement | Present | Partial | Absent | Full reporting |
|---|---|---|---|---|---|
R2d | README or equivalent run instructions supplied | 1 | 5 | 127 | 0.8% |
R2e | Reproduction script or workflow supplied | 12 | 9 | 112 | 9.0% |
R2b | Code and code version identified | 5 | 24 | 104 | 3.8% |
R3f | Missing data frequency reported | 15 | 30 | 81 | 11.9% |
R7b | Uncertainty estimates reported with calculation method | 30 | 24 | 79 | 22.6% |
R6b | Dependencies or duplicates across train/test partitions addressed | 15 | 28 | 78 | 12.3% |
R8a | External validity evidence reported | 30 | 40 | 63 | 22.6% |
R4b | Impossible or corrupt sample handling described | 39 | 31 | 60 | 30.0% |
R6a | Preprocessing and modeling restricted to training information where required | 30 | 40 | 49 | 24.8% |
R4a | Sample exclusions identified and justified | 55 | 34 | 42 | 42.0% |
R5f | Baseline comparisons justified | 69 | 28 | 36 | 51.9% |
R5e | Hyperparameter tuning reported | 44 | 55 | 28 | 34.6% |
Shown: the 12 items with the largest absent counts. The downloadable item table contains all 32 items, including unclear and not-applicable counts.
Study goals were reported more consistently than reproducibility details
Module-level scores show the contrast. Authors generally explained their scientific goals and models, but often did not provide the artifacts needed to rerun the analysis or the evidence needed to judge transportability.
| Module | Present | Partial | Absent | Full reporting | Half-credit score |
|---|---|---|---|---|---|
| Computational Reproducibility | 82 | 195 | 387 | 12.3% | 27.0% |
| Data Leakage | 129 | 113 | 131 | 34.3% | 49.3% |
| Metrics Uncertainty | 143 | 116 | 86 | 41.4% | 58.3% |
| Generalizability Limitations | 130 | 66 | 70 | 48.9% | 61.3% |
| Data Preprocessing | 209 | 81 | 104 | 53.0% | 63.3% |
| Data Quality | 561 | 258 | 102 | 60.9% | 74.9% |
| Modeling | 516 | 183 | 86 | 65.7% | 77.4% |
| Study Goals | 383 | 16 | 0 | 96.0% | 98.0% |
Explicit REFORMS adoption
None of the 133 eligible papers cited REFORMS, and none supplied a REFORMS checklist. That is unsurprising for a small early sample: publication lags, indirect adoption through journal policy, and broader reproducibility norms can all affect reporting without producing a citation.
Human validation has not started
Codex performed the screening and item-level assessment from the retrieved paper texts. Human validation checks eligibility decisions and individual item scores so that the project can estimate AI error and place uncertainty bounds around later results. The current 456 screening decisions come from a frozen 1,200-candidate sample.