456 candidates reviewed from the core sample

Preliminary results

The current AI review suggests that basic study goals are usually reported, while computational reproducibility, uncertainty, and external validation are much less consistent. Human validation has not started, so these figures describe the reviewed subset rather than the literature.

456lexical candidates screened
133eligible full ML-based-science papers
51.9%applicable items fully reported
0human validation checks completed

51.9% of applicable items were fully reported

Each of the 133 eligible papers was assessed against 32 reporting items derived from Princeton's REFORMS. Of 4,150 applicable item judgments, 2,153 were fully reported, 1,028 were partly reported, and 969 were absent or unclear. The secondary score gives half credit to partial reporting.

The 51.9% figure is an item-level reporting rate; papers were not assigned pass/fail results. Every eligible paper fully reported at least 2 items.

133 candidates met the full scoring criteria

The candidate frame includes papers whose titles or abstracts use explicit ML language. Screening then asks whether model performance or output is used as evidence for a scientific claim, and whether a complete original paper is available.

Eligibility outcome

OutcomePapersShare
Eligible full ML-based-science papers13329.2%
Not ML-based science31769.5%
ML-based science, but no complete paper body40.9%
Unresolved20.4%

Scientific-use classification

ClassPapersShare
ML-based science13730.0%
ML mentioned, but not used as evidence10322.6%
ML methods research8618.9%
Predictive analytics8318.2%
Non-ML quantitative analysis459.9%
Unresolved20.4%

The 137 ML-based-science candidates comprise 133 complete original papers and 4 records without a complete paper body, which could not be scored against the full reporting standard.

Publication form in the candidate set

Publication form is recorded separately from scientific use. This keeps reviews, editorials, responses, and conference abstracts visible as sample outcomes without treating them as complete empirical papers.

Publication formPapersShare
Original Research34575.7%
Review4710.3%
Editorial Opinion173.7%
Systematic Review153.3%
Other102.2%
Conference Abstract92.0%
Commentary Response51.1%
Bibliometric Review40.9%
Meta Analysis20.4%
Protocol10.2%
Uncertain10.2%

Pre/post comparison

The currently eligible papers include 63 before REFORMS and 70 after it. Full reporting was 52.0% before and 51.8% after, a change of -0.2 percentage points. The half-credit score changed by -0.6 points. These figures remain descriptive: the comparison contains 133 papers, and human validation has not started.

PeriodPapersApplicable itemsPresentPartialAbsent or unclearFull reportingHalf-credit score
Before REFORMS631,9651,02249345052.0%64.6%
After REFORMS702,1851,13153551951.8%64.0%

Paper-level reporting varied substantially

Every eligible paper reported at least some REFORMS-aligned information. The table treats each paper equally; the preceding comparison treats each applicable item equally.

PeriodPapersMean paper scoreMedianRangePapers with zero present items
Before REFORMS6352.0%53.1%19.4% to 93.8%0
After REFORMS7051.7%53.2%6.2% to 90.3%0

Run instructions and reproduction workflows were rarely reported

Across the 133 papers, run instructions and end-to-end reproduction workflows were almost never supplied. Code versions, missing-data frequencies, uncertainty estimates, external validity, exclusions, and safeguards against leakage were also often absent or incomplete.

ItemReporting requirementPresentPartialAbsentFull reporting
R2dREADME or equivalent run instructions supplied151270.8%
R2eReproduction script or workflow supplied1291129.0%
R2bCode and code version identified5241043.8%
R3fMissing data frequency reported15308111.9%
R7bUncertainty estimates reported with calculation method30247922.6%
R6bDependencies or duplicates across train/test partitions addressed15287812.3%
R8aExternal validity evidence reported30406322.6%
R4bImpossible or corrupt sample handling described39316030.0%
R6aPreprocessing and modeling restricted to training information where required30404924.8%
R4aSample exclusions identified and justified55344242.0%
R5fBaseline comparisons justified69283651.9%
R5eHyperparameter tuning reported44552834.6%

Shown: the 12 items with the largest absent counts. The downloadable item table contains all 32 items, including unclear and not-applicable counts.

Study goals were reported more consistently than reproducibility details

Module-level scores show the contrast. Authors generally explained their scientific goals and models, but often did not provide the artifacts needed to rerun the analysis or the evidence needed to judge transportability.

ModulePresentPartialAbsentFull reportingHalf-credit score
Computational Reproducibility8219538712.3%27.0%
Data Leakage12911313134.3%49.3%
Metrics Uncertainty1431168641.4%58.3%
Generalizability Limitations130667048.9%61.3%
Data Preprocessing2098110453.0%63.3%
Data Quality56125810260.9%74.9%
Modeling5161838665.7%77.4%
Study Goals38316096.0%98.0%

Explicit REFORMS adoption

None of the 133 eligible papers cited REFORMS, and none supplied a REFORMS checklist. That is unsurprising for a small early sample: publication lags, indirect adoption through journal policy, and broader reproducibility norms can all affect reporting without producing a citation.

Human validation has not started

Codex performed the screening and item-level assessment from the retrieved paper texts. Human validation checks eligibility decisions and individual item scores so that the project can estimate AI error and place uncertainty bounds around later results. The current 456 screening decisions come from a frozen 1,200-candidate sample.

JSON summaryScreening CSVPeriod CSVModule CSVAll items CSV