Brief
Review automation's evaluation gap: paper counts unreported evaluations, proposes PRISMA-LLM reporting split
An analysis of 888 review-automation papers finds wide variation in whether automated review steps were evaluated at all, and proposes a framework separating implementation disclosure from consequence-sensitive evaluation. The evidence covers reporting patterns only: the source gives no implementation detail, tooling, or adoption for the proposed framework.

The paper analyzes SciLitBench, a corpus of 888 review-automation papers carrying 14,726 annotations, to characterize changes in methods, review-stage use, evaluation and reported limitations. Automation, it reports, has shifted toward LLM- and software-facing workflows, including stages that can alter the evidence base.
By the paper's count, 38.0% of software and product papers reported no evaluation since 2023, against 9.3% of LLM papers. Reporting coverage increased with LLM workflow complexity, yet 52% of positive-only LLM evaluations still reported an unmet reliability or performance requirement.
From those patterns the authors introduce PRISMA-LLM, which separates implementation disclosure from consequence-sensitive evaluation and limitation reporting.
Our reading
For anyone running an AI-assisted step inside a review or evidence pipeline, the useful signal is that evaluation reporting tracks workflow complexity rather than stage risk: the source says automation has moved into stages that can alter the evidence base, while a large share of software-facing work carries no evaluation report at all. The proposed split matters operationally because it treats d…
What to do or watch
If your workflow includes an LLM-facing or software-facing review stage, check whether that stage currently has any evaluation report and whether unmet reliability or performance requirements are recorded alongside positive results. The unresolved question is whether PRISMA-LLM has been implemented or tested anywhere: the evidence describes the framework's intent but reports no deployment.
Source details and supporting facts
Each line is stated by the page named above it.
Stated by arXiv
- The analysis uses SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations.
- Since 2023, 38.0% of software/product papers reported no evaluation, compared with 9.3% of LLM papers.
- Reporting coverage increased with LLM workflow complexity, yet 52% of positive-only LLM evaluations still reported an unmet reliability or performance requirement.
- PRISMA-LLM separates implementation disclosure from consequence-sensitive evaluation and limitation reporting.
Sources
- arXivText stored 16 September 2026
How this story was checked. Written from the 1 page listed above, stored 16 September 2026; claims checked against that stored text on 16 September 2026.
What that means
- 4 of 4 reported statements were confirmed against the page that carries them; the rest were removed rather than published.
- Figures in the text were required to appear in the stored source text: yes. Identifiers: yes.
- The check reads stored text only: no claim rests on a fresh look that did not happen.
- Where the reporting was silent, the text says so instead of filling the gap.