Using AI in systematic reviews: the moreyou use it, the more your records matter
The workload of a systematic review is, frankly, enormous. Thousands of records to screen, dozens of papers to read closely, extraction tables, bias assessment—wanting help from AI is entirely reasonable. And there are places where leaning on AI is justified. But this review format has one distinctive property: the more you use AI, the more important it becomes to record why each decision was made. This article draws the boundaries and explains how to keep those records.
Published: 2026-07-22
Why records are everything
The value of a systematic review lies in its transparency—anyone should be able to retrace the same procedure. Peer reviewers will always ask how the papers were selected and appraised.
Using AI in the process is nothing to hide. But "the AI made the selection" is not, by itself, an adequate account. Only "a human reviewed the AI's suggestions and made the final call" stands up as a valid record.
Tasks you can hand to AI
Suggesting search terms. Assisting first-pass reading and prioritization of titles and abstracts. Structuring per-paper summaries. Extracting PICO candidates. Drafting extraction tables. Tidying prose.
The common thread: all of these are intermediate products a human can review and correct afterward. Saving time here doesn't lower the review's quality.
Decisions that must not be handed over
The final inclusion or exclusion call. Confirming each exclusion reason. Risk of bias assessment. Judging the certainty of the evidence and the strength of the conclusions.
These decisions shape the review's conclusion directly. There is a line between consulting AI output and letting AI decide—and it must not be crossed. The person who has read the original PDF makes the call and records it, reasons included.
Record format: keep AI suggestions and human judgments separate
A practical tip: split the record. Keep "the AI's candidate assessment and its rationale" and "the human's final decision and its reasoning" as separate entries.
Separating them lets you audit yourself later—are you being steered by the AI?—and lets you describe your methods honestly when you write them up.
Team rules come first
Rigorous reviews run on team procedures: independent screening by two reviewers, consensus checks on disagreements. The role of AI should likewise be a team decision, not an individual one.
Check with your research team, advisor, and target journal on whether to report AI use in your methods, and how much use is acceptable. Norms in this area are still taking shape.
Paperfy's role: a place your decision notes can come home to
Paperfy is not a dedicated screening-management tool. Where it helps is the stage after full text: reading, verifying, and recording.
Each PDF shows its AI summary and figures side by side, and you can edit the summary by hand and leave judgment notes like "Excluded — reason: inconsistent outcome definition (checked Methods, p. 4)". Months later, when you revisit a decision, jumping from your note straight back to the original text is worth a great deal.
What to do with the time you save
Put the hours AI saves you into rereading the originals and into team discussion.
A review's credibility ultimately rests on time spent with the original texts and on the quality of the decision record. AI is a tool for creating that time—not a substitute for it.
Key takeaways
- Limit AI's role to first-pass reading, suggestions, and drafting.
- Record AI suggestions and human final decisions separately.
- Confirm the acceptable scope of AI use with your team, advisor, and journal.
- Keep decision notes linked back to the original PDFs.
Turn AI's time savings into better decisions
With Paperfy, you can rework the AI summary into your own notes—while the original PDF that backs each judgment stays one click away.
Systematic review
Explore This Topic
Read the full cluster for this workflow, from first steps to practical use.
Read Next
These guides continue the same workflow.