How is a systematic review conducted?

A step-by-step method for a researcher running their first systematic review: protocol, search, screening, risk of bias, and the synthesis decision.

What is a systematic review?

A systematic review is the searching, selecting, and evaluating of the literature according to a predefined and written method in order to answer a specific question. The search query, the inclusion criteria, and the form of data extraction are fixed before screening begins; this way, selection is made according to a rule announced in advance, not whichever studies happen to catch the researcher's attention.

What makes a review systematic is not the search itself but the order of the decisions: in a free-form review the researcher can change criteria while reading, in a systematic review the criteria are written down and any change goes into the report.

It was argued long ago that a literature review should rest on a structured method rather than on the author's memory (Webster and Watson, 2002). A systematic review is the strictest form of this idea: the method is written in enough detail that another team following the same steps could arrive at the same set of studies.

What makes a review systematic is the following:

  • The question to be answered is written in a single sentence before screening begins.
  • The search query is reported exactly as run, together with the database and the date.
  • Inclusion and exclusion criteria are announced in advance and applied the same way at every elimination.
  • Selection decisions are made independently by at least two people.
  • How many records were eliminated at each stage is recorded with a number.
  • The risk of bias of the included studies is assessed one by one.

What is the difference between a systematic review and a traditional review?

A traditional review describes studies the author selected based on their own field knowledge, and which sources were left out and why is often not written down; in a systematic review these steps are defined and reported in advance. The difference is not the topic but the method: in one, selection rests on the author's judgment; in the other, on an announced rule.

Reviews are distinguished from one another by focus, goal, perspective, coverage, organization, and audience (Cooper, 1988); this classification shows that every review carries a coverage claim. A traditional review is content to be representative in its coverage; a systematic review claims to be complete within the criteria it has announced.

The difference between them is not a ranking of quality; a traditional review can be well suited to introducing a new field, opening a conceptual debate, or drawing up a history. But if a study claims to have gathered evidence for a question, the selection rule needs to be open to the reader.

DistinctionTraditional reviewSystematic review
QuestionTakes shape as the writing progressesFixed before screening begins
SearchMostly not describedQuery, database, and date are reported
SelectionAuthor's judgmentCriteria announced in advance
ReviewerUsually a single personAt least two, independent
Eliminated recordsInvisibleRecorded with a number and a reason
ReproductionNot practically possibleSteps can be repeated

What is PRISMA and what is it for?

PRISMA is a guideline that describes how systematic reviews should be reported (Page et al., 2021). It has two parts: a checklist listing what information the author needs to put in the report, and a flow diagram showing how many studies were eliminated at which stage. It governs reporting, not quality: every item being marked shows the work was described completely, not that the review was done well.

The checklist lists the items of information that must be present in the report: study type, abstract, the full text of the search query, the screening process, the assessment of bias, the presentation of results, and deviations from the protocol.

The flow diagram is a story told in numbers: how many records came from the databases, how many were removed as duplicates or eliminated, and how many full texts and final studies remained are all shown in a single diagram. If the diagram cannot be filled in, the process was not recorded.

A common mistake is using the PRISMA checklist as a quality scale. A review carried out with a weak search strategy can still tick every box, as long as it honestly reports that weak strategy; quality assessment is a separate task and runs through the risk of bias of the included studies.

The reverse also holds: a review that was well conducted but poorly reported cannot be verified by the reader. The condition for reproducibility is that the steps leading to the result be shared, not the result itself (Peng, 2011).

Why is a protocol written in advance?

A protocol sets down the question, the search strategy, the inclusion criteria, and the analysis plan in writing before the study begins; changing criteria after seeing the results makes it easier to keep studies that came out in the desired direction. Writing the protocol in advance and putting it on a dated record reduces this risk and makes teams doing the same work twice visible to each other.

A protocol contains at least the following:

  • The question to be answered and its components
  • Inclusion and exclusion criteria
  • The databases to be searched and a draft of the search string
  • Date range, language, and publication type restrictions
  • How many reviewers will work and how disagreement will be resolved
  • Which data will be extracted and how it will be synthesized
  • Which tool will be used to assess risk of bias

Putting the protocol on a dated record before the study begins is called protocol preregistration; it makes changes made after seeing the results impossible to hide, since the registered text can be compared against the published text. Public registries exist that keep this record for systematic reviews.

Deviating from the protocol is not forbidden, hiding it is: if you narrowed a criterion after encountering an unexpected type of study, you report that with its justification. PRISMA 2020 already asks where the protocol is registered and what, if any, deviations there were (Page et al., 2021).

How is a search strategy built?

A search strategy starts by translating the concepts of the question into search terms: for each concept, synonyms and subject headings are combined with OR, and concepts are linked with AND. Databases, date range, and language restrictions are then chosen, and the query is published exactly as run, together with the date.

Search string

Each concept of the question becomes a separate block; within it, that concept's synonyms, spelling variants, and the database's own subject headings are combined with OR, and the blocks are linked to each other with AND. Because truncation and proximity operators differ from database to database, the same query does not work identically everywhere; a version adapted for each database is written and reported separately.

Publishing the query is not a courtesy, it is the method itself. Without the search string a review cannot be reproduced; even if the number of records found is given, how that number was obtained remains unknown (Peng, 2011).

Choosing databases

A single database is not enough; coverage varies from field to field and no index contains every journal. Depending on the topic, several general indexes and the field's own database are searched together, with the reasoning behind the choice written down so the reader can judge the coverage. The reference lists of the included studies are also hand-searched.

Date range and language restriction

The date range needs a justification: the year a method was first published, or the date a regulation took effect, is a good starting point, picking a round number of years is not. A language restriction is also an exclusion criterion; if imposed, its justification and its likely effect on the result are written down.

Grey literature

Grey literature is theses, reports, conference proceedings, and institutional documents that fall outside commercial publishing. Including it is contested: it is hard to search because it is not indexed, and it may not have gone through peer review. On the other hand, looking only at published studies carries into the report the bias that arises because studies with positive results are more likely to be published (Higgins et al., Cochrane Handbook); whatever the decision, it is written in the protocol.

How are inclusion and exclusion decisions recorded?

Every record is screened independently by at least two reviewers: first the title and abstract, then the full text. Disagreement is resolved by discussion or by consulting a third person. The steps are listed below.

  1. Records are collected from the databases and duplicates are removed. The number of records found and the number removed are recorded separately.
  2. Title and abstract screening is done. Two reviewers work independently and do not see each other's decisions.
  3. The full text of the remaining records is obtained. Records that cannot be retrieved are reported as a separate number.
  4. Full texts are read against the criteria. A single exclusion reason is written for every study eliminated.
  5. Disagreements are resolved by discussion, or, if unresolved, by the decision of a third reviewer.
  6. All the numbers are carried into the flow diagram. The totals between stages must be consistent with each other.

The purpose of two independent reviewers is not division of labor but catching errors; when a single person screens, both lapses in attention and selection that fits expectations go unnoticed. Reporting the agreement between reviewers also shows how clearly the criteria were written (Higgins et al., Cochrane Handbook).

A log that keeps decisions together with their justification makes a review auditable. Sharing the extracted data and the decision record in a machine-readable form also makes it possible for other teams who want to work with the same data (Wilkinson et al., 2016).

scimind · doğrulama konsoluörnek veri
İddia
Aynı girdi, iki koşu0/6
A
0.84·····
B
0.84·····

Karşılaştırma

bekliyor

Kanıt

koşu bitmedi

denetim
The same input runs twice, the two outputs are compared cell by cell, and the match is stamped.

How is risk of bias assessed?

Risk of bias asks whether the result of a study systematically deviates from the true effect because of how it was designed and conducted. The assessment is done per study and per domain (Higgins et al., Cochrane Handbook); the result is not a single score but a reasoned judgment.

Risk of bias and reporting quality are different things: a study can be cleanly written and still have an inflated effect estimate because it was not blinded. Conversely, a study that under-reports important details has an uncertain risk, and uncertainty does not mean low risk.

The Cochrane tool for randomized studies queries risk domain by domain: the randomization process, deviations from the intended intervention, missing outcome data, how the outcome was measured, and selective reporting; for each domain both the judgment and its basis are written down. In non-randomized designs the main threat is confounding variables, so separate tools are used, and which one is stated in the protocol.

Here too, the assessment is done independently by two people and the results are presented study by study. A sensitivity analysis that excludes high-risk studies shows how much the overall result depends on them.

Is a meta-analysis always carried out?

No: a meta-analysis is meaningful when the studies' questions, participants, interventions, and measurement forms are similar enough to combine. If heterogeneity crosses this threshold, a single pooled effect size is misleading. In that case the findings are described through a qualitative synthesis and the review still remains systematic.

Heterogeneity comes from three sources: differences in participants, interventions, and outcome measures; differences in designs and risk of bias; and effect estimates being scattered more than can be explained by chance (Higgins et al., Cochrane Handbook).

The decision to pool comes before the statistics: if the studies are not asking the same question, their averages have no meaning, and a low numerical dispersion measure does not fix that. The decision is made first with the content, then tested with the number.

If a meta-analysis cannot be done, a qualitative synthesis is written: the studies are presented in a structured table, the direction and magnitude of the findings are compared, and the likely reasons for conflicting results are discussed. A qualitative synthesis is not a weaker version of a systematic review; the discipline of protocol, search, and screening applies here exactly the same. When a meta-analysis is done, the result is reported together with a confidence interval and sensitivity analyses, not reduced to a single number.

Which tool carries out these steps?

The steps of a systematic review can also be carried out by hand; a written protocol, a reference manager, and a properly kept spreadsheet are enough. What tools that bring the process together in one place contribute is recording inclusion decisions with their justification, automatically carrying the numbers into the flow diagram, and keeping the search query as part of the method itself.

Lacuna is a platform built for this work: it reads a field's literature and marks the gap left between clusters, writes inclusion and exclusion decisions to a decision log with their justification, and freezes the method when the analysis is closed, so the same corpus gives the same result.

Related guides

References

  • Cooper, H. M. (1988). Organizing knowledge syntheses: a taxonomy of literature reviews. Knowledge in Society, 1(1), 104-126.
  • Higgins, J. P. T. et al. (Ed.). Cochrane Handbook for Systematic Reviews of Interventions. Cochrane.
  • Page, M. J. et al. (2021). The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ, 372, n71.
  • Peng, R. D. (2011). Reproducible research in computational science. Science, 334(6060), 1226-1227.
  • Webster, J. and Watson, R. T. (2002). Analyzing the past to prepare for the future: writing a literature review. MIS Quarterly, 26(2), xiii-xxiii.
  • Wilkinson, M. D. et al. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3, 160018.

Do you need to review a field from start to finish?

For a workflow that reads the literature, records decisions with their justification, and seals the method, see the Lacuna page or write to us.