CoRATES is in early access. We're actively building and welcome your feedback.

AMSTAR 2 vs ROBIS: which tool for appraising systematic reviews?

Choosing an appraisal tool

AMSTAR 2 and ROBIS are the two tools most often used to appraise systematic reviews, whether in an umbrella review, in guideline development, or in a health technology assessment that leans on an existing review. They cover much of the same ground and their verdicts usually agree, but they ask different questions. AMSTAR 2 asks how well the review was conducted; ROBIS asks whether the review's conclusions are likely to be biased. The Cochrane Handbook's chapter on overviews does not recommend one over the other, citing a lack of empirical evidence, so the choice needs to be made and justified in your protocol.

This page sets out when each tool fits, the structural differences, what the empirical comparisons found, and the mistakes that show up most often in practice.

The short answer

Overview or umbrella review of intervention reviews

AMSTAR 2, usually

Written for reviews of healthcare interventions, quicker to apply, and its overall confidence rating is designed for exactly this use. The Handbook notes it may be preferred for future Cochrane overviews.

Reviews of diagnostic accuracy, prognosis or aetiology

ROBIS

AMSTAR 2 is scoped to intervention reviews and several of its items assume that context. ROBIS was built to cover interventions, diagnosis, prognosis and aetiology.

A guideline or HTA decision resting on one review

ROBIS, or both

ROBIS was designed with guideline developers in mind, and its final phase judges whether the review's interpretation of its findings is trustworthy, which is the question a panel is actually asking.

Side by side

Published2017 (Shea et al., BMJ). A substantive revision of the 2007 AMSTAR.2016 (Whiting et al., Journal of Clinical Epidemiology).
What it assessesMethodological quality: how well the review was conducted and, on some items, reported.Risk of bias: whether the review process and the interpretation of its findings could have distorted the conclusions.
Review typesSystematic reviews of healthcare interventions that include randomized trials, non-randomized studies, or both.Systematic reviews of interventions, diagnosis, prognosis and aetiology.
Structure16 items, seven of which are designated critical.Three phases: assess relevance (optional); identify concerns with the review process across four domains; judge the overall risk of bias in the review.
DomainsNot domain-based. The critical items cover protocol registration, search adequacy, justification of exclusions, risk-of-bias assessment of included studies, meta-analytical methods, consideration of risk of bias in interpretation, and publication bias.Study eligibility criteria; identification and selection of studies; data collection and study appraisal; synthesis and findings.
ResponsesPer item: Yes, No, and Partial Yes on some items.Per signalling question: Yes, Probably yes, Probably no, No, No information. Per domain: Low, High or Unclear concern.
OutputOverall confidence in the results of the review: High, Moderate, Low or Critically Low.Overall risk of bias in the review: Low, High or Unclear.
How the output is reachedDecision rules based on the pattern of critical and non-critical weaknesses. One critical flaw gives Low; more than one gives Critically Low.Reviewer judgement across the domain concerns and the phase 3 questions about whether the interpretation addressed them. No decision rules.
ScoreExplicitly none. The developers warn against reporting an item count or percentage.None.
Intended usersBroad: clinicians, guideline developers, overview authors, methodologists, journal editors.Primarily guideline developers and overview authors, plus review authors who want to avoid bias in their own reviews.
EffortLower. Items are concrete and most can be answered from the review report. Perry et al. (2021) found it more straightforward to use.Higher. Each domain judgement is a reasoned decision, and reviews without a formal synthesis are harder to rate.
In CoRATESSupported.Not currently supported.

Quality and bias are different questions

A review can follow good practice on every item and still reach a biased conclusion, if it interprets its findings without regard to the limitations of the included studies. Equally, a review with several process weaknesses can still land on the right answer. AMSTAR 2 counts weaknesses in conduct and reporting and turns them into a confidence rating. ROBIS asks, for each weakness, whether it could actually have distorted the result, and then in phase 3 whether the review authors took it into account when they interpreted their findings.

AMSTAR 2 is not blind to interpretation. Its thirteenth item asks whether risk of bias in the primary studies was accounted for when interpreting the results, and it is one of the seven critical items. But in ROBIS the appropriateness of the conclusions is a full phase of the assessment, and the relevance question from phase 1 feeds into it. If your purpose is to decide whether to act on a review, that phase is what you are paying for.

Scope: what each tool was written for

AMSTAR 2 is written for reviews of healthcare interventions. Its items about randomized and non-randomized designs, meta-analytical methods and the funding of included studies assume that context. Applied to a review of diagnostic accuracy or prognostic factors, several items do not fit, and adapting them silently produces a rating that readers will misread as a standard AMSTAR 2 result.

ROBIS was built to apply across question types. That flexibility is why guideline programmes with mixed evidence often reach for it, and why some overview authors find it vaguer than AMSTAR 2 for a pure intervention question.

The critical-item mechanism in AMSTAR 2

AMSTAR 2's rating rules are severe by design. Any one critical flaw caps the review at Low confidence, and two cap it at Critically Low, regardless of how many other items were satisfied. In published overviews the large majority of reviews land in those two categories, most often because of a missing protocol or an incomplete list of excluded studies.

The developers allow assessors to adjust which items are treated as critical for a particular field, but only if the choice is made in advance and reported. Changing the critical set after seeing the ratings is the AMSTAR 2 equivalent of outcome switching.

What the empirical comparisons found

Head-to-head studies find the two tools closely related rather than identical. Lorenz et al. (2019) reported high concordance between overall AMSTAR 2 and ROBIS ratings, with inter-rater reliability moderate for AMSTAR 2 and fair for ROBIS. Perry et al. (2021) applied both to the same 31 reviews and found identical median agreement between raters, with AMSTAR 2 more straightforward to use and ROBIS easier to apply to reviews that included a meta-analysis. Buhn et al. (2017) found ROBIS had fair reliability and good construct validity, and that reliability depended on the raters' experience.

The practical lesson is the same for either tool: pilot on a handful of reviews, write down how you will handle the items that caused disagreement, and use two independent assessors with a reconciliation step. Neither tool is reliable in the hands of a single untrained reviewer.

Using both, and what neither does

Applying both tools to the same reviews is common in methods research and occasionally in overviews. If you do, present both sets of results and decide in advance how a discordant verdict will be reported; do not merge them into a single rating. If you only have capacity for one, choose it on scope and purpose, state the choice in the protocol, and apply it to every included review.

Neither tool is a reporting checklist: PRISMA tells authors what to report and should not be used to appraise a review. And neither rates the certainty of the evidence: that is GRADE, applied to the body of evidence within a review, not to the review as a document.

Appraise with AMSTAR 2 in CoRATES

Open an appraisal in your browser and work through it item by item. Nothing to install and no account needed. ROBIS is not currently available in CoRATES.

Frequently asked questions

Which tool is more widely used?
AMSTAR 2, by a wide margin, in published overviews and umbrella reviews. ROBIS is used more by guideline developers and in methods research. The Cochrane Handbook chapter on overviews describes both, says it cannot currently recommend one over the other, and notes that AMSTAR 2 may be preferred for future Cochrane overviews.
Can I report an AMSTAR 2 result as a score out of 16?
No. The developers state that AMSTAR 2 is not intended to produce an overall score, because the items are not of equal importance. Report the overall confidence rating and, ideally, the item-level responses so readers can see which weaknesses drove it.
Can I use AMSTAR 2 on a review that includes both randomized and non-randomized studies?
Yes. That is the main reason AMSTAR 2 was developed. Several items ask separately about randomized and non-randomized designs, and the risk-of-bias item expects an appropriate tool to have been used for each.
Is ROBIS suitable for reviews of diagnostic test accuracy?
Yes. ROBIS was designed for reviews of interventions, diagnosis, prognosis and aetiology. AMSTAR 2 was not.
Can I appraise the primary studies with AMSTAR 2 or ROBIS?
No. Both appraise the review. The studies inside it need a tool for their own design: RoB 2 for randomized trials, ROBINS-I for non-randomized studies of interventions.
Should I use PRISMA to assess the quality of a review?
No. PRISMA is a reporting guideline. A review can be fully PRISMA-compliant and still be at high risk of bias, and a poorly reported review may have been well conducted. Use AMSTAR 2 or ROBIS for appraisal.
Do the two tools give the same verdict?
Usually, but not always. Overall ratings correlate strongly, and disagreements tend to arise where AMSTAR 2 assigns Critically Low for a critical flaw that ROBIS treats as a concern the review's interpretation addressed. Expect some discordant reviews and decide beforehand how to report them.
Does CoRATES support ROBIS?
Not at present. CoRATES supports AMSTAR 2, including the official decision rules for the overall confidence rating, multi-reviewer appraisal and reconciliation. If ROBIS support matters for your work, tell us through the contact page.

Further reading

  • Shea BJ, Reeves BC, Wells G, et al. (2017). AMSTAR 2: a critical appraisal tool for systematic reviews that include randomised or non-randomised studies of healthcare interventions, or both. BMJ 2017;358:j4008. View
  • Whiting P, Savovic J, Higgins JPT, et al. (2016). ROBIS: A new tool to assess risk of bias in systematic reviews was developed. Journal of Clinical Epidemiology, 69, 225-234. View
  • Perry R, Whitmarsh A, Leach V, Davies P. (2021). A comparison of two assessment tools used in overviews of systematic reviews: ROBIS versus AMSTAR-2. Systematic Reviews, 10, 273. View
  • Lorenz RC, Matthias K, Pieper D, et al. (2019). A psychometric study found AMSTAR 2 to be a valid and moderately reliable appraisal tool. Journal of Clinical Epidemiology, 114, 133-140. View
  • Buhn S, Mathes T, Prengel P, et al. (2017). The risk of bias in systematic reviews tool showed fair reliability and good construct validity. Journal of Clinical Epidemiology, 91, 121-128. View

This page describes the tools and cites their official sources. It does not reproduce signalling questions, items or scoring tables, which are the intellectual property of their original authors. Consult the official publications and guidance linked above when applying any tool.