RoB 2 vs ROBINS-I: which risk-of-bias tool for your study?
Choosing an appraisal tool
RoB 2 and ROBINS-I are companion tools from the Cochrane Bias Methods Group, and the line between them is drawn by study design rather than by topic, by quality, or by what the study calls itself. If participants were allocated to interventions by a genuinely random process, use RoB 2. If they were not, use ROBINS-I. Most of the questions people bring to this choice are about designs near that line, or about what to do when a review includes both.
This page gives the short answer first, then a design-by-design guide, a side-by-side comparison, and the differences that matter once you start reporting and synthesising results. It describes the tools; it does not reproduce their signalling questions, which belong to the tool developers and are linked at the end.
The short answer
Randomized trial
RoB 2
Parallel-group trials use the main version. Cluster-randomized and crossover trials use the RoB 2 variants written for those designs, which add design-specific signalling questions.
Non-randomized study of an intervention
ROBINS-I V2
Cohort-type follow-up studies in which intervention groups were not formed by randomization: prospective and retrospective cohorts, registry and database analyses, non-randomized controlled trials, and quasi-randomized designs.
Review that includes both
Both, kept separate
Apply each tool to the studies it was designed for, present the assessments separately, and analyse the two groups separately. The two judgement scales are related but not interchangeable.
Choose by study design
The Cochrane Handbook scopes RoB 2 to randomized trials and ROBINS-I to non-randomized studies of interventions. Work down the questions below, then use the table for the designs reviewers most often ask about, including the ones that fall outside both tools.
Were participants allocated to interventions by a genuinely random process?
- Yes
Which randomized design?
- Parallel-group
RoB 2
The main version of the tool, assessed per result.
- Cluster-randomized
RoB 2, cluster variant
Adds a domain for bias from the timing of identification and recruitment.
- Crossover
RoB 2, crossover variant
Adds questions on carry-over and period effects.
- Parallel-group
- No, including quasi-randomized
Is there a comparison between intervention groups?
- No
Neither tool
Single-arm studies and case series have no effect estimate whose bias can be assessed.
- Yes
Is it a follow-up (cohort-type) design?
- Yes
ROBINS-I V2
Cohorts, registry and database analyses, non-randomized controlled trials.
- No
Outside the current V2 scope
Case-control and before-after designs. Document the approach you take.
- Yes
- No
| Study design | Tool | Why |
|---|---|---|
| Parallel-group randomized trial | RoB 2 | The main version of the tool. Assess each result of interest separately: a result is one outcome, at one time point, from one analysis. |
| Cluster-randomized trial | RoB 2, cluster variant | Adds a domain for bias arising from the timing of identification and recruitment of participants, because people recruited after their cluster was allocated may have been selected with knowledge of the intervention. Do not use the parallel-group version. |
| Crossover trial | RoB 2, crossover variant | Adds signalling questions about carry-over and period effects. Do not use the parallel-group version. |
| Quasi-randomized trial (allocation by alternation, date of birth, record number) | ROBINS-I V2 | A predictable allocation rule is not randomization, so the groups cannot be assumed free of confounding. Cochrane treats these as non-randomized controlled trials, which fall within the scope of ROBINS-I. |
| Non-randomized controlled trial | ROBINS-I V2 | Investigator-assigned groups without randomization. The Cochrane Handbook lists this design among the follow-up studies ROBINS-I covers, however the study is labelled. |
| Prospective or retrospective cohort study | ROBINS-I V2 | The design ROBINS-I V2 is written for. Confounding is usually the decisive domain, and the confounders to look for should be listed at protocol stage. |
| Registry, electronic health record or claims database analysis | ROBINS-I V2 | Treated as a cohort study. Classification of intervention, immortal time and prevalent-user bias are the usual trouble spots, and V2 asks about them explicitly. |
| Controlled before-after study, interrupted time series | ROBINS-I, with care | The Cochrane Handbook discusses these designs under ROBINS-I, but the V2 document currently released covers follow-up (cohort) studies and the developers have signalled variants for other designs. Say in your protocol which version and guidance you followed. |
| Case-control study | Outside current scope | Neither tool is written for case-control designs. ROBINS-I variants for further designs are anticipated; until then, justify and report whatever approach you take rather than applying the cohort tool silently. |
| Single-arm study, case series | Neither | Both tools assess the result of a comparison between intervention groups. Without a comparator there is no effect estimate whose bias can be assessed. |
| Systematic review | AMSTAR 2 or ROBIS | Neither RoB 2 nor ROBINS-I appraises a review. See the AMSTAR 2 vs ROBIS comparison linked below. |
Side by side
| RoB 2 | ROBINS-I V2 | |
|---|---|---|
| Published | August 2019 (Sterne et al., BMJ). Replaces the original Cochrane tool of 2008 and 2011. | October 2016 (Sterne et al., BMJ). Version 2 first released November 2024; the current document was posted on 20 November 2025 and is still marked by the developers as a draft subject to change. |
| Study designs | Randomized trials. Variants for cluster-randomized and crossover trials. | Non-randomized follow-up (cohort) studies of interventions. Variants for further designs anticipated. |
| Unit of assessment | A specific result: one outcome, time point and analysis. | A specific result, judged against a target randomized trial that the study is taken to emulate. |
| Before you start | Decide whether you are assessing the effect of assignment to intervention or the effect of adhering to it. Domain 2 differs between the two. | List the important confounders at protocol stage, specify the target trial, and decide whether the analysis estimates an intention-to-treat or per-protocol effect. Domain 1 differs between the two. |
| Bias domains | Five: the randomization process; deviations from intended interventions; missing outcome data; measurement of the outcome; selection of the reported result. | Six: confounding; classification of intervention; selection of participants into the study or analysis; missing data; measurement of the outcome; selection of the reported result. |
| Triage | None. Every result receives the full assessment. | Preliminary questions can send a result straight to Critical risk of bias, for example when the authors made no attempt to control confounding, so the full assessment is skipped. |
| Signalling question responses | Yes, Probably yes, Probably no, No, No information. | The same five, plus graded options on some questions that separate a substantial problem from a minor one. |
| Domain judgements | Low risk, Some concerns, High risk. | Low, Moderate, Serious, Critical, or No information. In V2 the confounding domain can also return Low except for concerns about uncontrolled confounding. |
| What Low means | The result is comparable to a well-conducted trial for that domain. | Comparable to a well-performed randomized trial for that domain. The Handbook expects this to be rare for confounding in non-randomized studies. |
| Overall judgement | High if any domain is High, or if Some concerns in several domains substantially lowers confidence. Low only if every domain is Low. | Driven by the worst domain. Critical means the result is too biased to be useful and should not be included in synthesis. |
| Algorithms | Published algorithms propose the domain judgement from the responses. Reviewers may override with a documented reason. | Added in V2. The 2016 version gave guidance but left the mapping to the reviewer. |
| Expertise | Methodological. Knowledge of the clinical area helps with the deviations and measurement domains. | Methodological and content expertise together. The Cochrane Handbook recommends involving both methodologists and health professionals who know the prognostic factors. |
| Cochrane status | Recommended tool for randomized trials in Cochrane Reviews. | Recommended tool for non-randomized studies of interventions in Cochrane Reviews. |
| In CoRATES | Supported (parallel-group version). | Supported (V2). |
Confounding is the reason ROBINS-I exists
RoB 2 has no confounding domain because a properly generated and concealed random allocation deals with confounding by design; its first domain therefore asks whether the randomization actually worked. ROBINS-I starts from the opposite premise. Its first domain asks whether the study measured and adjusted for the confounders that matter, which means the review team must know the clinical area well enough to name those confounders before assessment begins. This is the main reason ROBINS-I assessments take longer, and the reason the Cochrane Handbook asks for content expertise on the team.
The scales look similar but do not line up
RoB 2 has three judgement levels and ROBINS-I has four. ROBINS-I anchors Low to a well-performed randomized trial, so a non-randomized study rated Low overall is unusual, and the Handbook says as much. Moderate means the study is sound for a non-randomized study; it is not a synonym for Some concerns. Critical has no counterpart in RoB 2 at all: it means the result is too biased to be informative and should be left out of the synthesis rather than downweighted.
Keep the two sets of judgements in separate tables and figures, and describe each scale in your methods. A single traffic-light plot that colours Moderate and Some concerns the same yellow invites readers to treat them as equivalent.
The target trial
ROBINS-I asks you to describe the hypothetical randomized trial the study is emulating: the eligible participants, the intervention strategy, the comparator strategy, and whether the analysis estimates the effect of starting the intervention or of starting and adhering to it. Every judgement is then made relative to that trial rather than to an idealised observational study. RoB 2 contains a lighter version of the same idea. You state whether you are assessing the effect of assignment or of adherence, and the second domain changes accordingly.
Triage in ROBINS-I V2
V2 adds a short preliminary section before the full assessment. If the authors made no attempt to control for confounding and the potential for confounding is sufficient, or if the method of measuring the outcome was inappropriate, the result is at Critical risk of bias and no further assessment is required. RoB 2 has no equivalent; every result receives the full assessment. In a review built on large database studies the triage step can save a great deal of time, but note that it asks about any attempt to control confounding, so a crude adjustment still leads to the full assessment.
Time, training and agreement
Both tools are more demanding than their predecessors. An evaluation of RoB 2 found only slight inter-rater agreement on the overall judgement among experienced reviewers without calibration, at roughly half an hour per result (Minozzi et al., 2020). A follow-up study found agreement improved substantially once the team wrote review-specific implementation instructions (Minozzi et al., 2022). The Handbook describes ROBINS-I as more involved again.
Whichever tool you use, plan a calibration exercise on a few studies, write down how you will answer the questions that gave trouble, and use two independent reviewers with a documented reconciliation step.
Reviews that include both randomized and non-randomized studies
Apply each tool to the studies it was designed for and say so in the protocol. If a study reports both a randomized comparison and a non-randomized comparison, for example a patient-preference or comprehensive cohort design, assess the randomized result with RoB 2 and the non-randomized result with ROBINS-I as separate results.
Present the two sets of assessments separately, in separate risk-of-bias tables and separate summary figures. The Cochrane Handbook advises that randomized trials and non-randomized studies should not be combined in a single meta-analysis and that their results should be presented and analysed separately, so the risk-of-bias presentation should follow the same split.
The choice of tool also reaches into GRADE. Under the GRADE guidance written for ROBINS-I (Schunemann et al., 2019), a body of non-randomized evidence assessed with ROBINS-I starts at high certainty and is rated down for the bias actually found, instead of starting at low certainty because of its design. The Handbook notes that the final rating is still usually low or very low. That approach only holds up if the ROBINS-I assessment was done rigorously, with confounders pre-specified and a target trial stated.
Appraise with RoB 2 and ROBINS-I V2 in CoRATES
Open an appraisal in your browser and work through it item by item. Nothing to install and no account needed.
Frequently asked questions
- Can I use RoB 2 for a non-randomized study that is designed like a trial?
- No. The first RoB 2 domain assumes a random allocation process existed and asks whether it was carried out properly. Without randomization there is no way within RoB 2 to consider confounding, which is usually the largest source of bias in a non-randomized comparison. Use ROBINS-I.
- Which tool should I use for a quasi-randomized trial?
- ROBINS-I. Allocation by alternation, date of birth, hospital number or day of the week is predictable, so recruiters can foresee the next assignment and the groups cannot be assumed comparable. Cochrane treats such studies as non-randomized controlled trials.
- Which tool should I use for a pilot or feasibility randomized trial?
- RoB 2, as for any randomized trial. A small sample is a question of imprecision, which GRADE handles separately; it is not a source of bias.
- Which tool should I use for a cohort study?
- ROBINS-I V2. Cohort and other follow-up designs are exactly what V2 is written for. Start by listing the confounders you expect to matter and specifying the target trial, then assess one result at a time.
- Can I use ROBINS-I V2 for a case-control study?
- The V2 document currently released covers follow-up (cohort) studies, and the developers have indicated that variants for other designs are in preparation. If you must assess a case-control study now, say in your protocol what you did and why, and treat the resulting judgements with caution.
- Is Moderate in ROBINS-I the same as Some concerns in RoB 2?
- No. Moderate in ROBINS-I means the study is sound for a non-randomized study but not comparable to a well-performed trial; it is the expected result for a good cohort study in the confounding domain. Some concerns in RoB 2 is a judgement about a randomized trial that falls short of Low for a specific reason. Report each on its own scale.
- When do I need ROBINS-E instead?
- When the question is about an exposure rather than an intervention: a pollutant, a diet, an occupational hazard. ROBINS-E (Higgins et al., 2024) is the companion tool for those studies and shares the same family structure. ROBINS-I is for interventions that could, at least in principle, be assigned in a trial.
- Does CoRATES support both tools?
- Yes. CoRATES implements RoB 2 for randomized trials and ROBINS-I V2 for non-randomized studies, applies the official algorithms to propose domain and overall judgements, and supports independent assessment by several reviewers followed by reconciliation. Each tool keeps its own judgement scale.
Reference Documents
- RoB 2 official tool and templates (riskofbias.info)
- ROBINS-I V2 official tool and guidance (riskofbias.info)
- Cochrane Handbook Chapter 8: Assessing risk of bias in a randomized trial
- Cochrane Handbook Chapter 24: Including non-randomized studies on intervention effects
- Cochrane Handbook Chapter 25: Assessing risk of bias in a non-randomized study
Further reading
- Sterne JAC, Savovic J, Page MJ, et al. (2019). RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ 2019;366:l4898. View
- Sterne JA, Hernan MA, Reeves BC, et al. (2016). ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions. BMJ 2016;355:i4919. View
- Higgins JPT, Morgan RL, Rooney AA, et al. (2024). A tool to assess risk of bias in non-randomized follow-up studies of exposure effects (ROBINS-E). Environment International, 186, 108602. View
- Schunemann HJ, Cuello C, Akl EA, et al. (2019). GRADE guidelines: 18. How ROBINS-I and other tools to assess risk of bias in nonrandomized studies should be used to rate the certainty of a body of evidence. Journal of Clinical Epidemiology, 111, 105-114. View
- Minozzi S, Cinquini M, Gianola S, Gonzalez-Lorenzo M, Banzi R. (2020). The revised Cochrane risk of bias tool for randomized trials (RoB 2) showed low interrater reliability and challenges in its application. Journal of Clinical Epidemiology, 126, 37-44. View
- Minozzi S, Dwan K, Borrelli F, Filippini G. (2022). Reliability of the revised Cochrane risk-of-bias tool for randomised trials (RoB2) improved with the use of implementation instruction. Journal of Clinical Epidemiology, 141, 99-105. View
This page describes the tools and cites their official sources. It does not reproduce signalling questions, items or scoring tables, which are the intellectual property of their original authors. Consult the official publications and guidance linked above when applying any tool.