What's new in ROBINS-I V2: version 1 vs version 2
Choosing an appraisal tool
ROBINS-I was published in 2016 and became the recommended risk-of-bias tool for non-randomized studies of interventions in Cochrane Reviews. Version 2 was first released in November 2024, and the current document, posted on 20 November 2025, is still described by the developers as a draft subject to change. It keeps the ideas that defined the original: assessment of one result at a time against a target trial, confounding as the central concern, and a four-level scale anchored to a well-performed randomized trial. But it changes enough of the structure that V1 and V2 assessments are not interchangeable.
This page lists what changed, explains the reasoning behind the larger changes, and sets out what a review team should do if it started under V1.
The short answer
Starting a new review
V2
It is the current version at riskofbias.info and the version CoRATES implements. Note the draft status in your methods and record the document date you used.
Review in progress with V1 assessments
Finish with V1, or re-assess everything
Do not mix versions within a review. Re-assessing under V2 is a substantial job, because the domains, questions and response options differ.
Designs other than cohort studies
Check the released documents
The V2 document currently available covers follow-up (cohort) studies. The Cochrane Handbook notes that variants for several other designs are in preparation.
How the domains map
Five of the seven 2016 domains carry over with the same scope. Classification of intervention and selection of participants swap places, and the deviations domain is absorbed into the per-protocol variant of the confounding domain.
- Confounding becomes Domain 1, Confounding, variant A or B
- Selection of participants becomes Domain 3, Selection of participants
- Classification of interventions becomes Domain 2, Classification of intervention
- Deviations (from intended interventions) in part becomes Domain 1, Confounding, variant A or B
- Missing data becomes Domain 4, Missing data
- Measurement of outcomes becomes Domain 5, Measurement of the outcome
- Selection of the reported result becomes Domain 6, Selection of the reported result
Dashed line: when the analysis estimates a per-protocol effect, the concerns of the deviations domain are assessed as time-varying confounding within Domain 1, variant B.
What changed
| ROBINS-I (2016) | ROBINS-I V2 (2024 onwards) | |
|---|---|---|
| Bias domains | Seven: confounding; selection of participants; classification of interventions; deviations from intended interventions; missing data; measurement of outcomes; selection of the reported result. | Six. The deviations domain is gone, and classification of intervention now comes before selection of participants. |
| Where deviations went | A separate domain covering co-interventions, switches and adherence. | Handled through the effect of interest. If the analysis estimates a per-protocol effect, the confounding domain expands to cover time-varying confounding. The requirement to pre-specify co-interventions was removed. |
| Confounding domain | One set of questions covering baseline and time-varying confounding. | Two variants chosen by the effect of interest: one for the intention-to-treat effect, where only baseline confounding needs to be addressed, and one for the per-protocol effect, where baseline and time-varying confounding are both assessed. |
| Planning | List important confounders and co-interventions at protocol stage; describe the target trial. | List confounders at protocol stage. For each result: specify the numerical result and outcome, answer the triage questions, and describe the target trial including whether the analysis accounted for switches and deviations during follow-up. |
| Triage | None. Every result received the full assessment. | New preliminary questions send a result straight to Critical risk of bias when the authors made no attempt to control confounding and the potential for confounding is sufficient, or when the outcome measurement was inappropriate. |
| Responses | Yes, Probably yes, Probably no, No, No information. | The same five, plus graded options on some questions that distinguish a substantial problem from a minor one, for example No but not substantial versus No and probably substantial. |
| How domain judgements are reached | Guidance described patterns of responses, but the reviewer decided. | Algorithms map the responses to a proposed judgement for each domain. The reviewer can override with a recorded reason. |
| Judgement categories | Low, Moderate, Serious, Critical, No information. | The same, plus a qualified judgement, Low except for concerns about uncontrolled confounding, which recognises that unmeasured confounding can never be ruled out in a non-randomized study. |
| Immortal time and prevalent users | Discussed in the guidance. | Explicit signalling questions in the classification and selection domains. |
| Missing data | Focused on the amount of missing data and the analysis used. | Substantially reconceived and expanded, following the approach taken in RoB 2. |
| Scope | Cohort-type designs, with Handbook guidance extending to controlled before-after and interrupted time series designs. | Follow-up (cohort) studies. Variants for other designs anticipated. |
| Information sources | Recorded in free text alongside the judgements. | A structured list of the sources used: journal articles, protocol, statistical analysis plan, registry records, regulatory documents, individual participant data, correspondence with investigators. |
Why the deviations domain went away
In V1 the deviations domain overlapped with confounding whenever the effect of interest was the effect of starting and adhering to the intervention, and it was frequently applied to intention-to-treat analyses where it did not belong. V2 resolves this by making the effect of interest a switch that changes the tool. When the analysis estimates the effect of assignment, only baseline confounding is assessed. When it estimates the effect of adhering, the confounding domain also covers time-varying confounding, which is where switches, co-interventions and adherence actually bias a non-randomized comparison.
Reviewers who learned V1 will look for the deviations domain and not find it. The concern has not been dropped; it has been placed where the causal structure says it belongs.
Triage changes how you plan
Under V1 a study with no adjustment for confounding received a Critical judgement only after a full pass through every domain. Under V2 the preliminary questions can end the assessment early. This matters most in reviews built on large numbers of registry or database studies, where a substantial fraction may fail triage. The questions ask whether the authors made any attempt to control confounding and whether the potential for confounding is sufficient to set the result aside, so a crude adjustment still leads to the full assessment. Do not use triage as a shortcut for studies that adjusted badly; that is what the confounding domain is for.
Algorithms and graded responses
V1 gave detailed guidance on how patterns of responses should map to judgements but left the final step to the reviewer. V2 follows RoB 2 in proposing the judgement algorithmically, which improves consistency across reviewers and makes overrides visible, because an override has to be recorded with a reason. The graded response options serve the same end. Being able to answer that a confounder was not controlled but the omission is probably not substantial, rather than a bare No, lets the algorithm distinguish a Moderate from a Serious judgement for reasons a reader can follow.
Immortal time and prevalent users
Immortal time arises when the period between entry into a cohort and the start of treatment is counted as exposed time, or is excluded from only one group, so that treated participants appear to survive longer by construction (Suissa, 2008). Prevalent-user designs compare people already established on a treatment with non-users, which conditions on having tolerated it. Both are common in database studies and both were discussed in the V1 guidance. V2 adds explicit signalling questions in the classification and selection domains so that they are asked about every time.
Scope narrowed to follow-up studies
V1 was written for cohort-type designs, and the Cochrane Handbook extended it, with caveats, to controlled before-after and interrupted time series designs. The V2 document currently released is scoped to follow-up (cohort) studies, and the Handbook notes that a new version is under preparation with variants for several types of non-randomized design. Until those variants appear, a review including other designs should state which document it followed and why, and should not present a V2 cohort assessment of a case-control or before-after study as if it were routine.
What this means for a review in progress
Pick one version and apply it to every result. If your protocol said ROBINS-I without a version, state the version in the methods now. Mixing V1 and V2 assessments within one review produces judgements with different domain structures, and a reader cannot tell which studies were assessed which way.
Re-assessing under V2 is not a light touch. The domains differ, the confounding domain depends on a per-result decision about the effect of interest, the response options differ, and the triage step may remove studies from full assessment. Budget for it as new work. If the review is close to completion, finishing under V1 and saying so is the honest choice.
Record the document date. V2 remains a draft and the developers have revised it since the first release. A review that reports the ROBINS-I V2 version dated 20 November 2025 is reproducible; one that reports only ROBINS-I V2 is not.
Appraise with ROBINS-I V2 in CoRATES
Open an appraisal in your browser and work through it item by item. Nothing to install and no account needed.
Frequently asked questions
- Is ROBINS-I V2 mandatory?
- No published Cochrane guidance requires it. V2 is the current version at riskofbias.info and the natural choice for a new review. Follow your protocol and any journal or editorial requirement, and state the version and document date you used.
- Are V1 assessments now wrong?
- No. They are valid for the tool that produced them. Report the version and do not convert them to V2.
- Can I combine V1 and V2 assessments in one review?
- No. The domain structures differ, so the two sets of judgements cannot be presented in one table without misleading readers. Use one version throughout.
- Does V2 still assess one result at a time?
- Yes. Each assessment starts by specifying the numerical result and the outcome it relates to, and the target trial is described for that result.
- Is there a V2 for case-control studies?
- Not in the currently released documents, which cover follow-up (cohort) studies. The Cochrane Handbook notes that variants for several other designs are in preparation.
- Why is there a Low except for concerns about uncontrolled confounding judgement?
- Because a non-randomized study can never demonstrate that all confounding has been controlled. The qualified judgement lets a well-designed, well-adjusted study be recognised as such without claiming the equivalence to a randomized trial that a plain Low would imply.
- How is ROBINS-I V2 related to ROBINS-E?
- ROBINS-E (Higgins et al., 2024) is the companion tool for non-randomized studies of exposures rather than interventions. It was developed alongside the V2 work and shares the same family design, but it is a separate tool with its own domains and should not be substituted for ROBINS-I.
- How does CoRATES implement V2?
- CoRATES follows the V2 structure: the planning list of confounders, the result and outcome specification, the triage questions, the target trial description including the intention-to-treat or per-protocol choice, the confounding variant that follows from it, the six domains with graded responses, and the algorithms that propose domain and overall judgements. Several reviewers can assess independently and reconcile. CoRATES does not modify the official algorithms.
Further reading
- Sterne JA, Hernan MA, Reeves BC, et al. (2016). ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions. BMJ 2016;355:i4919. View
- Higgins JPT, Morgan RL, Rooney AA, et al. (2024). A tool to assess risk of bias in non-randomized follow-up studies of exposure effects (ROBINS-E). Environment International, 186, 108602. View
- Hernan MA, Robins JM. (2016). Using Big Data to Emulate a Target Trial When a Randomized Trial Is Not Available. American Journal of Epidemiology, 183, 758-764. View
- Suissa S. (2008). Immortal time bias in pharmaco-epidemiology. American Journal of Epidemiology, 167, 492-499. View
- Schunemann HJ, Cuello C, Akl EA, et al. (2019). GRADE guidelines: 18. How ROBINS-I and other tools to assess risk of bias in nonrandomized studies should be used to rate the certainty of a body of evidence. Journal of Clinical Epidemiology, 111, 105-114. View
Related guides
This page describes the tools and cites their official sources. It does not reproduce signalling questions, items or scoring tables, which are the intellectual property of their original authors. Consult the official publications and guidance linked above when applying any tool.