Retinopathy of prematurity (ROP) remains one of the most important preventable causes of childhood visual impairment. Screening and timely treatment can preserve vision, but ROP diagnosis is not always straightforward. In particular, “plus disease” – the retinal vascular abnormality that often tips a case toward treatment – is known to be subjective in ROP grading. Experts may differ in how much vessel dilation and tortuosity they require before calling normal, pre-plus, or plus disease.
Such diagnostic variability has raised an important question in Norway. Previous national data showed regional differences in the frequency of treated ROP, including one region with a higher treatment rate that could not be explained by known ROP risk factors. So could differences in how ophthalmologists grade plus disease be driving those regional patterns?
Austeng and colleagues set out to test that possibility by inviting every Norwegian ophthalmologist who makes independent ROP treatment decisions to take part in an image-grading study. All 15 eligible clinicians participated, representing the country’s five ROP treatment centers. Thirteen were pediatric ophthalmologists and two were vitreoretinal surgeons; all had at least five years’ experience in ROP screening.
The participants graded 180 posterior pole retinal images through a web-based platform. The images came from the US Imaging and Informatics in Retinopathy of Prematurity (iROP) study and included normal eyes, eyes with pre-plus disease, and eyes with plus disease. Thirty repeated images were included to assess intra-expert variation. For each image, clinicians assigned one of three diagnoses: normal, pre-plus, or plus disease. The study also used an AI-derived Vascular Severity Score, a nine-level scale designed to capture the continuum of vascular abnormality more finely than the conventional three-category International Classification of Retinopathy of Prematurity (ICROP).
The results suggest that while disagreement exists, it is not enough to explain Norway’s regional treatment differences. Crucially, the center corresponding to the Norwegian region with the previously reported highest frequency of treated ROP did not appear to be over-calling plus disease. If anything, it tended to under-call plus disease compared with both the Norwegian reference standard and the US reference standard. That finding weakens the idea that a lower diagnostic threshold at that center explains the higher historical treatment rate.
The study also confirmed where disagreement is most likely to arise. Agreement was high at the ends of the disease spectrum – normal and plus disease – but lower in the gray zones between normal and pre-plus, and between pre-plus and plus. The same pattern appeared when diagnoses were compared with the AI-derived Vascular Severity Score, with the greatest discrepancies around boundary scores.
Compared with the US reference standard, Norwegian clinicians appeared, on average, to call fewer cases as pre-plus or plus disease. The study authors suggest that Norwegian ophthalmologists may be relative under-callers compared with US experts, although agreement with the US standard still ranged from moderate to substantial at center level.
The study is limited in that it focused on plus disease and did not include ROP stage, another component of treatment decision-making. The authors also acknowledge possible diagnostic bias, as participants knew about Norway’s regional treatment variation before testing. Still, the study’s major strength lies in its national completeness: every Norwegian ophthalmologist making ROP treatment decisions took part in the study.
The conclusion states that the low inter-observer variability found in Norway means it is unlikely to explain the country’s existing regional differences in frequency of treated ROP. For the time being, it remains unclear as to what is causing this regional variability.