Unintentionally retaining foreign objects during minimally invasive surgery can cause serious complications requiring reoperation (Badiee et al., 2025). SAVE FOCUS addresses this patient safety challenge by benchmarking AI on tasks surgeons face intraoperatively: continuous counting and tracking of surgical items, real-time localization of unaccounted objects, and verification of complete retrieval before closure.
At the same time, the challenge targets a critical technical frontier: can vision-language models (VLMs) maintain accurate understanding across hours-long procedures when safety decisions depend on remembering what happened much earlier in the operation?
Lack of access to essential surgery results in 1.5 million deaths per year (Mock et al., 2015). An additional 143 million operations are needed annually, most urgently in low-resource settings where over one-third of the world's population lives without sufficient surgical workforce (Rose et al.).
Minimally invasive surgery holds promise for expanding access by reducing complications and training time. The SAVE (Surgery: Assess / Validate / Expand) program aims to double the number of surgical providers trained each year, creating an additional 100,000 within a decade. By advancing AI systems that provide scalable quality assurance for foreign object tracking and retrieval verification, SAVE FOCUS supports this mission to democratize surgical safety by creating tools that can scale to underserved regions everywhere surgical care is delivered.
Recent progress in general-domain VLMs has enabled increasingly strong temporal reasoning over extended video streams (Bai et al., 2025). However, it is so far unclear whether these emerging capabilities translate to real clinical workflows, where critical events unfold over tens of minutes to hours. This gap is important: many clinically meaningful questions in minimally invasive surgery require persistent memory, temporal consistency, and reasoning across long time horizons. And it is particularly urgent in low-resource settings where surgical teams face higher patient volumes and fewer support personnel, thus making reliable AI assistance even more valuable.
The ORena SAVE FOCUS challenge addresses this unmet need by providing a structured benchmark that targets an urgent patient-safety problem in minimally invasive procedures: ensuring the retrieval of foreign objects, such as sponges and needles, from the abdomen at the end of the operation.
FOCUS aims to generate scientific progress by tackling two fundamental research questions:
Clinical utility: Can VLMs generate clinically meaningful and safety-relevant information about surgical foreign objects that supports quality assurance in diverse healthcare settings?
Technical limits: What are the current limitations of VLMs in surgical scene reasoning, and how can we advance long-context understanding to meet the demands of real surgical workflows?
By establishing a rigorous benchmark on this critical safety problem, SAVE FOCUS aims to accelerate AI capabilities that can democratize surgical assistance—making advanced quality assurance accessible not only in high-income countries, but in the regions of Asia, Africa, and Latin America where surgical workforce expansion is most urgently needed.