SAVE FOCUS – Foreign Object Contextual Understanding for Safe Surgical AI – is organized in three tracks.
The FRAME Track evaluates a model’s ability to answer clinically relevant questions from a single image. This track targets core surgical scene understanding skills such as foreign object detection, identification, attribute recognition, and spatial localization within a single moment in time.
The SEGMENT Track focuses on short video segments (up to 5 min), requiring models to incorporate local temporal context to answer questions about foreign objects and their interactions with anatomy and instruments.
The PROCEDURE Track challenges models with long surgical video contexts, ranging from extended segments to full-length laparoscopic procedures, to assess their capacity for long-term memory, persistent tracking, and global reasoning.
Together, these tracks enable a systematic characterization of where current VLMs succeed and fail as task complexity transitions from instantaneous perception to long-context intraoperative reasoning.
When was the first sponge inserted in the abdomen?
Please return the time-point.