Challenge Design

The challenge has been designed according to the Equator guideline BIAS. The full challenge design can be found here.

Three Competiton Tracks

SAVE FOCUSForeign Object Contextual Understanding for Safe Surgical AI – is organized in three tracks.

Frame Track

The FRAME Track evaluates a model’s ability to answer clinically relevant questions from a single image. This track targets core surgical scene understanding skills such as foreign object detection, identification, attribute recognition, and spatial localization within a single moment in time.

Segment Track

The SEGMENT Track focuses on short video segments (up to 5 min), requiring models to incorporate local temporal context to answer questions about foreign objects and their interactions with anatomy and instruments.

Procedure Track

The PROCEDURE Track challenges models with long surgical video contexts, ranging from extended segments to full-length laparoscopic procedures, to assess their capacity for long-term memory, persistent tracking, and global reasoning.

Together, these tracks enable a systematic characterization of where current VLMs succeed and fail as task complexity transitions from instantaneous perception to long-context intraoperative reasoning.

PRIZES & ALLOCATION

The ORena SAVE FOCUS Challenge is backed by a substantial prize pool to recognize outstanding contributions to safe, clinically grounded surgical AI.

Total Prize Pool
$60,000.00

Track Allocation

FRAME Track ~20%
SEGMENT Track ~40%
PROCEDURE Track (50/50 Technical & Clinical) ~40%

Podium Ranking Split

1st Place 50%
2nd Place 30%
3rd Place 20%

EVALUATION & RANKING (SUMMARY)

Submissions are evaluated following Metrics Reloaded recommendations, balancing interpretability, category fairness, and robustness against out-of-distribution (OOD) cases.

Dual Leaderboards (PROCEDURE)

Technical: Computed on all test VQA cases to assess general multi-modal capability.
Clinical: Computed solely on expert-flagged, high-stakes safety questions.

Scoring & LLM Judge

Accuracy: Primary metric across format-verified exact matches (binary, number, time).
LLM-as-a-Judge: Majority vote across independent judge LLMs for open-ended answers.

Copeland Aggregation

Accuracies are stratified into 10 capability × robustness buckets. The final rank is determined by pairwise bucket dominance (Copeland score) with bootstrap tie-breaking.

Baseline Qualification

Two organizer baselines (Frontier VLM & Fine-tuned Open-source VLM) are provided. Final qualification and prizes require beating both baselines.

CHALLENGE RULES

General Participation

  • Open to individuals, academia, and industry teams (subject to standard sanctions compliance).
  • All submitted inference pipelines must be fully automated without manual intervention.
  • Teams must register as unified entities; multiple accounts to circumvent submission quotas are strictly prohibited.

Allowed Training Resources

  • PROCEDURE Track: Challenge data, public/private datasets, and open/closed foundation models.
  • FRAME & SEGMENT Tracks: Challenge data, public datasets, and public pre-trained models. Any external manually added annotations must be made public by the final deadline.

Pre-Evaluation & Final Qualification

  • Up to 14 submissions per team during pre-evaluation on 20 unseen validation videos.
  • To qualify for the final phase, teams must outperform both baselines on the corresponding leaderboard.

Conflict of Interest & Authorship

  • Teams outperforming the baselines in the final test phase may nominate up to 3 co-authors for the post-challenge publication.
  • Members of organizer labs are listed on the leaderboard for reference but are ineligible for prize money.
EXAMPLE QUESTION

"When was the first sponge inserted in the abdomen? Please return the time-point."

This video includes surgical scenes. Viewer discretion is advised.
This video includes surgical scenes. 
Viewer discretion is advised.