We will release 200 annotated full length laparoscopic videos for training.
We will release >50,000 question-answer pairs along with the training videos. Questions will cover the following capabilities:
1. Object recognition and identity matching (Which object (type)? In which state? Where?):
Recognition of object instances, their semantic type, attributes, spatial context, and identity consistency across time.
2. Temporal grounding (When? How long?):
Localization of events or object occurrences within the video timeline, including their exact temporal position and duration.
3. Aggregation (How many?):
Combination of information on foreign objects and events across multiple objects, instances, and/or time points.
4. Event and procedural understanding (Which action?):
Recognition and interpretation of actions, events, and their procedural structure over time.
5. Complex reasoning (Why? What happens if?):
Inference of functional, causal, or outcome-related information beyond direct observation.
