About
Welcome to the Workshop on Medical Video Understanding (MedVidU), held in conjunction with ECCV 2026.
Medical video understanding is a rapidly emerging frontier at the intersection of computer vision and healthcare. Medical procedures generate massive volumes of video data, yet the ability to automatically interpret this data — recognizing phases, detecting critical anatomical structures, grounding temporal events, and generating rich granular textual descriptions — remains an open and challenging problem. Recent advances in large multimodal models and video foundation models have created unprecedented opportunities to tackle these tasks, but standardized benchmarks that span diverse tasks and community-wide evaluation protocols are still lacking.
This workshop addresses this gap by bringing together researchers working on medical AI and video understanding. To catalyze progress, the workshop hosts the MedVidU Challenge, centered on the diverse MedVidBench dataset (CVPR 2026), spanning four domains: laparoscopic surgery, open surgery, robotic surgery, and nursing.
Keynote Speakers
Yueming Jin
National University of Singapore
Yueming Jin is an Assistant Professor (Presidential Young Professorship) in the Departments of Biomedical Engineering and Electrical & Computer Engineering at the National University of Singapore, where she leads the IMVR Lab (Intelligent Medical Vision and Robotics). Her research focuses on AI for healthcare, with an emphasis on medical image computing and surgical data science, including multimodal medical AI and intelligent agents for robotic surgery. She was named to the Forbes 30 Under 30 Asia list (2024) and received the IJCARS-MICCAI 2021 Best Paper Award and the ICRA 2021 Best Paper Award in Medical Robotics.
Lalithkumar Seenivasan
Lalithkumar Seenivasan is an Assistant Research Professor in the Department of Computer Science at Johns Hopkins University. Prior to this, he was a Postdoctoral Research Fellow in the ARCADE Lab at Johns Hopkins. His research focuses on vision-language models for surgical applications, including surgical visual question answering and scene graph generation, with the goal of building assistive and automation technologies for clinical training, surgical assistance, and clinical operations. He received his Ph.D. in Biomedical Engineering from the National University of Singapore and won the MICCAI 2023 STAR Award for SurgicalGPT.
Workshop Program
| Time | Session | Presenter(s) |
|---|---|---|
| 13:30 – 13:45 | Opening Remarks | |
| 13:45 – 14:25 | Invited Talk | Yueming Jin |
| 14:25 – 14:40 | Oral Presentation: OphEdit: Training-Free Text-Guided Editing of Ophthalmic Surgical Video | Ritul Jangir, Arkya Jyoti Bagchi, Mangalton Okram, Saurabh Seetaram Korgaonkar, Aiman Farooq, Deepak Mishra |
| 14:40 – 14:55 | Oral Presentation: Beyond Instrument Motion: Recognizing Tissue Tension Toward Surgical Skill Assessment | Marko Haralovic, Zhiqi Miao, Alexander Machiel Bont, Jiapan Guo, Frans van Workum, Estefania Talavera |
| 14:55 – 15:35 | Coffee Break / Poster Session | |
| 15:35 – 16:15 | Invited Talk | Lalithkumar Seenivasan |
| 16:15 – 16:30 | MedVidU Challenge: Introduction | |
| 16:30 – 16:40 | Oral Presentation: Task-Conditioned Model Selection for Heterogeneous Medical Video Understanding | Mengkang Lu, Yuntian Dong, Yong Xia |
| 16:40 – 16:50 | Oral Presentation: SPIRAL: Structured Supervision Harvesting and Self-Refining Inference for Heterogeneous Medical Video Understanding | Podakanti Satyajith Chary, Nagarajan Ganapathy |
| 16:50 – 17:00 | Challenge Awards & Closing Remarks |
Organizers
United Imaging Intelligence · University of Strasbourg / IHU Strasbourg · TUM
Technical Committee
Call for Papers
Please submit your manuscript formatted according to the ECCV 2026 author guidelines. Accepted papers are intended to be published in the ECCV 2026 workshop proceedings.
We encourage submissions reporting novel theories, methods, and applications of video understanding in medical and surgical settings — see the Topics of Interest below.
Note: MedVidU 2026 Challenge participants are required to submit a paper. However, you can submit a paper without taking part in the MedVidU Challenge.
Topics of Interest
- Temporal action grounding in medical video
- Dense video captioning for medical procedures
- Surgical tool detection and tracking
- Medical video foundation models and multimodal large language models
- Surgical phase and action recognition
- Medical scene understanding from videos and 3D/4D reconstruction
- Privacy-preserving analysis of surgical recordings
- Explainability and trustworthiness in AI for medical videos
- Efficient video representation learning for long medical procedures
Important Dates
| Paper Submission | Loading… | |
| Paper Acceptance Notification | August 5, 2026 | Loading… |
| Camera Ready | August 15, 2026 | Loading… |
*All deadlines are Anywhere on Earth (AoE). Timelines are subject to change. Submit via OpenReview.
MedVidU Challenge
The MedVidU Challenge is built on the MedVidBench dataset (CVPR 2026), providing a unified evaluation across diverse tasks including Critical View of Safety (CVS) assessment, next action prediction, skill assessment, temporal action grounding, dense video captioning, and video summary & region captioning. The benchmark spans four domains: laparoscopic surgery, open surgery, robotic surgery, and nursing.
Participation in the Challenge is optional. You are welcome to take part in the workshop by submitting a paper on medical video understanding without entering the Challenge — see the Call for Papers below for paper submission details.
Challenge Format
See the Challenge Guide for detailed step-by-step instructions.
- Download the MedVidU ECCV 2026 train/val split from UII-AI/MedVidU_ECCV2026_TrainVal and train your model.
- Run inference on the MedVidBench test set and submit your predictions to the MedVidBench Leaderboard.
- Challenge participants are required to submit a report on OpenReview for the MedVidU Challenge.
Important Dates
| Registration Opens & Dataset Release (HuggingFace) | May 20, 2026 | Live now |
| Public Validation Leaderboard Opens | May 20, 2026 | Live now |
| Workshop Paper Submission (Mandatory for Challenge Participants) | Loading… | |
| Challenge Deadline (Public Leaderboard Closes) | Loading… |
*Deadlines are Anywhere on Earth (AoE). Timelines are subject to change.
Top qualifying winners will get a chance to present their method in the workshop.
We plan to have an extended journal submission with the top qualifying teams.
Prizes: 1st place — $800, 2nd place — $200. Prizes are sponsored by United Imaging Intelligence (UII), Boston, MA.
Venue & Location
The workshop will be held in conjunction with ECCV 2026.
Workshop Location
Malmö, Sweden
9 September 2026 (PM)
MedVidU is co-located with the European Conference on Computer Vision (ECCV 2026). Please refer to the main ECCV 2026 website for details on travel and accommodation.
Contact
For general inquiries about the workshop, please email medvidu@googlegroups.com.