ECCV 2026 Workshop On Medical Video Understanding

(MedVidU @ ECCV2026)

Advancing video understanding for surgical and clinical procedures.

9 September 2026 (PM) · Malmö, Sweden

About

Welcome to the Workshop on Medical Video Understanding (MedVidU), held in conjunction with ECCV 2026.

Medical video understanding is a rapidly emerging frontier at the intersection of computer vision and healthcare. Medical procedures generate massive volumes of video data, yet the ability to automatically interpret this data — recognizing phases, detecting critical anatomical structures, grounding temporal events, and generating rich granular textual descriptions — remains an open and challenging problem. Recent advances in large multimodal models and video foundation models have created unprecedented opportunities to tackle these tasks, but standardized benchmarks that span diverse tasks and community-wide evaluation protocols are still lacking.

This workshop addresses this gap by bringing together researchers working on medical AI and video understanding. To catalyze progress, the workshop hosts the MedVidU Challenge, centered on the diverse MedVidBench dataset (CVPR 2026), spanning four domains: laparoscopic surgery, open surgery, robotic surgery, and nursing.

Keynote Speakers

Yueming Jin

Yueming Jin

National University of Singapore

Yueming Jin is an Assistant Professor (Presidential Young Professorship) in the Departments of Biomedical Engineering and Electrical & Computer Engineering at the National University of Singapore, where she leads the IMVR Lab (Intelligent Medical Vision and Robotics). Her research focuses on AI for healthcare, with an emphasis on medical image computing and surgical data science, including multimodal medical AI and intelligent agents for robotic surgery. She was named to the Forbes 30 Under 30 Asia list (2024) and received the IJCARS-MICCAI 2021 Best Paper Award and the ICRA 2021 Best Paper Award in Medical Robotics.

Lalithkumar Seenivasan

Lalithkumar Seenivasan

Johns Hopkins University

Lalithkumar Seenivasan is an Assistant Research Professor in the Department of Computer Science at Johns Hopkins University. Prior to this, he was a Postdoctoral Research Fellow in the ARCADE Lab at Johns Hopkins. His research focuses on vision-language models for surgical applications, including surgical visual question answering and scene graph generation, with the goal of building assistive and automation technologies for clinical training, surgical assistance, and clinical operations. He received his Ph.D. in Biomedical Engineering from the National University of Singapore and won the MICCAI 2023 STAR Award for SurgicalGPT.

Workshop Program

Time Session Presenter(s)
13:30 – 13:45Opening Remarks
13:45 – 14:25Invited TalkYueming Jin
14:25 – 14:40Oral Presentation: OphEdit: Training-Free Text-Guided Editing of Ophthalmic Surgical VideoRitul Jangir, Arkya Jyoti Bagchi, Mangalton Okram, Saurabh Seetaram Korgaonkar, Aiman Farooq, Deepak Mishra
14:40 – 14:55Oral Presentation: Beyond Instrument Motion: Recognizing Tissue Tension Toward Surgical Skill AssessmentMarko Haralovic, Zhiqi Miao, Alexander Machiel Bont, Jiapan Guo, Frans van Workum, Estefania Talavera
14:55 – 15:35Coffee Break / Poster Session
15:35 – 16:15Invited TalkLalithkumar Seenivasan
16:15 – 16:30MedVidU Challenge: Introduction
16:30 – 16:40Oral Presentation: Task-Conditioned Model Selection for Heterogeneous Medical Video UnderstandingMengkang Lu, Yuntian Dong, Yong Xia
16:40 – 16:50Oral Presentation: SPIRAL: Structured Supervision Harvesting and Self-Refining Inference for Heterogeneous Medical Video UnderstandingPodakanti Satyajith Chary, Nagarajan Ganapathy
16:50 – 17:00Challenge Awards & Closing Remarks

Organizers

United Imaging Intelligence · University of Strasbourg / IHU Strasbourg · TUM

Technical Committee

Call for Papers

Please submit your manuscript formatted according to the ECCV 2026 author guidelines. Accepted papers are intended to be published in the ECCV 2026 workshop proceedings.

We encourage submissions reporting novel theories, methods, and applications of video understanding in medical and surgical settings — see the Topics of Interest below.

Note: MedVidU 2026 Challenge participants are required to submit a paper. However, you can submit a paper without taking part in the MedVidU Challenge.

Submit on OpenReview

Topics of Interest

  • Temporal action grounding in medical video
  • Dense video captioning for medical procedures
  • Surgical tool detection and tracking
  • Medical video foundation models and multimodal large language models
  • Surgical phase and action recognition
  • Medical scene understanding from videos and 3D/4D reconstruction
  • Privacy-preserving analysis of surgical recordings
  • Explainability and trustworthiness in AI for medical videos
  • Efficient video representation learning for long medical procedures

Important Dates

Paper Submission July 13, 2026 July 27, 2026 (23:59 AoE) Loading…
Paper Acceptance Notification August 5, 2026 Loading…
Camera Ready August 15, 2026 Loading…

*All deadlines are Anywhere on Earth (AoE). Timelines are subject to change. Submit via OpenReview.

MedVidU Challenge

The MedVidU Challenge is built on the MedVidBench dataset (CVPR 2026), providing a unified evaluation across diverse tasks including Critical View of Safety (CVS) assessment, next action prediction, skill assessment, temporal action grounding, dense video captioning, and video summary & region captioning. The benchmark spans four domains: laparoscopic surgery, open surgery, robotic surgery, and nursing.

Participation in the Challenge is optional. You are welcome to take part in the workshop by submitting a paper on medical video understanding without entering the Challenge — see the Call for Papers below for paper submission details.

Challenge Format

See the Challenge Guide for detailed step-by-step instructions.

Important Dates

Registration Opens & Dataset Release (HuggingFace) May 20, 2026 Live now
Public Validation Leaderboard Opens May 20, 2026 Live now
Workshop Paper Submission (Mandatory for Challenge Participants) July 13, 2026 July 27, 2026 (23:59 AoE) Loading…
Challenge Deadline (Public Leaderboard Closes) July 13, 2026 July 27, 2026 (23:59 AoE) Loading…

*Deadlines are Anywhere on Earth (AoE). Timelines are subject to change.

Top qualifying winners will get a chance to present their method in the workshop.

We plan to have an extended journal submission with the top qualifying teams.

Prizes: 1st place — $800, 2nd place — $200. Prizes are sponsored by United Imaging Intelligence (UII), Boston, MA.

Venue & Location

The workshop will be held in conjunction with ECCV 2026.

Workshop Location

Malmö, Sweden

9 September 2026 (PM)

MedVidU is co-located with the European Conference on Computer Vision (ECCV 2026). Please refer to the main ECCV 2026 website for details on travel and accommodation.

Malmö, Sweden

Contact

For general inquiries about the workshop, please email medvidu@googlegroups.com.