Surgical Video Phase Detection via Deep Learning Annotation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The large volume of surgical videos recorded during robotic surgeries makes it impractical for highly skilled medical professionals to manually identify and annotate various phases and activities, hindering efficient analysis and real-time feedback for surgeons.
Innovation Solution
Implementing video processing techniques that divide surgical videos into blocks, apply prediction models to predict phases and activities, and generate aggregated predictions to annotate the videos, allowing for real-time or near-real-time identification and alerting of unusual activities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If highly skilled medical professionals manually examine surgical videos to identify phases and mark events, then annotation accuracy is improved, but time consumption and resource requirements increase significantly
Solution Approach 1:
The patent introduces an intermediary system consisting of deep learning models (phase prediction model and activity recognition model) that act as mediators between the surgical videos and the final annotations. These models pre-process and analyze the videos, generating preliminary annotations that medical professionals can then review and verify, significantly reducing their time consumption while maintaining high accuracy through the intermediary's automated analysis
Solution Approach 2:
The system performs preliminary action by automatically generating phase predictions and activity annotations before medical professionals examine the videos. The deep learning models pre-analyze the surgical videos, identify phases, detect activities, and create preliminary annotation sets that reduce the workload of medical professionals while ensuring accuracy through their subsequent review
2Loss of information
If manual annotation of all surgical videos is performed, then comprehensive analysis is improved, but productivity decreases due to the large volume of videos
Solution Approach 1:
The patent replaces the mechanical system of manual annotation by medical professionals with an automated electronic system based on deep learning models. The phase prediction model and activity recognition model automatically process surgical videos, extracting phases and activities without manual intervention, thereby dramatically increasing annotation throughput while maintaining comprehensive analysis through the models' ability to process all videos in the dataset
3Reliability
If real-time phase recognition is implemented during surgery, then surgical safety is improved through immediate alerts, but system complexity increases
Solution Approach 1:
The patent implements feedback by having the phase prediction model continuously analyze video frames in real-time during surgery and provide immediate feedback through alerts and notifications when specific phases or unusual activities are detected. This feedback loop enables real-time phase recognition that improves surgical safety by alerting surgeons to critical moments without requiring complex manual monitoring systems
Solution Approach 2:
The system performs self-service by automatically detecting phases and activities without requiring complex external monitoring or manual intervention. The deep learning models independently process video streams, identify phases, detect activities, and generate alerts autonomously, reducing system complexity compared to manual monitoring approaches while maintaining high reliability through automated real-time analysis
Data Source
AI summary
One example method for detecting phases of a surgical procedure via video processing includes accessing a video of the surgical procedure and dividing the video into one or more blocks, each of the blocks containing one or more video frames. For each of the blocks, the method includes applying a prediction model on the video frames of the respective block to obtain a phase prediction for each of the video frames. The prediction model is configured to predict, for an input video frame, one of the plurality of phases of the surgical procedure. The method further includes generating an aggregated phase prediction for the respective block by aggregating the phase predictions of the video frames, and modifying the video of the surgical procedure to include an indication of a predicted phase of the respective block based on the aggregated phase prediction.


