Frame-to-Video Encoder for Medical Image Annotation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current AI technologies for medical image analysis require extensive frame-level annotations, which are time-consuming and costly, and using video-level annotations is technically challenging due to incomplete annotation information for training ML models.

Innovation Solution

The system employs a frame-to-video feature encoder jointly trained with frame-level and video-level annotations to generate accurate video-level predictions, allowing for improved frame-level and video-level prediction accuracy without the need for extensive frame-by-frame annotations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If frame-level annotations are used for training ML models, then prediction accuracy is improved, but annotation time and cost increase significantly

Engineering Contradiction:
Improveprediction accuracyVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the annotation task into two levels: video-level annotations (coarse-grained, less time-consuming) and frame-level annotations (fine-grained, more accurate). By using video-level annotations as primary training data and frame-level annotations as supplementary guidance, the system achieves accurate feature localization without requiring extensive frame-by-frame annotation, thus resolving the contradiction between prediction accuracy and annotation time

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary component (the frame-to-video feature encoder and multi-scale feature fusion mechanism) that bridges video-level annotations and frame-level predictions. This intermediary enables the model to infer frame-level features from video-level context, reducing the need for direct frame-level annotation while maintaining prediction accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If video-level annotations are used for training, then annotation efficiency is improved, but training complexity increases due to incomplete annotation information

Engineering Contradiction:
Improveannotation efficiencyVSAvoidtraining complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent transitions from traditional 2D frame-level annotation space to a temporal dimension by utilizing video-level annotations that span multiple frames. This dimensional change allows the model to learn temporal patterns and contextual relationships, enabling accurate predictions without requiring annotations for every individual frame, thus improving efficiency while managing training complexity through temporal context utilization

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Reliability

If extensive frame-level annotations are collected, then model training quality is improved, but data preparation cost and time increase

Engineering Contradiction:
Improvemodel training qualityVSAvoiddata preparation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary action by collecting video-level annotations first, which are less time-consuming and can be obtained more efficiently. These video-level annotations serve as a foundation for model training, and frame-level annotations are only collected selectively where needed to enhance specific features. This preliminary annotation strategy maintains model training quality while significantly reducing data preparation time

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240404254A1Video-level medical image annotation
Publication Date: 2024.12.05 KONINKLIJKE PHILIPS NV
  • US20240404254A1 patent drawing
  • US20240404254A1 patent drawing
  • US20240404254A1 patent drawing

AI summary

Techniques for training models, using video-level annotations as additional supervision, to generate both video-level and frame-level predictions based on medical images are disclosed. In some examples, medical imaging data is received including frame-level annotations and video-level annotations. A training dataset may be generated comprising frame-level ground truth data and video-level ground truth data, and the model is trained, using the training dataset, to generate frame-level feature localizations/segmentations and/or video-level feature predictions on new medical imaging data. In some examples, a model includes a frame-to-video feature encoder that learns to generate video-level predictions from frame-level predictions. The frame-to-video feature encoder may be jointly trained based on video-level annotations along with frame-level annotations during training.