Video Capture Annotation for Closed Imaging ML Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Proprietary imaging machines, such as MRI and X-ray scanners, operate as closed systems, limiting data sharing and integration with other systems, making it difficult to automate processes or integrate with third-party software.

Innovation Solution

A system that captures raw images from proprietary machines and human interactions to train a machine-learned model, allowing continuous improvement through user feedback and retraining.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If proprietary imaging machines operate as closed systems, then system reliability and specialized functionality are maintained, but data sharing and integration with other systems are limited

Engineering Contradiction:
Improvesystem reliabilityVSAvoiddata sharing capability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent introduces an intermediary system that captures video output from proprietary imaging machines and human operator interactions, serving as a mediator between the closed proprietary system and external machine learning systems. This intermediary enables data sharing without compromising the proprietary nature of the imaging machine itself.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system creates a copy of the video output stream and operator interactions for training purposes, rather than requiring direct access to the proprietary imaging machine's internal data. This copying approach enables data sharing while maintaining the closed system architecture of the original imaging machine.

Inventive Principle:
Principle #26Copying

2Measurement precision

If manual review of images by operators is performed, then detection accuracy for targeted subject matter is maintained, but productivity and processing speed are reduced

Engineering Contradiction:
Improvedetection accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent implements a feedback loop where operator interactions with the video stream are captured and used to train machine learning models. The models then provide automated detections that feed back into the system, creating an iterative improvement process that enhances both speed and accuracy over time.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system replaces the mechanical process of manual visual review with an automated machine learning-based detection system. Operators initially provide manual annotations that train the ML models, which then perform automated detection, substituting the manual mechanical process with an automated intelligent system.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If machine learning models are trained using operator interactions, then adaptability and continuous improvement are enhanced, but device complexity and data processing requirements increase

Engineering Contradiction:
Improvemodel adaptabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the system into distinct functional components: the proprietary imaging machine, the video capture system, the operator interaction interface, and the machine learning training system. This segmentation reduces overall complexity by making each component independent and easier to manage while enabling the complex adaptive functionality through their coordinated interaction.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20260072161A1Machine-Learning Models for Integrated Video Capture and Annotation System
Publication Date: 2026.03.12 MATROID INC
  • US20260072161A1 patent drawing
  • US20260072161A1 patent drawing
  • US20260072161A1 patent drawing

AI summary

A system accesses a first video stream from an internal scanning device (e.g., an X-ray scanner) that scans objects or individuals. It also accesses a second video stream from a capturing device that records a human operator reviewing and interacting with the first stream on a display to identify targeted subject matter. The system then identifies the targeted subject matter based on the operator's interactions and constructs a training dataset based on the identified targeted subject matter. Using this training dataset, the system trains a machine-learning model to identify the targeted subject matter in future video streams from scanning devices.