Image-to-Video Bootstrapping for Computer Vision Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern computer vision systems require large-scale annotated video datasets, which are time-consuming and resource-intensive to create, limiting their speed and accuracy in video understanding applications.

Innovation Solution

A novel image-to-video bootstrapping technique that reduces annotation time and computational resources by iteratively training a visual recognition model using image datasets and unlabeled video data, with machine-in-the-loop feedback for improving label accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If visual recognizers are trained on image datasets and applied to video domain, then training speed is improved, but accuracy deteriorates due to different visual characteristics between images and videos

Engineering Contradiction:
Improvetraining speedVSAvoidrecognition accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary representation called 'video-like images' that are generated from actual video frames. These intermediary images retain the compression artifacts and visual characteristics of video while being processable as images. The visual recognizer trained on these intermediary images acts as a bridge, allowing image-based training methods to work effectively on video data without sacrificing accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms the training data parameters by converting actual video frames into 'video-like images' through controlled compression and processing. This parameter transformation allows the training data to maintain the essential visual characteristics of video (including compression artifacts) while being compatible with image-based training pipelines, thus resolving the mismatch between training data and target application domain.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If video datasets are annotated by human labelers inspecting every frame, then label accuracy is improved, but time consumption and computational resources increase significantly

Engineering Contradiction:
Improvelabel accuracyVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by automatically generating 'video-like images' from video frames before the annotation process. These pre-processed images serve as ready-to-use training data that captures the essential visual characteristics of video. This preliminary processing eliminates the need for manual frame-by-frame inspection while maintaining annotation quality, as the generated images can be directly used for training visual recognizers.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system performs self-service by automatically transforming video frames into annotated training data through the video-like image generation process. The computational pipeline autonomously handles the conversion and preparation of training data without requiring human intervention for each frame, thereby reducing annotation time and computational resource requirements while maintaining accuracy.

Inventive Principle:
Principle #25Self-service

3Quantity of substance

If video frames are compressed using codecs to reduce file size, then storage efficiency is improved, but visual quality deteriorates with blurry frames

Engineering Contradiction:
Improvefile sizeVSAvoidvisual quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent converts the harmful effect of compression artifacts into a benefit by deliberately generating 'video-like images' that incorporate controlled compression. Instead of treating compression artifacts as noise to be eliminated, the methodology embraces them as essential characteristics that make the training data more representative of actual video conditions. This allows the visual recognizer to learn from data that closely mimics real-world video quality, improving robustness while working with compressed video files.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Data Source

PatentUS10740394B2Machine-in-the-loop, image-to-video computer vision bootstrapping
Publication Date: 2020.08.11 VERIZON PATENT & LICENSING INC
  • US10740394B2 patent drawing
  • US10740394B2 patent drawing
  • US10740394B2 patent drawing

AI summary

Disclosed are systems and methods for improving interactions with and between computers in content searching, hosting and/or providing systems supported by or configured with devices, servers and/or platforms. The disclosed systems and methods provide a novel machine-in-the-loop, image-to-video bootstrapping framework that harnesses a training set built upon an image dataset and a video dataset in order to efficiently produce an accurate training set to be applied to frames of videos. The disclosed systems and methods reduce the amount of time required to build the training dataset, and also provide mechanisms to apply the training dataset to any type of content and for any type of recognition task.