Automated Key Frame Selection for AI Training Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for identifying 'key frames' in image stacks or video sequences for training AI and ML systems in ADAS and AD are time-consuming, subjective, and lack generalizability, with existing techniques failing to effectively automate the process for selecting diverse and representative frames.

Innovation Solution

An automated process using a structural similarity metric and a multi-expert system to filter image frames, removing frames with high structural similarity and designating frames with significant disagreements among expert models as 'key frames, thereby forming a reduced dataset suitable for training AI/ML systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual identification of key frames is performed, then subjective control over frame selection is maintained, but the process becomes time-consuming and prone to human error

Engineering Contradiction:
Improveframe selection accuracyVSAvoidcuration time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs self-service by automatically identifying key frames through multi-expert model disagreement analysis, eliminating the need for manual frame selection while maintaining or improving selection quality through objective computational metrics

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The manual mechanical process of frame selection is replaced with an automated computational system that uses multiple expert models and disagreement metrics to identify key frames, substituting human judgment with algorithmic analysis

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Quantity of substance

If all image frames are used for training, then comprehensive data coverage is achieved, but the training dataset becomes unmanageably large and inefficient

Engineering Contradiction:
Improvedata coverageVSAvoidtraining efficiency
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The system extracts only the essential key frames from the complete image stack by identifying frames where expert models disagree, removing redundant frames while preserving the critical information needed for effective training

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of using all frames or a fixed percentage, the system applies partial action by selectively choosing only those frames that meet the disagreement threshold criteria, achieving sufficient data coverage with a manageable subset

Inventive Principle:
Principle #16Partial or excessive action

3Device complexity

If sequential search with root key frames is used, then a simplified filtering process is achieved, but the method lacks ability to identify diverse limiting conditions

Engineering Contradiction:
Improvefiltering process complexityVSAvoidlimiting condition identification
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The multi-expert system provides multi-functionality by simultaneously evaluating multiple aspects of frame content through different expert models, enabling the system to identify diverse limiting conditions that a single sequential search method would miss

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11062455B2Data filtering of image stacks and video streams
Publication Date: 2021.07.13 VOLVO CAR CORP
  • US11062455B2 patent drawing
  • US11062455B2 patent drawing
  • US11062455B2 patent drawing

AI summary

Filtering a data set including a plurality of image frames to form a reduced “key frame” data set including a reduced plurality of “key” image frames that is suitable for use in training an artificial intelligence (AI) or machine learning (ML) system, including: removing an image frame from the plurality of image frames of the data set if a structural similarity metric of the image frame with respect to another image frame exceeds a predetermined threshold, thereby forming a reduced data set including a reduced plurality of image frames; and analyzing an object/semantic content of each of the reduced plurality of images using a plurality of dissimilar expert models and designating any image frames for which the plurality of expert models disagree related to the object/semantic content as “key” image frames, thereby forming the reduced “key frame” data set including the reduced plurality of “key” image frames.