In-Cabin Human Detection Training Data Sampling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Training artificial neural networks for autonomous vehicles faces challenges in obtaining high-quality training data due to the limited range of driver movements and the high labeling cost associated with extracting images from raw video files, where many frames have little comparative value.

Innovation Solution

A method and system for efficiently sampling images from raw video files by generating feature graphs of target images, comparing them to baseline graphs to identify unexpected feature points, and using controller area network signals to augment valuable images for training datasets, thereby reducing labeling costs and improving computer vision algorithms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If every frame or indiscriminate sampling of frames from raw video feed is extracted, then the quantity of training images increases, but the labeling cost increases inordinately and many images have little comparable value

Engineering Contradiction:
Improvequantity of training imagesVSAvoidlabeling cost
Core Design Contradiction:
Quantity of substanceVSEase of manufacture

Solution Approach 1:

The system uses automated inference results to self-identify valuable images for training. The computer vision algorithm processes video frames and automatically selects mis-detected images (false positives and false negatives) as training samples, eliminating the need for manual labeling of all frames. This self-service approach reduces labeling costs while maintaining training data quality.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system extracts only the valuable subset of images from the raw video feed based on inference results. Instead of taking all frames, it selectively extracts mis-detected images that provide the most training value. This extraction principle filters out redundant frames and focuses on images that will improve the algorithm's performance.

Inventive Principle:
Principle #2Taking out (Extraction)

2Adaptability or versatility

If a wide range of driver pose samples is obtained for training data, then the training quality improves, but the difficulty of obtaining diverse samples increases due to limited driver movements in vehicle

Engineering Contradiction:
Improverange of driver posesVSAvoiddifficulty of obtaining diverse samples
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The system performs preliminary inference processing on video frames to identify potential training samples before actual training. By running initial detection and identifying mis-detected frames in advance, the system prepares a curated set of diverse pose samples without requiring manual collection. This preliminary action enables the system to overcome the limitation of limited driver movements by automatically finding diverse poses in the available video data.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses inference results as feedback to identify valuable training samples. Mis-detected images (false positives and false negatives) provide feedback about edge cases and diverse poses that the algorithm struggles with. This feedback mechanism guides the selection of training samples, ensuring that diverse and challenging driver poses are included in the training set.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10776642B2Sampling training data for in-cabin human detection from raw video
Publication Date: 2020.09.15 TOYOTA JIDOSHA KK
  • US10776642B2 patent drawing
  • US10776642B2 patent drawing
  • US10776642B2 patent drawing

AI summary

Systems and methods to efficiently and effectively train artificial intelligence and neural networks for an autonomous or semi-autonomous vehicle are disclosed. The systems and methods provide for the minimization of the labeling cost by sampling images from a raw video file which are mis-detected, i.e., false positive and false negative detections, or indicate abnormal or unexpected driver behavior. Supplemental information such as controller area network signals and data may be used to augment and further encapsulate desired images from video.