Edge Case Data Generation for Automotive ML Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current active learning approaches for machine learning models in automotive applications are inefficient, requiring prolonged time to obtain adequate additional training data, and suffer from scalability issues as the size of the training dataset increases, leading to decreased accuracy improvements.

Innovation Solution

The method involves identifying edge cases where machine learning models perform poorly, actively collecting and sampling relevant real-world data, and generating synthetic data to enhance training datasets, allowing for the retraining or training of new models with improved accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional active learning approaches are used to collect additional training data, then model accuracy can be improved, but the process requires prolonged time and extensive resource usage

Engineering Contradiction:
Improvemodel accuracyVSAvoidtime to obtain training data
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent generates synthetic data that copies the characteristics of real edge case data through simulation and augmentation techniques. This allows the model to learn from synthesized representations of rare scenarios without requiring extensive collection and labeling of actual real-world edge cases, significantly reducing the time required while maintaining accuracy improvements

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system proactively identifies potential edge cases and generates synthetic training data in advance before they are encountered in real deployment. By preparing synthetic examples of rare scenarios beforehand, the model is pre-trained to handle these cases, eliminating the need for prolonged data collection periods

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If the size of the training dataset is increased to improve model accuracy, then better performance is achieved, but scalability issues arise and accuracy improvements decrease

Engineering Contradiction:
Improvemodel accuracyVSAvoidscalability of accuracy improvement
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

Instead of uniformly increasing the entire training dataset, the patent applies data generation and augmentation specifically to edge case scenarios where the model performs poorly. This targeted approach focuses computational resources on improving accuracy in critical areas rather than diluting efforts across the entire dataset, maintaining scalability

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically adjusts training parameters including data sampling strategies, augmentation intensity, and synthetic data generation parameters based on model performance metrics. This allows the training process to adaptively optimize for accuracy improvements in edge cases without being constrained by fixed dataset size limitations

Inventive Principle:
Principle #35Parameter changes

3Reliability

If extensive real-world data collection is performed to address edge cases, then model robustness improves, but resource usage and complexity increase significantly

Engineering Contradiction:
Improvemodel robustnessVSAvoiddata collection and labeling complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent creates synthetic copies of edge case scenarios through simulation environments and data augmentation, avoiding the need for extensive real-world data collection and manual labeling. This maintains model robustness by providing diverse edge case examples while significantly reducing the complexity of data acquisition and preparation processes

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system automatically generates synthetic edge case data and integrates it into the training pipeline without requiring extensive manual intervention for data collection and labeling. This self-service approach to data preparation reduces operational complexity while maintaining the ability to improve model robustness through targeted edge case training

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20230351736A1Active data collection, sampling, and generation for use in training machine learning models for automotive or other applications
Publication Date: 2023.11.02 WHS ENERGY SOLUTIONS LLC
  • US20230351736A1 patent drawing
  • US20230351736A1 patent drawing
  • US20230351736A1 patent drawing

AI summary

A method includes identifying one or more edge cases associated with at least one trained machine learning model, where the at least one trained machine learning model is configured to perform at least one function related to one or more vehicles. The method also includes obtaining raw data associated with the one or more edge cases from at least one of the one or more vehicles and selecting a subset of the raw data. The method further includes generating synthetic data associated with the one or more edge cases. In addition, the method includes at least one of: retraining the at least one trained machine learning model and training at least one new machine learning model using the selected subset of raw data and the synthetic data.