Vehicle Component ML Training Data Balancing by Cluster Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning methods for vehicle components face inefficiencies due to unevenly distributed data sets, leading to poor performance and robustness issues, especially in safety-critical environments, where high-performance and rapid training are essential.

Innovation Solution

A computer-implemented method that generates a training data set by clustering multidimensional data points using a cluster algorithm, focusing on equal representation of underrepresented scenarios, reducing training time and resources, and using validation data to improve algorithm performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If large data sets are used to improve machine learning performance, then the performance and robustness are improved, but the computing effort, training time, and resource requirements increase significantly

Engineering Contradiction:
Improvemachine learning performanceVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent extracts and removes redundant data points from the training data set by comparing similarity between data points and identifying those that provide minimal additional information. This extraction process reduces the overall data set size while preserving the essential information needed for reliable machine learning performance.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameter of data set size by dynamically adjusting it based on redundancy analysis. Instead of using a fixed large data set, the system optimizes the data set size by removing redundant instances, thereby reducing training time while maintaining performance through intelligent parameter adjustment.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If large data sets are used to improve machine learning performance, then the performance and robustness are improved, but the computing resources and processing power required increase

Engineering Contradiction:
Improvemachine learning performanceVSAvoidcomputing resources
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent extracts and removes redundant data points from the training data set by comparing similarity between data points and identifying those that provide minimal additional information. This extraction process reduces the overall data set size while preserving the essential information needed for reliable machine learning performance.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameter of data set size by dynamically adjusting it based on redundancy analysis. Instead of using a fixed large data set, the system optimizes the data set size by removing redundant instances, thereby reducing training time while maintaining performance through intelligent parameter adjustment.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If iterative manual data set generation is performed to improve performance, then the data quality can be optimized, but the process becomes time-consuming and error-prone

Engineering Contradiction:
Improvedata qualityVSAvoiddata preparation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements self-service by enabling the system to automatically analyze data redundancy and remove duplicate instances without manual intervention. The automated redundancy detection and removal process eliminates the need for time-consuming manual data set generation while maintaining high data quality through objective similarity metrics.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical data preparation processes with automated computational methods. Instead of iterative manual data set generation, the system uses algorithmic redundancy detection and removal, substituting human effort with automated processing that is both faster and more consistent.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Reliability

If all available data is used for training to improve performance, then the model robustness is improved, but the training process becomes ineffective and time-consuming

Engineering Contradiction:
Improvemodel robustnessVSAvoidtraining efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent extracts and removes redundant data points from the training data set by comparing similarity between data points and identifying those that provide minimal additional information. This extraction process reduces the overall data set size while preserving the essential information needed for reliable machine learning performance.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameter of data set size by dynamically adjusting it based on redundancy analysis. Instead of using a fixed large data set, the system optimizes the data set size by removing redundant instances, thereby reducing training time while maintaining performance through intelligent parameter adjustment.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11970178B2Computer-implemented method for machine learning for operating a vehicle component, and method for operating a vehicle component
Publication Date: 2024.04.30 VOLKSWAGEN AG
  • US11970178B2 patent drawing
  • US11970178B2 patent drawing
  • US11970178B2 patent drawing

AI summary

The invention relates to a method for preparing or generating a training data set for machine learning to operate a vehicle component. Provided multidimensional data points are divided up in a first step by dividing up the plurality of data points into multidimensional clusters by using a cluster algorithm. Then a training data set is generated by selecting data points from the basic training data set. The selection comprises determining a smallest cluster among the plurality of clusters with the lowest number of data points. Furthermore, at least one subset of the data points of the smallest cluster is provided for the training data set. In another step, a subset of data points is selected from each of the other clusters for the training data set, wherein the number of selected data points of each other cluster corresponds to the number of selected data points of the smallest cluster.