Vehicle Component ML Training Data Balancing by Cluster Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning methods for vehicle components face inefficiencies due to unevenly distributed data sets, leading to poor performance and robustness issues, especially in safety-critical environments, where high-performance and rapid training are essential.
Innovation Solution
A computer-implemented method that generates a training data set by clustering multidimensional data points using a cluster algorithm, focusing on equal representation of underrepresented scenarios, reducing training time and resources, and using validation data to improve algorithm performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If large data sets are used to improve machine learning performance, then the performance and robustness are improved, but the computing effort, training time, and resource requirements increase significantly
Solution Approach 1:
The patent extracts and removes redundant data points from the training data set by comparing similarity between data points and identifying those that provide minimal additional information. This extraction process reduces the overall data set size while preserving the essential information needed for reliable machine learning performance.
Solution Approach 2:
The patent changes the parameter of data set size by dynamically adjusting it based on redundancy analysis. Instead of using a fixed large data set, the system optimizes the data set size by removing redundant instances, thereby reducing training time while maintaining performance through intelligent parameter adjustment.
2Reliability
If large data sets are used to improve machine learning performance, then the performance and robustness are improved, but the computing resources and processing power required increase
Solution Approach 1:
The patent extracts and removes redundant data points from the training data set by comparing similarity between data points and identifying those that provide minimal additional information. This extraction process reduces the overall data set size while preserving the essential information needed for reliable machine learning performance.
Solution Approach 2:
The patent changes the parameter of data set size by dynamically adjusting it based on redundancy analysis. Instead of using a fixed large data set, the system optimizes the data set size by removing redundant instances, thereby reducing training time while maintaining performance through intelligent parameter adjustment.
3Reliability
If iterative manual data set generation is performed to improve performance, then the data quality can be optimized, but the process becomes time-consuming and error-prone
Solution Approach 1:
The patent implements self-service by enabling the system to automatically analyze data redundancy and remove duplicate instances without manual intervention. The automated redundancy detection and removal process eliminates the need for time-consuming manual data set generation while maintaining high data quality through objective similarity metrics.
Solution Approach 2:
The patent replaces manual mechanical data preparation processes with automated computational methods. Instead of iterative manual data set generation, the system uses algorithmic redundancy detection and removal, substituting human effort with automated processing that is both faster and more consistent.
4Reliability
If all available data is used for training to improve performance, then the model robustness is improved, but the training process becomes ineffective and time-consuming
Solution Approach 1:
The patent extracts and removes redundant data points from the training data set by comparing similarity between data points and identifying those that provide minimal additional information. This extraction process reduces the overall data set size while preserving the essential information needed for reliable machine learning performance.
Solution Approach 2:
The patent changes the parameter of data set size by dynamically adjusting it based on redundancy analysis. Instead of using a fixed large data set, the system optimizes the data set size by removing redundant instances, thereby reducing training time while maintaining performance through intelligent parameter adjustment.
Data Source
AI summary
The invention relates to a method for preparing or generating a training data set for machine learning to operate a vehicle component. Provided multidimensional data points are divided up in a first step by dividing up the plurality of data points into multidimensional clusters by using a cluster algorithm. Then a training data set is generated by selecting data points from the basic training data set. The selection comprises determining a smallest cluster among the plurality of clusters with the lowest number of data points. Furthermore, at least one subset of the data points of the smallest cluster is provided for the training data set. In another step, a subset of data points is selected from each of the other clusters for the training data set, wherein the number of selected data points of each other cluster corresponds to the number of selected data points of the smallest cluster.


