Vehicle ML Data Set Selection Using Slot Error Overwrite

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating training, validation, and test data sets for machine learning in motor vehicles often result in non-causal correlations being learned, leading to false results during vehicle operation, due to inefficient data analysis and storage.

Innovation Solution

The method involves assigning memory areas with multiple slots to store data sets of measuring signal sequences, where each slot has an estimated slot error value. Data sets with lower estimated slot error values are overwritten, while those with higher errors are retained, optimizing data storage and selection for machine learning.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If all available measured data are stored and analyzed, then the completeness of training data sets is improved, but the memory capacity required and analysis time increase significantly

Engineering Contradiction:
Improvecompleteness of training data setsVSAvoidmemory capacity required
Core Design Contradiction:
ReliabilityVSVolume of stationary object

Solution Approach 1:

The patent segments the measured data into different operating variable ranges (e.g., speed ranges, temperature ranges) and assigns each range a dedicated memory area with multiple slots. This segmentation allows selective storage of representative data samples from each operating condition without storing all raw data, thereby reducing memory requirements while maintaining training data completeness across different operating conditions.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter of data representation by storing only essential measuring signal sequences and their associated operating variable ranges rather than complete raw data sets. By parameterizing the storage approach to keep only representative samples characterized by their operating conditions and key signal features, the system reduces memory capacity requirements while preserving the reliability needed for comprehensive machine learning training.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If all available measured data are analyzed, then the accuracy of machine learning models is improved, but the analysis time and computing resources increase significantly

Engineering Contradiction:
Improveaccuracy of machine learning modelsVSAvoidanalysis time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary organization of measured data during the data collection phase by automatically assigning data to appropriate memory areas based on their operating variable ranges and storing only representative samples. This preliminary action eliminates the need for time-consuming analysis of all raw data later, as the data are already structured and filtered for machine learning training, thereby reducing analysis time while maintaining model accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts only the essential and representative measuring signal sequences from the complete set of measured data by using automated selection based on operating variable ranges and error values. This extraction process removes redundant and non-informative data before the machine learning training phase, reducing the volume of data requiring analysis while preserving the accuracy needed for robust model training.

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If data sets are selectively stored based on error values, then the quality of training data is improved, but the device complexity for data management increases

Engineering Contradiction:
Improvequality of training dataVSAvoiddevice complexity for data management
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a self-service data management system where the control unit automatically assigns measured data to appropriate memory areas and determines which data sets to store or overwrite based on calculated error values and operating variable ranges. This automated self-service approach improves training data quality through error-based selection while minimizing the increase in device complexity by using rule-based automated decisions rather than complex manual management systems.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12346071B2Method and control unit for automatically selecting data sets for a method for machine learning
Publication Date: 2025.07.01 ROBERT BOSCH GMBH
  • US12346071B2 patent drawing
  • US12346071B2 patent drawing

AI summary

A method for the automated selection of data sets for a method for machine learning for detecting operating variables of a motor vehicle, in which methods, measuring signal sequences for particular operating variable ranges of the motor vehicle, are detected during the operation of the motor vehicle. In the method, the operating variable ranges are assigned memory areas, each including multiple slots, each of which is configured to store a data set containing a detected measuring signal sequence, and data sets already stored in the slots being overwritable with data sets newly detected instantaneously for the same memory area in each case. For each slot of a memory area in which a data set is stored, an estimated slot error value is formed and stored together with the measuring signal sequence. The data sets, whose estimated slot error value is comparatively low, are preferably overwritten.