Reference-Point Data Segmentation Across ML Training Phases

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning models for automated vehicle or robot control struggle with overfitting when trained on limited data, leading to inadequate generalization capabilities.

Innovation Solution

A method for dividing measurement data records across different phases of training, including optimization and test phases, using reference points to assign similar data records to the same phase, ensuring diverse data usage and reducing overfitting.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If a machine learning model is trained on a finite set of training examples to achieve accurate results, then the model can reproduce target outputs for training data, but the model suffers from overfitting and inadequate generalization capabilities when encountering unseen data

Engineering Contradiction:
Improvetraining accuracyVSAvoidgeneralization capability
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent segments the training process into multiple distinct phases (optimization phases and test phases) by introducing reference points that divide the training data into different groups. Each phase trains the model on specific subsets of data while testing on others, preventing the model from overfitting to a single homogeneous dataset and improving generalization to unseen data.

Inventive Principle:
Principle #1Segmentation

2Productivity

If measurement data records are used exclusively for optimization training, then the model can be trained efficiently, but the model cannot be properly evaluated on unseen measurement data

Engineering Contradiction:
Improvetraining efficiencyVSAvoidperformance evaluation accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the measurement data records into different phases using reference points, creating distinct optimization phases and test phases. This allows simultaneous pursuit of training efficiency (through dedicated optimization phases) and accurate performance evaluation (through separate test phases with unseen data), resolving the contradiction between productivity and measurement precision.

Inventive Principle:
Principle #1Segmentation

3Quantity of substance

If all measurement data records are used in the optimization phase, then complete data utilization is achieved, but the model memorizes training data rather than extracting general knowledge

Engineering Contradiction:
Improvedata utilizationVSAvoidknowledge generalization
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent segments the measurement data records across multiple phases using reference points, ensuring that not all data is used in any single optimization phase. This segmentation prevents memorization while maintaining overall data utilization across the complete training process, thereby improving the model's ability to generalize knowledge to unseen scenarios.

Inventive Principle:
Principle #1Segmentation

4Reliability

If measurement data records are divided into different phases using reference points, then overfitting is reduced and generalization is improved, but the training process complexity increases

Engineering Contradiction:
Improvegeneralization capabilityVSAvoidtraining process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces reference points as intermediary elements that facilitate the division of measurement data records into different phases. These reference points act as mediators between the raw data and the training phases, providing a systematic and automated mechanism for data segmentation that manages complexity while achieving improved generalization.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250272615A1Division of measurement data records across the phases of the training of a machine learning model
Publication Date: 2025.08.28 ROBERT BOSCH GMBH
  • US20250272615A1 patent drawing
  • US20250272615A1 patent drawing
  • US20250272615A1 patent drawing

AI summary

A method for dividing a specified set of measurement data records for the training of a machine learning model across different specified phases of the training. Each measurement data record contains values of one or more measurement variables. The method includes: ascertaining a sequence of reference points, which cover a space of the measurement data records and do not coincide with measurement data records; for one or more of the measurement data records, ascertaining, with a specified distance measure, to which reference point this measurement data record is closest, and assigning the measurement data record to this reference point; dividing the reference points across the specified phases of the training so that one or more reference points are assigned to each phase of the training; assigning the measurement data records assigned to a reference point also to the phase of the training to which this reference point is assigned.