ML Training Data Partitioning for Tolerance-Controlled Generalization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for building machine learning models from limited test data of rail vehicles fail to ensure that the sum of characteristic quantities after data partitioning remains within a tolerance value, affecting generalization performance.
Innovation Solution
A method that divides teaching data sets to maintain an error between pre- and post-division characteristic quantities within a tolerance value, followed by reducing similar data to enhance diversity and create a machine learning model with high generalization performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data partitioning is performed to increase the number of training data, then the quantity of training data is improved, but the manufacturing precision of characteristic quantities deteriorates because the sum of characteristic quantities after partitioning may differ significantly from the original
Solution Approach 1:
The patent divides teaching data into multiple partitioned datasets, where each partition contains a subset of the original teaching data. This segmentation increases the number of available training datasets while maintaining the integrity of characteristic quantities through controlled division strategies that preserve the sum of characteristic quantities across partitions.
Solution Approach 2:
The patent introduces a tolerance value parameter that controls the acceptable deviation in the sum of characteristic quantities after partitioning. By adjusting this parameter, the system balances between increasing data quantity through partitioning and maintaining the precision requirements of characteristic quantities, thus resolving the contradiction between quantity and precision.
2Quantity of substance
If data augmentation is performed to increase training data quantity, then the quantity of training data is improved, but the reliability of characteristic quantities deteriorates because physically meaningful relationships may be distorted
Solution Approach 1:
The patent performs data partitioning before machine learning model training, pre-organizing the teaching data into multiple valid training datasets. This preliminary action ensures that each partitioned dataset maintains the correct characteristic quantity relationships from the outset, preventing the distortion of physically meaningful relationships that would occur with post-hoc augmentation methods.
3Manufacturing precision
If the tolerance value is set strictly to ensure precision of characteristic quantities, then the manufacturing precision is improved, but the productivity of model building deteriorates due to more restrictive data partitioning constraints
Solution Approach 1:
The patent applies a tolerance value that allows for partial deviation from the exact sum of characteristic quantities, rather than requiring strict equality. This partial action approach (allowing small deviations within tolerance) significantly increases the number of valid data partitioning schemes available, thereby improving model building productivity while still maintaining sufficient precision for reliable predictions.
Data Source
AI summary
A machine learning model building device comprises an actual operation database that holds actual operation data. The machine learning model building device creates a teaching data set including one or more pieces of teaching data based on the actual operation data obtained from the actual operation database. The machine learning model building device creates a post-division teaching data set containing a plurality of pieces of teaching data after dividing the teaching data contained in the teaching data set by dividing the teaching data so that an error between a characteristic quantity of the teaching data before division and the sum of characteristic quantities of the plurality of pieces of teaching data after division becomes less than a tolerance value; and creates the machine learning model using the post-division teaching data set.


