ML Training Data Partitioning for Tolerance-Controlled Generalization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for building machine learning models from limited test data of rail vehicles fail to ensure that the sum of characteristic quantities after data partitioning remains within a tolerance value, affecting generalization performance.

Innovation Solution

A method that divides teaching data sets to maintain an error between pre- and post-division characteristic quantities within a tolerance value, followed by reducing similar data to enhance diversity and create a machine learning model with high generalization performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data partitioning is performed to increase the number of training data, then the quantity of training data is improved, but the manufacturing precision of characteristic quantities deteriorates because the sum of characteristic quantities after partitioning may differ significantly from the original

Engineering Contradiction:
Improvenumber of training dataVSAvoidprecision of characteristic quantities
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent divides teaching data into multiple partitioned datasets, where each partition contains a subset of the original teaching data. This segmentation increases the number of available training datasets while maintaining the integrity of characteristic quantities through controlled division strategies that preserve the sum of characteristic quantities across partitions.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a tolerance value parameter that controls the acceptable deviation in the sum of characteristic quantities after partitioning. By adjusting this parameter, the system balances between increasing data quantity through partitioning and maintaining the precision requirements of characteristic quantities, thus resolving the contradiction between quantity and precision.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If data augmentation is performed to increase training data quantity, then the quantity of training data is improved, but the reliability of characteristic quantities deteriorates because physically meaningful relationships may be distorted

Engineering Contradiction:
Improvenumber of training dataVSAvoidreliability of characteristic quantities
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent performs data partitioning before machine learning model training, pre-organizing the teaching data into multiple valid training datasets. This preliminary action ensures that each partitioned dataset maintains the correct characteristic quantity relationships from the outset, preventing the distortion of physically meaningful relationships that would occur with post-hoc augmentation methods.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If the tolerance value is set strictly to ensure precision of characteristic quantities, then the manufacturing precision is improved, but the productivity of model building deteriorates due to more restrictive data partitioning constraints

Engineering Contradiction:
Improveprecision of characteristic quantitiesVSAvoidefficiency of model building
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent applies a tolerance value that allows for partial deviation from the exact sum of characteristic quantities, rather than requiring strict equality. This partial action approach (allowing small deviations within tolerance) significantly increases the number of valid data partitioning schemes available, thereby improving model building productivity while still maintaining sufficient precision for reliable predictions.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250371420A1Machine Learning Model Building Device, Machine Learning Model Building Method, and Non-Transitory Computer-Readable Storage Medium
Publication Date: 2025.12.04 HITACHI LTD
  • US20250371420A1 patent drawing
  • US20250371420A1 patent drawing
  • US20250371420A1 patent drawing

AI summary

A machine learning model building device comprises an actual operation database that holds actual operation data. The machine learning model building device creates a teaching data set including one or more pieces of teaching data based on the actual operation data obtained from the actual operation database. The machine learning model building device creates a post-division teaching data set containing a plurality of pieces of teaching data after dividing the teaching data contained in the teaching data set by dividing the teaching data so that an error between a characteristic quantity of the teaching data before division and the sum of characteristic quantities of the plurality of pieces of teaching data after division becomes less than a tolerance value; and creates the machine learning model using the post-division teaching data set.