Machine Learning Data Segmentation for Product Design Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The accuracy of machine learning models in product design depends on comprehensive training and verification data, particularly when predicting performance across multiple product types, as existing methods often fail to ensure that verification data adequately represents the target product type, leading to inaccurate predictions.
Innovation Solution
A learning device and method that classify product data into predetermined categories, distribute it into training and verification data sets to maintain proportional representation, and use various index values to verify prediction accuracy, allowing for both overall and individual product type evaluations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If verification data is used to verify prediction accuracy of a machine learning model, then prediction accuracy can be assessed, but if the verification data does not comprehensively include data of the target product type, accurate verification cannot be performed
Solution Approach 1:
The patent segments the verification process into two distinct parts: (1) first verification that assesses overall prediction accuracy using all verification data, and (2) second verification that assesses prediction accuracy for each product type separately. This segmentation allows comprehensive verification while ensuring each product type is adequately represented and evaluated, resolving the contradiction between overall assessment and type-specific accuracy.
2Adaptability or versatility
If a machine learning model is used to design a plurality of different product types, then the model can handle diverse products, but the training data must comprehensively include data for all target product types
Solution Approach 1:
The patent segments training data into product-type-specific subsets and allocates them proportionally to both training and verification data. This ensures that each product type has adequate representation in the training data, enabling the model to learn diverse product characteristics while maintaining manageable data requirements through structured organization.
Solution Approach 2:
The patent creates a universal verification framework that handles multiple product types simultaneously. The same verification process and metrics are applied across all product types, allowing the system to verify prediction accuracy for diverse products using a unified approach that maintains consistency and comprehensiveness.
3Measurement precision
If verification is performed for each product type individually, then accurate assessment of each type can be achieved, but additional verification steps are required
Solution Approach 1:
The patent segments verification into hierarchical levels: an overall verification step that assesses global prediction accuracy, and product-type-specific verification steps that assess accuracy for each category. This segmented approach enables precise measurement of prediction accuracy at multiple levels while organizing the complexity into a structured, manageable process.
Data Source
AI summary
A learning device includes a processor configured to execute a program to generate learning source data by classifying product data, which includes a plurality of data sets that are pairs of explanatory variables indicating product materials or design matters and target variables indicating performance of the product, into each predetermined classification, distribute the learning source data into training data and verification data, generate a machine learning model on the basis of the training data, and compare a target variable included in the verification data and a target variable predicted by the machine learning model on the basis of explanatory variables of the verification data. The processor is configured to execute the program to distribute the learning source data to make a proportion of a data set for each classification in the learning source data correspond to a proportion of the data set for each classification in the training data and a proportion of the data set for each classification in the verification data.


