Teaching Data Extending Device for ML
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Collecting a large amount of data in a real environment for machine learning is time-consuming and burdensome, particularly for deep learning methods, which require extensive data management over long periods.
Innovation Solution
A teaching data extending device and method that analyze the relationships between features in existing teaching data, select uncorrelated features, and generate new teaching data by replacing the values of selected features with values from other data classified in the same class, thereby reducing the need for extensive data collection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a large amount of data is collected in a real environment for machine learning, then the analysis accuracy of the model is improved, but the data collection time and management burden increase significantly
Solution Approach 1:
The patent creates artificial teaching data by copying and transforming existing real teaching data. The feature selecting unit identifies important features, and the teaching data extending unit generates new data samples by replacing selected feature values with those from other teaching data of the same class, thereby copying successful data patterns to expand the dataset without physical collection
Solution Approach 2:
The patent transforms the parameters of existing teaching data by selectively replacing feature values. The system changes specific feature parameters while maintaining class labels, creating varied but valid teaching data samples that expand the dataset while preserving the underlying data distribution and relationships
2Productivity
If data is collected over a long period to ensure sufficient quantity, then the model learning effectiveness is improved, but the equipment management burden and operational complexity increase
Solution Approach 1:
The system performs self-service data generation by automatically selecting features and generating artificial teaching data without requiring external data collection equipment or human intervention. The feature selecting unit and teaching data extending unit work autonomously to expand the dataset, eliminating the need for continuous equipment operation and management
3Quantity of substance
If more teaching data is generated through feature replacement, then the data quantity is increased, but the risk of losing original data characteristics may increase
Solution Approach 1:
The patent applies local quality by selectively replacing only certain feature values while preserving others. The feature selecting unit identifies which features should be replaced based on their importance and relationships, ensuring that critical data characteristics are maintained while still generating data variation. This localized transformation approach maintains data fidelity while expanding quantity
Data Source
AI summary
A teaching data extending device includes: a relationship acquiring unit that obtains a relationship between a plurality of features included in each of a plurality of teaching data; a feature selecting unit that selects any one or more of the plurality of features based on the relationship; and a teaching data extending unit that generates, for one or more teaching data, new teaching data in which a value of the feature selected by the feature selecting unit is replaced with a value of the feature in another teaching data classified in a same class.


