Synthesizing Recognition Target Data for Machine Learning Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for generating training data for machine learning models are inefficient, leading to difficulties in improving the training efficiency of individual models, as they often include data that is not beneficial for the learning process.
Innovation Solution
An information processing method that synthesizes recognition target data into a determined region of sensing data, generating composite data with similar human-perceived characteristics, and using this composite data to identify and select training data based on recognition accuracy, thereby improving the training efficiency of the model.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional methods are used to generate training data by combining different sensors or using difference images, then the quantity of training data increases, but the individual training efficiency of the model does not improve because the generated data is not always beneficial to the model
Solution Approach 1:
The patent implements a feedback mechanism where the model evaluates generated composite data through recognition processing, and based on the recognition results, determines whether to use the composite data as training data. This feedback loop ensures that only beneficial data is selected for training, improving individual training efficiency while maintaining data quantity.
Solution Approach 2:
The system enables the model to self-evaluate the quality of generated training data through its own recognition processing. The model autonomously determines which composite data items are beneficial for its training by comparing recognition results with reference data, eliminating the need for external quality assessment.
2Quantity of substance
If composite data is generated and used as training data without verification, then the number of training data items increases, but the recognition accuracy may not improve due to inclusion of non-beneficial data
Solution Approach 1:
The patent uses recognition processing results as feedback to determine whether composite data should be used as training data. By verifying whether the model can correctly recognize the synthesized recognition target in the composite data, the system ensures that only data contributing to recognition accuracy improvement is included in the training set.
Solution Approach 2:
The patent replaces manual or mechanical data quality assessment with automated recognition processing. The model itself evaluates the quality of composite data through recognition tasks, substituting traditional quality control mechanisms with an intelligent, model-based assessment system.
3Productivity
If all generated composite data is used for training, then data accumulation speed increases, but training time and computational resources are wasted on processing non-beneficial data
Solution Approach 1:
The patent extracts only the beneficial portion of generated composite data by using recognition results as a filter. Composite data items where the model successfully recognizes the synthesized target are extracted and used as training data, while non-beneficial items are discarded, avoiding waste of training time and computational resources.
Solution Approach 2:
The model autonomously identifies and selects beneficial training data through self-evaluation of recognition results. This self-service mechanism eliminates the need for external filtering or manual selection, maintaining fast data accumulation speed while ensuring efficient use of training resources.
Data Source
AI summary
An information processing method includes: obtaining sensing data; determining a synthesis region in the sensing data in which recognition target data is to be synthesized with the sensing data; generating composite data by synthesizing the recognition target data into the synthesis region, the recognition target data having same or similar characteristics perceived by a human sensory system as the sensing data; obtaining recognition result data by providing the composite data to a model which has been trained using machine learning to recognize a recognition target; making a second determination based on the recognition result data and reference data including at least the synthesis region, the second determination being to determine whether to make a first determination, the first determination being to determine training data for the model based on the composite data; and making the first determination when it is determined in the second determination to make the first determination.


