Analytical Data Analyzer Simulated Training Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing analytical data analysis methods using machine learning face challenges in achieving high accuracy due to limited data availability, particularly in scientific analysis, where data variation can lead to reduced discrimination accuracy and over-fitting.
Innovation Solution
The method generates simulated data with controlled variations within specific ranges that do not affect identification, allowing for increased data quantity and improved machine learning accuracy by accounting for measurement-related factors such as analyzer variations, sample characteristics, and environmental differences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the amount of training data is increased to improve machine learning accuracy, then discrimination accuracy is improved, but in data-scarce scenarios like biological sample analysis, it is difficult to acquire sufficient typical data for effective machine learning
Solution Approach 1:
The patent creates simulated data copies by adding controlled variations to existing measurement data. These simulated data copies mimic real measurement scenarios including noise and variations, allowing the machine learning model to train on expanded datasets without requiring additional physical samples. This resolves the contradiction by generating sufficient training data through systematic duplication and modification of limited original data.
Solution Approach 2:
The patent applies parameter changes by systematically varying measurement parameters such as adding different levels of noise, adjusting baseline offsets, and modifying spectral characteristics within realistic ranges. These parameter transformations generate diverse simulated training data from limited original measurements, enabling effective machine learning training while maintaining data fidelity and avoiding the need to acquire large quantities of rare biological samples.
2Quantity of substance
If data variation is added to increase data quantity for machine learning, then the amount of training data increases, but adding excessive variation may reverse discrimination results and reduce accuracy
Solution Approach 1:
The patent applies partial action by adding only the necessary amount of variation to training data - enough to improve model robustness and prevent overfitting, but not so much that it reverses discrimination results. The variation is carefully controlled within ranges that reflect actual measurement conditions, applying just sufficient transformation to enhance learning without exceeding the threshold that would corrupt the underlying patterns needed for accurate discrimination.
3Productivity
If machine learning is performed with limited data, then the analysis can be completed with available resources, but the accuracy of machine learning is easily reduced due to data variation
Solution Approach 1:
The patent applies preliminary action by pre-processing existing measurement data to create expanded simulated datasets before performing machine learning. By generating varied training samples in advance through controlled parameter modifications and noise additions, the system prepares robust training data that accounts for potential variations, enabling accurate discrimination while maintaining analysis efficiency with limited original data resources.
Data Source
AI summary
This analytical data analysis method uses machine learning of analysis result data (31) measured by an analyzer (1), and includes generating simulated data (32) in which a data variation has been added to the analysis result data (31) within a range that does not affect identification, performing the machine learning using the generated simulated data (32), and performing discrimination using a discrimination criterion (23b) obtained through the machine learning.


