Multichannel Detector Analysis Data Processing for Overfitting Prevention
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Analysis data mapped to lower dimensions often exhibit spurious correlations due to peaks from additives or lifestyle habits, leading to overfitting when used in neural networks or SVMs for regression or discrimination analysis, requiring a large number of samples to ignore noise components effectively.
Innovation Solution
Calculating the contribution value of each channel's output to the regression or discrimination function using partial differentiation, allowing for the identification and exclusion of channels with low correlation, thereby reducing noise and preventing overfitting by selecting only highly contributing channels for analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If analysis data mapped to lower dimensions is used in neural networks or SVMs for regression or discrimination analysis, then the analysis can be performed with reduced data complexity, but spurious correlations due to peaks from additives or lifestyle habits cause overfitting
Solution Approach 1:
The patent applies preliminary action by calculating contribution values and identifying noise channels before performing the actual regression or discrimination analysis. This preprocessing step removes channels that would cause spurious correlations and overfitting, ensuring that only high-quality channels expressing true sample characteristics are used in the subsequent machine learning models.
2Reliability
If a large number of samples are collected to ignore noise components effectively, then overfitting can be prevented, but the time and resources required for data collection increase
Solution Approach 1:
The patent extracts and removes noise channels from the analysis data by calculating contribution values for each channel and identifying those with low contribution. This extraction process eliminates the need to collect large numbers of samples to overcome noise, as the harmful channels are removed beforehand, significantly reducing data collection requirements while maintaining analysis reliability.
3Loss of information
If all channels of the detector are used for analysis, then comprehensive data is obtained, but noise components from channels not expressing sample characteristics reduce analysis accuracy
Solution Approach 1:
The patent applies local quality by differentiating between high-quality channels that express sample characteristics and low-quality noise channels. Instead of treating all channels uniformly, the method calculates contribution values to identify and selectively use only the high-quality channels, thereby maintaining data completeness from relevant sources while eliminating noise that would reduce analysis accuracy.
Data Source
AI summary
An analysis data processing method for processing analysis data collected with an analyzing device for each of a plurality of samples, by applying an analytical technique using statistical machine learning to multidimensional analysis data formed by output values obtained from a plurality of channels of a multichannel detector provided in the analyzing device, the method including: acquiring a non-linear regression or non-linear discrimination function expressing analysis data obtained for known samples; calculating a contribution value of each of the output values obtained from the plurality of channels forming the analysis data of the known samples, to the acquired non-linear regression or non-linear discrimination function, based on a differential value of the non-linear regression function or non-linear discrimination function; and identifying one or more of the plurality of channels of the detector, which are to be used for processing analysis data obtained for an unknown sample, based on the contribution value.


