Correlated Variable Grouping in Mass Spectrometry Data Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In mass spectrometry data analysis, identifying and interpreting correlated variables is challenging due to the large number of variables generated, which complicates the analysis and often leads to the removal of valuable information, such as unpredictable fragments or known origins like isotopes, adducts, and charge states, prior to principal component analysis (PCA).
Innovation Solution
A computer-implemented method that groups variables after PCA by selecting a subset of principal components, defining a spatial angle around a selected variable, and assigning variables within that angle to a group, allowing for the retention and interpretation of correlated features, including the use of Pareto scaling to reduce intensity dominance and enhance interpretation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If correlated variables are removed prior to PCA, then analysis complexity is reduced, but valuable information is lost
Solution Approach 1:
The patent performs PCA first to transform the data and identify correlated variables through their spatial relationships in the transformed space, then groups them based on angular proximity. This preliminary transformation allows for informed grouping decisions that preserve valuable information while reducing complexity in subsequent analysis steps.
Solution Approach 2:
The patent segments correlated variables into distinct groups based on their spatial angles in the PCA-transformed space. By dividing the large set of variables into smaller, correlated groups, the method reduces analysis complexity while maintaining the integrity and interpretability of each group's information.
2Loss of information
If all variables are retained for interpretation, then information completeness is improved, but analysis complexity increases
Solution Approach 1:
The patent segments the complete set of variables into correlated groups based on spatial angles in PCA space. This segmentation allows all variables to be retained for interpretation while organizing them into manageable groups that reduce the overall analysis complexity and improve interpretability.
Solution Approach 2:
The patent transforms variables from the original high-dimensional space into PCA-transformed space, where correlated variables exhibit distinct spatial angular relationships. This dimensionality change enables efficient grouping and interpretation of all variables without being overwhelmed by the original data complexity.
3Use of energy by moving object
If intense variables dominate the analysis, then signal strength is improved, but interpretation accuracy decreases
Solution Approach 1:
The patent applies Pareto scaling to transform the intensity values of variables, reducing the dominance of intensely varying variables while preserving their relative relationships. This parameter transformation improves interpretation accuracy by preventing signal dominance, while the subsequent PCA and spatial grouping maintain the essential signal strength information in an organized manner.
Data Source
AI summary
According to various embodiments, variables are grouped in an unsupervised manner after principal component analysis of a plurality of variables from a plurality of samples. A number of principal components are selected. A subset principal component space is created for those components. A starting variable is selected. A spatial angle is defined around a vector extending from the origin to the starting variable. A set of one or more variables is selected within the spatial angle. The set is assigned to a group. The set is removed from further analysis. The process is repeated starting with the selection of a new starting variable until all groups are found.


