Convolved Peak Identification in Mass Spectrometry
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In mass spectrometry data analysis, identifying correlated variables is challenging due to the complexity of interpreting convolved peaks, which can result from multiple components, complicating biomarker discovery and metabolomics studies.
Innovation Solution
A computer-implemented method is developed to identify correlated variables after principal component analysis (PCA) by selecting a subset of principal components, defining spatial angles, and grouping variables based on significance, allowing for the interpretation of correlated features and simplification of data for further analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If multivariate statistical techniques are applied to mass spectrometry data to identify correlated variables, then the interpretation of mass spectral data is improved, but the complexity of analyzing hundreds or thousands of correlated variables increases
Solution Approach 1:
The patent segments the complex set of hundreds or thousands of correlated variables into smaller, manageable groups through hierarchical clustering. By organizing variables into clusters based on correlation patterns, the method divides the overwhelming analytical task into smaller units that can be analyzed individually, thus reducing analysis complexity while preserving interpretative quality.
Solution Approach 2:
The patent extracts key correlated variable groups from the large dataset and removes redundant variables. By identifying and extracting the most significant correlated clusters, the method eliminates unnecessary variables that complicate analysis, thereby reducing analysis complexity while maintaining the essential interpretative information.
2Device complexity
If correlated variables are removed prior to principal component analysis, then analysis is simplified, but valuable information about unpredictable fragments and compound identification is lost
Solution Approach 1:
The patent applies local quality by treating different variable groups differently based on their characteristics. Instead of uniformly removing all correlated variables, the method identifies specific clusters with different properties (e.g., unpredictable fragments vs. known isotopes) and retains those that provide valuable interpretative information while simplifying others, thus balancing complexity reduction with information retention.
Solution Approach 2:
The patent introduces asymmetry in the analysis process by applying different handling strategies to different correlated variable groups. Rather than symmetric removal of all correlated variables, the method asymmetrically preserves certain groups (like unpredictable fragments) that provide unique interpretative value while removing or simplifying others, thereby maintaining information retention while reducing overall complexity.
3Ease of operation
If a peak is analyzed as a single component, then the analysis is straightforward, but convolved peaks from multiple components cannot be properly identified or interpreted
Solution Approach 1:
The patent merges multiple correlated variables that correspond to different components contributing to a convolved peak into unified clusters. By combining these related variables based on their correlation patterns, the method enables the identification of convolved peaks as distinct groups, allowing for accurate peak identification while maintaining analytical simplicity through the clustered representation.
Data Source
AI summary
A method for identifying a convolved peak is described. A plurality of spectra is obtained. A multivariate analysis technique is used to assign data points from the plurality of spectra to a plurality of groups. A peak is selected from the plurality of spectra. If the peak includes data points assigned to two or more groups of the plurality of groups, the peak is identified as a convolved peak. Principal component analysis is one multivariate analysis technique that is used to assign data points. A number of principal components are selected. A subset principal component space is created. A data point in the subset principal component space is selected. A vector is extended from the origin of the subset principal component space to the data point. One or more data points within a spatial angle around the vector are assigned to a group.


