Principal Component Analysis Variable Grouping for Mass Spectrometry Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The high dimensionality of mass spectrometry data from a large number of samples poses challenges for data processing, limiting the use of techniques like independent component analysis and linear discriminant analysis and requiring significant computer resources.

Innovation Solution

The implementation of principal component analysis (PCA) followed by variable grouping, which reduces dimensionality and identifies correlated variables, allowing for efficient processing and interpretation of mass spectrometry data through the creation of group representations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If principal component analysis followed by variable grouping is applied, then dimensionality is reduced and processing efficiency is improved, but information loss may occur during dimensionality reduction

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidinformation loss
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent segments the high-dimensional variable space into multiple clusters or groups of correlated variables. Each cluster represents a subset of variables that are highly correlated with each other, allowing the system to process data in manageable segments rather than as a single high-dimensional entity, thus improving processing efficiency while preserving the correlation structure within each segment

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent merges highly correlated variables into representative variables or clusters. By combining variables that convey similar information, the system reduces dimensionality while maintaining the essential information content, as the merged variables capture the collective behavior of the original correlated variables

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If high dimensionality data is processed using traditional methods, then comprehensive analysis is achieved, but significant computer resources are required and processing time increases

Engineering Contradiction:
Improveanalysis comprehensivenessVSAvoidcomputer resources
Core Design Contradiction:
Measurement precisionVSUse of energy by stationary object

Solution Approach 1:

The patent extracts and removes redundant information from the high-dimensional data by identifying and eliminating highly correlated variables. By extracting only the essential, non-redundant variables that contribute uniquely to the analysis, the system maintains comprehensive analysis capability while significantly reducing the computational burden on processing systems

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If high dimensionality data is processed using traditional methods, then complete variable analysis is possible, but processing time increases and timely processing becomes difficult

Engineering Contradiction:
Improvevariable analysis completenessVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary variable grouping and dimensionality reduction before conducting the main analysis. By pre-organizing variables into correlated clusters and selecting representative variables in advance, the system prepares the data structure to enable faster subsequent processing while ensuring that no important variable relationships are lost, thus achieving both completeness and timeliness

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8180581B2Systems and methods for identifying correlated variables in large amounts of data
Publication Date: 2012.05.15 MDS ANALYTICAL TECH A BUSINESS UNIT OF MDS INS
  • US8180581B2 patent drawing
  • US8180581B2 patent drawing
  • US8180581B2 patent drawing

AI summary

Groups of correlated representations of variables are identified from a large amount of spectrometry data. A plurality of samples is analyzed and a plurality of measured variables is obtained from a spectrometer. A processor executes a number of steps. The plurality of measured variables is divided into a plurality of measured variable subsets. Principal component analysis followed by variable grouping (PCVG) is performed on each measured variable subset, producing one or more group representations for each measured variable subset and a plurality of group representations for the plurality of measured variable subsets. While the total number of the plurality of group representations is greater than a maximum number, the plurality of group representations is divided into a plurality of representative subsets and PCVG is performed on each subset. PCVG is performed on the remaining the plurality of group representations, producing a plurality of groups of correlated representations of variables.