Probabilistic Clustering for LC-MS Data Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing clustering methods for liquid chromatography mass spectrometry (LC-MS) data, such as k-means clustering, face limitations including the need for pre-specifying the number of clusters and inability to integrate over all possible cluster center locations, and fail to properly normalize distance criteria for probability of association between data points.

Innovation Solution

A probabilistic or Bayesian approach is employed to cluster LC-MS data by determining physico-chemical properties like mass or mass-to-charge ratio and chromatographic retention time, allowing for probabilistic association and grouping of data points without pre-specifying cluster numbers, and using error determination and calibration functions to refine clustering.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If k-means clustering is used to cluster LC-MS data, then clustering can be performed, but the number of clusters must be pre-specified and it cannot integrate over all possible cluster center locations

Engineering Contradiction:
Improveflexibility in clustering approachVSAvoidcomplexity of clustering algorithm
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent changes the fundamental parameters of the clustering approach by transitioning from deterministic k-means to probabilistic Bayesian clustering. This allows the number of clusters to be inferred from the data rather than pre-specified, and integrates over all possible cluster center locations using probability distributions, thereby resolving the contradiction between adaptability and complexity.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the mechanical iterative assignment process of k-means clustering with a probabilistic Bayesian framework. Instead of deterministically assigning points to nearest centroids, the system uses probability models to estimate cluster memberships and cluster parameters, providing greater flexibility while maintaining computational tractability through Bayesian inference.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If k-means clustering uses ad hoc distance criterion, then clustering assignments can be computed, but it cannot properly normalize to give probability of association between data points

Engineering Contradiction:
Improveprobability of associationVSAvoidsimplicity of distance measure
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent replaces the simple Euclidean distance metric of k-means with a probabilistic distance model that naturally provides normalized association probabilities. The Bayesian framework transforms distance measurements into probability estimates through likelihood functions, enabling precise quantification of association between data points and clusters while maintaining ease of operation through automated probability calculations.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the nature of the distance criterion from a fixed ad hoc metric to a flexible probabilistic measure. By modeling distances in terms of probability distributions and using Bayesian inference, the system achieves proper normalization of association probabilities without sacrificing operational simplicity, as the probabilistic framework automatically handles normalization.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If probabilistic Bayesian approach is used to cluster LC-MS data, then accurate probability of association is obtained, but computational complexity increases compared to k-means

Engineering Contradiction:
Improveaccuracy of association probabilityVSAvoidcomputational time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial Bayesian inference by focusing computational efforts on the most critical parameters and using iterative approximation methods. Rather than computing exact posterior distributions for all parameters, the system uses practical approximations that achieve sufficient accuracy for LC-MS data analysis while significantly reducing computational time compared to full Bayesian treatment.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The probabilistic Bayesian framework is designed to be self-calibrating, automatically adjusting its computational complexity based on the data characteristics. The system uses evidence from the data itself to determine the appropriate level of modeling detail required, performing more complex calculations only when necessary to achieve accurate association probabilities, thereby balancing accuracy and computational efficiency.

Inventive Principle:
Principle #25Self-service

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enables accurate and flexible clustering of LC-MS data, improving the identification of analytes with differential expression levels between samples by integrating over all possible cluster center locations and providing a more accurate probability of association, thus enhancing the detection of changes in relative concentrations and intensities of peptides or proteins.

Implementation Method 1

determining a first physico-chemical property and a second physico-chemical property of components, molecules or analytes in a first sample, wherein the first physico-chemical property comprises the mass or mass to charge ratio

Methodology Applied
Scientific EffectMass spectrometry:

Implementation Method 2

the second physico-chemical property comprises the elution time, hydrophobicity, hydrophilicity, migration time, or chromatographic retention time

Methodology Applied
Scientific EffectChromatography: Chromatography

Data Source

PatentUS8515685B2Method of mass spectrometry, a mass spectrometer, and probabilistic method of clustering data
Publication Date: 2013.08.20 MICROMASS UK LTD
  • US8515685B2 patent drawing
  • US8515685B2 patent drawing
  • US8515685B2 patent drawing

AI summary

A mass spectrometer and a method of spectrometry are disclosed wherein liquid chromatography mass spectral data are probabilistically clustered on the basis of mass to charge ratio and retention time.