Machine Learning Sweat Classification via Mass Spectrometry

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The identification of specific molecules in trace samples like fingerprints is challenging, making it difficult to determine characteristics such as age, gender, ethnicity, and disease states from chemical analysis.

Innovation Solution

A machine learning model is trained on raw m/z mass spectrometry data associated with quantities of interest, allowing direct determination of characteristics without identifying specific molecules present, leveraging the complexity of the data to classify individuals into groups like gender, ethnicity, age, and disease states from sweat samples.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If specific molecules are identified in trace samples like fingerprints, then chemical analysis accuracy is improved, but the complexity and difficulty of the analysis process increases significantly

Engineering Contradiction:
Improvechemical analysis accuracyVSAvoidanalysis process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the essential information needed for classification (presence/absence of molecules) while discarding the complex task of identifying specific molecular structures. This is achieved by using machine learning models that work directly with mass spectrometry data patterns without requiring molecular identification, thus reducing analysis complexity while maintaining classification accuracy

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the approach from molecular identification (qualitative analysis) to pattern recognition in mass spectrometry data (quantitative analysis). By transforming the problem from identifying what molecules are present to recognizing patterns in m/z ratios and peak intensities, the complexity is reduced while preserving the ability to determine characteristics of interest

Inventive Principle:
Principle #35Parameter changes

2Productivity

If machine learning models are trained on raw mass spectrometry data without molecular identification, then analysis time is reduced, but the reliability of determining characteristics of interest may be compromised

Engineering Contradiction:
Improveanalysis speedVSAvoidclassification reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent performs preliminary training of machine learning models on extensively labeled datasets where molecular compositions are known. This preliminary action creates pre-trained models that have learned reliable patterns associating mass spectrometry data with characteristics like age, gender, and ethnicity. When applied to new samples, these pre-trained models can quickly classify without requiring time-consuming molecular identification, thus maintaining both speed and reliability

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent incorporates feedback mechanisms where the machine learning model's predictions are continuously refined based on comparison with known reference data. The model learns from the relationship between mass spectrometry patterns and verified characteristics, adjusting its internal parameters to improve accuracy. This feedback loop ensures that even without explicit molecular identification, the classification remains reliable

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11416704B2Group classification based on machine learning analysis of mass spectrometry data from sweat
Publication Date: 2022.08.16 THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
  • US11416704B2 patent drawing
  • US11416704B2 patent drawing
  • US11416704B2 patent drawing

AI summary

Machine learning analysis of mass spectrometry spectra from human sweat samples is used to determine characteristics of interest such as age, ethnicity, gender drug use and disease state directly from the m/z data. This avoids the difficult problem of performing a full chemical analysis of human sweat samples to determine the characteristics of interest.