Chromatogram Peak Separation via Modified EM Algorithm

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing chromatogram data processing methods struggle to automatically separate peaks originating from multiple components or those with shoulder peaks, requiring manual expertise and being unsuitable for peaks with complex waveform profiles.

Innovation Solution

A modified Expectation Maximization (EM) algorithm for Gaussian mixture models is applied to three-dimensional chromatogram data, allowing for automatic peak separation by iteratively estimating chromatogram and spectrum waveforms, correcting parameters, and filtering components, even in cases of overlapping peaks or shoulder peaks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual peak separation methods are used, then expertise and skill are required for accurate separation, but the process becomes cumbersome and time-consuming

Engineering Contradiction:
Improvepeak separation accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs automatic peak separation through the EM algorithm, allowing the data processing system to serve itself without requiring manual expert intervention. The algorithm automatically identifies peak boundaries, separates overlapping peaks, and determines component concentrations based on the chromatogram data and spectral information.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The manual mechanical process of peak separation by expert analysts is replaced with an automated computational system using the EM algorithm. This substitutes human expertise with a mathematical model that iteratively estimates peak parameters and separates components based on spectral and chromatographic data.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If existing automatic peak separation methods are used, then processing speed is improved, but they fail to handle peaks with complex waveform profiles such as shoulder peaks

Engineering Contradiction:
Improveprocessing speedVSAvoidseparation accuracy for complex peaks
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system uses dynamic iterative estimation through the EM algorithm that adapts to different peak shapes and complexities. The algorithm dynamically adjusts peak parameters including position, width, and amplitude based on the actual data, allowing it to handle various waveform profiles including shoulder peaks, asymmetric peaks, and overlapping peaks of different intensities.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes and optimizes multiple parameters including peak position, width, amplitude, and spectral characteristics through iterative estimation. By adjusting these parameters based on the chromatogram data and spectral information, the system accurately separates peaks with complex waveform profiles that static methods cannot handle.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If manual peak purity determination is performed, then accurate identification of target components is achieved, but the process requires analyst expertise and is not automated

Engineering Contradiction:
Improvepeak purity determination accuracyVSAvoidautomation level
Core Design Contradiction:
Measurement precisionVSExtent of automation

Solution Approach 1:

The system automatically determines peak purity by comparing spectral information across different time points and using the EM algorithm to identify pure component spectra. The system self-evaluates whether peaks originate from single or multiple components without requiring manual analyst determination, achieving both automation and accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Spectral information serves as an intermediary that enables automatic peak purity determination. By analyzing spectral characteristics and comparing them across different retention times, the system uses spectral data as a mediator to automatically identify whether peaks are pure or mixed, replacing manual visual or computational assessment.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Extent of automation

If deconvolution processing is used for peak separation, then some automated separation is achieved, but it cannot handle shoulder peaks or peaks with tailing

Engineering Contradiction:
Improveautomated separation capabilityVSAvoidhandling of complex peak shapes
Core Design Contradiction:
Extent of automationVSAdaptability or versatility

Solution Approach 1:

The EM algorithm provides dynamic adaptation to various peak shapes through iterative parameter estimation. Unlike fixed deconvolution methods, the system dynamically adjusts peak model parameters based on the actual chromatogram data, enabling it to handle shoulder peaks, asymmetric peaks, and peaks with tailing by continuously refining the peak representation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system incorporates spectral information as an additional dimension beyond the traditional chromatographic time dimension. By using spectral data from multichannel detectors, the system gains another dimension for peak separation and purity determination, enabling it to distinguish overlapping peaks that have different spectral characteristics even when they overlap in the time dimension.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS10416134B2Chromatogram data processing method and chromatogram data processing apparatus
Publication Date: 2019.09.17 SHIMADZU CORP
  • US10416134B2 patent drawing
  • US10416134B2 patent drawing
  • US10416134B2 patent drawing

AI summary

The EM algorithm for a Gaussian mixture model is applied to the separation of peaks that overlap one another on a chromatogram. However, the number of overlapping components is unknown. Thus, a suitable number of models is set, and the fitting of model parameters is performed while an actually measured signal is appropriately divided for each model by the EM algorithm. Then, when a solution converges, a determination is made as to whether a peak-like waveform is present in a residue signal that is not divided. When the peak-like waveform is present, a peak model is added. The EM algorithm is executed again. In the M step, optimization is performed using, not only a simple Gaussian function, but also a modified Gaussian function assuming a tailing. In the M step, the estimation of a spectrum assuming a chromatogram and the estimation of a chromatogram assuming a spectrum are repeatedly performed.