Hierarchical Clustering for Chromatogram Peak Grouping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing methods for generating tables summarizing peaks in chromatograms require analysts to manually observe and compare multiple chromatograms, leading to significant workload and variability in accuracy depending on the analyst's knowledge and experience.
Innovation Solution
A data processing device and method that acquire detection data from a detection device, generate analysis data with peak information, and use hierarchical clustering to group peaks while prohibiting the grouping of peaks satisfying specific conditions, such as being in the same chromatogram or having different spectra, into the same cluster.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual observation and comparison of multiple chromatograms is used to generate peak summary tables, then the analyst can identify components, but the workload becomes enormous and the accuracy varies depending on analyst knowledge and experience
Solution Approach 1:
The patent replaces the manual mechanical process of observing and comparing chromatograms with an automated computational system. The computing unit automatically extracts peak information from multiple chromatograms, performs clustering analysis, and generates summary tables without human intervention, thereby eliminating the time-consuming manual work while maintaining or improving accuracy through consistent algorithmic processing.
Solution Approach 2:
The system enables self-service data analysis by automatically processing chromatogram data without requiring analyst intervention. The computing unit independently performs peak detection, information extraction, clustering, and table generation, allowing the data processing system to serve itself rather than relying on external human analysts.
2Productivity
If manual observation and comparison of multiple chromatograms is used to generate peak summary tables, then the analyst can identify components, but the accuracy varies depending on analyst knowledge and experience
Solution Approach 1:
The patent replaces the manual mechanical process of observing and comparing chromatograms with an automated computational system. The computing unit automatically extracts peak information from multiple chromatograms, performs clustering analysis, and generates summary tables without human intervention, thereby eliminating the time-consuming manual work while maintaining or improving accuracy through consistent algorithmic processing.
Solution Approach 2:
The system incorporates feedback mechanisms where the computing unit iteratively processes peak information, performs clustering based on extracted features, and refines the grouping of peaks across multiple chromatograms. This automated feedback loop ensures consistent and reproducible results that are not subject to the variability of human analyst judgment.
3Ease of operation
If hierarchical clustering is used to group peaks from multiple chromatograms, then the workload of analysts is reduced, but peaks satisfying specific conditions may be incorrectly grouped into the same cluster
Solution Approach 1:
The patent extracts and isolates specific conditions that should prevent peak grouping, such as peaks belonging to the same chromatogram or peaks with different spectral characteristics. By explicitly identifying and separating these cases before clustering, the system ensures that such peaks are excluded from being grouped together, thereby maintaining grouping accuracy while still utilizing automated hierarchical clustering for the remaining peaks.
Solution Approach 2:
The system applies preliminary anti-action by pre-identifying and marking peaks that should not be grouped together based on specific conditions (e.g., same chromatogram origin, different spectra). This preliminary restriction prevents erroneous grouping before the hierarchical clustering process begins, ensuring that the automation does not compromise the reliability of peak identification.
Data Source
AI summary
The data processing device includes a data acquisition unit that acquires detection data indicating signal intensity corresponding to a component in a sample detected by the detection device, and a computing unit that processes the detection data acquired by the data acquisition unit. The computing unit is configured to: generate a plurality of analysis data including a peak of the signal intensity based on the detection data; generate a plurality of clusters by grouping the peak included in each of the plurality of analysis data using hierarchical clustering based on peak information corresponding to the peak; and prohibit grouping a plurality of peaks included in the same analysis data, into the same cluster in the hierarchical clustering.


