GC-MS (Gas Chromatography-Mass Spectrometer) co-elution peak automatic enhanced analysis method for complex plant matrix sample

By combining the DEMCR-ALS algorithm with Gaussian smoothing and clustering strategies, the signal interference problem of co-elution peaks in complex plant matrix samples was solved, achieving efficient compound separation and accurate quantification, and improving the automation level of GC-MS analysis.

CN121499680APending Publication Date: 2026-02-10ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511615297.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing technologies suffer from severe interference in co-elution peak signals when processing complex plant matrix samples, making it difficult to separate target compounds, affecting detection sensitivity and quantitative accuracy. Furthermore, existing tools are highly dependent on initial estimates and have low automation levels.

Method used

The Dynamic Elimination Multivariate Curve-Resolved Alternating Least Squares (DEMCR-ALS) algorithm is adopted, combined with Gaussian smoothing peak detection, EIC clustering, ICA and ITTFA strategies, to automatically resolve TIC peaks, reduce dependence on initial estimates and improve the separation of co-elution peaks.

Benefits of technology

It enables efficient and automated analysis of complex plant matrix samples, improves the separation and detection accuracy of compounds, reduces dependence on initial estimates, and enhances the performance of quantitative and qualitative assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121499680A_ABST
    Figure CN121499680A_ABST
Patent Text Reader

Abstract

The invention provides an automatic enhanced analysis method for a GC-MS co-elution peak of a complex plant matrix sample. The automatic enhanced analysis method comprises the following steps: determining a peak analysis range of a single TIC peak based on a Gaussian smoothing peak detection method; four strategies are developed based on EIC clustering, independent component analysis (ICA), iterative target conversion factor analysis (ITTFA) and a model peak to construct an initial chromatogram; component separation is realized through a dynamic elimination multivariate curve resolution alternating partial least squares (DEMCR-ALS) algorithm, so that pure chromatography and mass spectrum are extracted. According to the method, a plurality of initialization strategies are adopted to improve the separation degree of a co-elution peak; the dependence on an inherent initial estimation value in a traditional method is reduced; especially, an initial chromatogram construction strategy based on a model peak shows excellent performance in both qualitative and quantitative evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of gas chromatography-mass spectrometry data analysis technology, and particularly relates to an automatic enhancement analysis method for GC-MS co-elution peaks of complex plant matrix samples. Background Technology

[0002] Gas chromatography-mass spectrometry (GC-MS), due to its high separation efficiency and sensitive detection capabilities, has become one of the most widely used analytical platforms for high-throughput identification and quantification of semi-volatile and volatile compounds in complex mixtures. However, despite its versatility, GC-MS still faces significant challenges in analyzing highly complex plant matrices. These complex plant matrices typically refer to plant extracts containing hundreds or even thousands of chemical components (such as sugars, amino acids, organic acids, and volatile aroma components), with significant differences in physicochemical properties and a wide concentration range among these components. Due to the limited separation capabilities of chromatographic columns and the inherent complexity of samples, co-elution of analytes remains common, leading to mutual interference of target compound signals, difficulty in separation, and significant matrix effects, thus affecting detection sensitivity and quantification accuracy. Solving the problem of accurate analysis of complex plant matrices is crucial for promoting its in-depth application in fields such as food traceability and metabolomics.

[0003] To address the challenge of separating co-elution peaks, various computational tools such as AMDIS, eRah, ADAP-GC, and MS-DIAL have been developed to extract components from the raw GC-MS dataset, most of which employ model-peak-based strategies. However, their performance is highly dependent on the accurate identification of representative model peaks, which is an inherently challenging task when dealing with complex co-elution signals.

[0004] Another more robust strategy for total ion chromatography (TIC) peak resolution is multivariate curve resolution (MCR), which utilizes the bilinear structure of GC-MS data to recover pure chromatographic and spectral curves. Several MCR-based algorithms have been developed, including evolutionary factor analysis (EFA), window factor analysis (WFA), heuristic evolutionary latent projection (HELP), sub-window factor analysis (SFA), iterative target transformation factor analysis (ITTFA), and the widely used multivariate curve resolution alternating least squares (MCR-ALS). While these methods have shown great potential in separating co-eluted signals, their practical application remains limited due to the need for extensive preprocessing, manual parameter tuning, and sensitivity to initial estimates and the number of components.

[0005] Furthermore, to address the retention time offset problem, some existing tools incorporate basic retention time alignment functionality based on one-dimensional (1D) chromatographic signals. However, due to the lack of prior knowledge about the chemical composition of each peak in 1D chromatographic signals, this correlation-based approach often leads to inaccurate alignment of TICs, especially in complex mixtures.

[0006] These challenges underscore the urgent need for automated, accurate component resolution and effective retention time offset correction to ensure the robustness and reproducibility of GC-MS-based analyses. Summary of the Invention

[0007] To address the shortcomings of existing technologies, this invention provides an automated enhancement and resolution method for co-eluting peaks in GC-MS of complex plant matrix samples. This method can perform TIC peak resolution, retention time shift correction, component registration, chemometric analysis, and compound identification. The TIC peak resolution module utilizes a Dynamically Eliminated Alternating Equal-Least Squares (DEMCR-ALS) algorithm. This algorithm employs multiple initialization strategies to improve the separation of co-eluting peaks and reduce reliance on the initial estimates inherent in traditional methods.

[0008] To achieve the above objectives, this invention provides an automated enhancement and analysis method for GC-MS co-elution peaks in complex plant matrix samples, comprising the following steps: S1. GC-MS analysis was performed on complex plant matrix samples containing multiple chemical components to obtain data, and TIC plots were generated based on the data; S2. Based on the TIC plot of each peak, a Gaussian smoothing peak detection algorithm is used to determine the peak resolution range of a single TIC peak; S3. Based on the determined peak resolution range, construct an initial chromatogram; S4. Based on the initial chromatogram, the dynamic elimination multivariate curve resolution alternating partial least squares algorithm is used to automatically perform peak resolution for each TIC peak to obtain the mass spectrum of each component; S5. Based on the obtained mass spectra of each component, automatically export them to an MSP file, and then import them into the National Institute of Standards and Technology (NIST) library for component identification.

[0009] Preferably, in step S2, the determination of the peak resolution range of a single TIC peak using a Gaussian smoothing peak detection algorithm specifically involves: A Gaussian smoothing-based peak detection method is applied to detect TIC peaks. The following equations are used to determine the starting boundary BL and ending boundary BR of the peak resolution range for each TIC peak: BL = pce - wL * rL, BR = pce + wR * rR. Here, pce represents the vertex position of the TIC peak; wL and wR represent the widths of the left and right portions of the TIC peak, respectively, calculated as follows: wL = pce - pbe and wR = ped - pce, where pbe and ped are the start and end positions of the TIC peak value; rL and rR represent the amplification factors of the left and right portions of the TIC peak, respectively, calculated as follows: rL = hce / (hce - hbe) and rR = hce / (hce - hed), where hbe, hed, and hce represent the responses at the start, end, and vertex of the TIC peak, respectively.

[0010] Preferably, the starting boundary BL and the ending boundary BR are adjusted by moving five data points forward and backward respectively: BL=BL-5 and BR=BR+5; if the TIC mass spectrum has a correlation coefficient ≥0.95 with the nearest TIC mass spectrum on the left, then BL=pbe; if the TIC mass spectrum has a correlation coefficient ≥0.95 with the nearest TIC mass spectrum on the right, then BR=ped.

[0011] Preferably, the method for constructing the initial chromatogram in step S3 is as follows: using four strategies to construct the initial chromatogram: EIC clustering, independent component analysis (ICA), iterative target transformation factor analysis (ITTFA), and model peaks.

[0012] Preferably, the dynamic elimination multivariate curve-resolved alternating partial least squares algorithm is used to solve for the C and S matrices through an iterative optimization process; the algorithm is based on the bilinear structure of GC-MS data and can be described as: X = CST + E; where X represents the matrix of all detected EIC peaks included in the resolution range of a TIC peak, C represents the matrix of the chromatogram, S represents the matrix of the mass spectrum, and E represents the instrument noise.

[0013] Preferably, the iterative optimization process of the dynamic elimination of multivariate curve-resolved alternating partial least squares algorithm includes the following steps: Step 1: Perform ill-condition analysis on the chromatographic matrix C to delete chromatograms with a condition number < 30, and generate a new C chromatographic matrix; Step 2: Using the chromatographic matrix C generated in Step 1, calculate the mass spectrometry matrix S: ST = C-1X; use the Pearson correlation coefficient method to calculate the correlation coefficient RS between mass spectra, eliminate mass spectra with RS greater than 0.90 to obtain a new S matrix; where X represents the matrix of all detected EIC peaks included in the resolution range of a TIC peak, C represents the chromatographic matrix, S represents the mass spectra matrix, and T represents the retention time characteristic matrix; Step 3: Update the chromatographic matrix C using the mass spectrometry matrix S, with the formula: Cnew=X(ST)-1; where Cnew represents the non-negative least squares function calculation with a fixed chromatographic matrix, X represents the matrix of all detected EIC peaks within the resolution range of a TIC peak, S represents the mass spectrometry matrix, and T represents the retention time feature matrix. Step 4: Using the updated chromatographic matrix Cnew and mass spectrometry matrix S, calculate the new X matrix: Xnew = CnewST; use the Pearson correlation coefficient calculation method to calculate the correlation coefficient RE between X and Xnew and the residual v to evaluate the convergence criterion; if the minimum value of the correlation coefficient RE is greater than 0.99 and the residual v < 1 × 10-6, then the iterative optimization has converged. Step 5: Repeat steps 1-4 until convergence is achieved or the number of iterations exceeds 300.

[0014] Compared with the prior art, the present invention has the following advantages and technical effects: This invention provides an automated enhanced resolution method for co-eluting peaks in GC-MS of complex plant matrix samples. This method develops a novel Dynamic Elimination Multivariate Curve-Resolved Alternating Partial Least Squares (DEMCR-ALS) algorithm. First, a Gaussian smoothing-based peak detection method is used to determine the peak resolution range of individual TIC peaks. Then, four strategies are developed based on EIC clustering, ICA, ITTFA, and model peaks to construct initial chromatograms. Finally, the DEMCR-ALS algorithm is used to achieve component separation, thereby extracting pure chromatograms and mass spectrometers. This method employs multiple initialization strategies to improve the resolution of co-eluting peaks and reduces dependence on the inherent initial estimates in traditional methods. In particular, the model peak-based strategy demonstrates excellent performance in both qualitative and quantitative evaluation. Attached Figure Description

[0015] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an undue limitation of this application.

[0016] Figure 1 The diagrams show a comparison between the TIC peak analysis workflow of AntDAS-CPR (a data analysis tool) and traditional methods. Figure a shows the workflow of the traditional method, and Figure b shows the workflow of AntDAS-CPR. Figure 2The diagrams show the peak resolution of TIC (Total Ion Chromatography) in AntDAS-CPR. A1-A3 represent the peak resolution range determination method, with A1 being the TIC peak extraction diagram and A2 representing the resolution range of TIC peak #83. A3 is the EIC (Extracted Ion Chromatography) diagram under TIC peak #83. B1-B3 represent the initial chromatogram construction method, with B1 being the clustered EIC diagram, B2 being the chromatogram constructed based on the clustered EIC, and B3 being the initial chromatogram. C1-C3 represent the component resolution method using DEMCR-ALS (Dynamic Elimination Multivariate Curve Resolution Alternating Partial Least Squares Algorithm), with C1 being the dynamic elimination curve diagram, B2 being the resolved chromatogram, and C3 being the retrieved component diagram. D1-D2 represent compound identification diagrams, with D1 being the compound identification diagram for component #1 and D2 being the compound identification diagram for component #2. Here, intensity represents signal strength, and Number of components represents the number of components.

[0017] Figure 3 Examples of TIC peak resolution performed by AntDAS-CPR under different resolution scenarios are shown. Figure a shows the TIC chromatogram of a complex plant matrix sample, with approximately 273 peaks detected. The inset shows a magnified area, highlighting overlapping peaks (e.g., 81-84, 118-120, 138-142, 157-160). Figure b shows six typical resolution cases for selected TIC peaks. Each case (from top to bottom) includes: the TIC peak, the EIC spectrum within the determined resolution interval, and the chromatogram of the resolved components (labeled as 1# and 2#). Figure c shows the mass spectra of the resolved components 1# and 2#. The spectral similarity is indicated above each set of figures.

[0018] Figure 4 This is a comparison of AntDAS-CPR with advanced GC-MS data analysis tools on Dataset 1, where the data analysis tools are AMDIS, ADAP-GC, MS-DIAL, and eRah. Dataset 1 consists of data obtained by GC-MS analysis after adding 22 standards to a chrysanthemum matrix. Figure a shows the number of target compounds detected at seven concentration levels (40-160 ng / mL), Figure b shows the molecular factor distribution of the detected target compounds, and Figure c shows the coefficient of determination R calibrated at different concentration levels. 2 Evaluate quantitative performance.

[0019] Figure 5This section compares the non-targeting identification performance of four GC-MS data analysis tools—AntDAS-CPR, ADAP-GC, MS-DIAL, and eRah—on Dataset 2 (specifically, GC-MS data of volatile organic compounds from processed flaxseed collected under different capture methods). Figure a shows the batch analysis output of each tool: total detected components (cyan), library matches satisfying a matching factor (MS) > 700 and a retention index (RI) tolerance < 30 (orange), and unique identification results confirmed manually after removing duplicates with similar retention times (red). Figure b shows the intersection analysis of compounds identified by the tools. Bar charts show the number of compounds identified by single, dual, triple, and quadruple methods; the junction matrix (bottom) indicates the specific intersections of AntDAS-CPR, ADAP-GC, MS-DIAL, and eRah. The largest intersection corresponds to compounds detected by all four methods (n=35).

[0020] Figure 6 Sample discrimination and Monte Carlo evaluation of PLS-DA models constructed for the differential components of AntDAS-CPR, ADAP-GC, MS-DIAL and eRah on dataset 2, where Figure a is the result of AntDAS-CPR, Figure b is the result of ADAP-GC, Figure c is the result of MS-DIAL and Figure d is the result of eRah. Detailed Implementation

[0021] It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and embodiments.

[0022] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0023] Example 1 The data required for this invention were obtained based on GC-MS.

[0024] Chrysanthemum and flaxseed were used as the substrates. Chrysanthemum was supplemented with 22 organic compounds (as shown in Table 1) as the sample. Flaxseed was used as the sample by three different VOCs capture methods during the roasting process. GC-MS analysis was performed using an Agilent 7890B-5977A GC-MS instrument.

[0025] This embodiment specifically includes the following steps: (1) Preparation of mixed standards: The stock solutions of the 22 organic compounds listed in Table 1 were mixed to obtain a stock solution containing 2 mg / mL of each compound. A series of standard compound solutions with concentrations ranging from 400 to 1600 ng / mL and increments of 200 ng / mL were prepared by diluting the stock solutions with CH2Cl2. Each concentration was repeated 3 times to obtain a total of 7 working solutions of organic compounds with different concentrations. The above series of standard compound solutions with different concentrations were diluted with chrysanthemum plant extract (the preparation process is described in the experiment below) at a volume ratio of 1:9 to obtain additive mass concentration levels ranging from 40 ng / mL to 160 ng / mL, with increments of 20 ng / mL.

[0026] Table 1: 22 organic compounds added to the chrysanthemum substrate

[0027] (2) Preparation of Chrysanthemum Plant Extract: Approximately 20 mg of chrysanthemum sample powder from Henan, Hubei, and Jiangsu provinces were weighed and mixed uniformly to prepare the chrysanthemum QC sample. Approximately 20 mg of the chrysanthemum QC sample was weighed into a centrifuge tube, and then 1.5 ml of an extraction solution consisting of methanol and chloroform (3:2, v / v) was added. The centrifuge tube was vortexed for 2 min, sonicated at room temperature for 30 min, and then centrifuged at 13000 r / min for 10 min. 200 mL of the supernatant solution was transferred to a chromatographic vial and dried in a metal bath under carbon dioxide protection at 50 °C. The dried sample was redissolved in 100 mL of methoxyamine pyridine hydrochloride solution (20 mg / mL). After vortexing the redissolved solution for 2 min, it was placed in a metal bath at 70 °C for 100 min, and then 100 mL of bis(trimethylsilyl)trifluoroacetamide (BSTFA) was added. Subsequently, the chromatographic vial was incubated at 70 °C for 100 min. Transfer the solution to a 1.5 ml centrifuge tube and centrifuge at 13000 r / min for 5 min. Transfer 100 mL of the supernatant to another chromatographic vial and incubate at 20 °C for 6 h.

[0028] According to the dilution method in (1), the chrysanthemum plant extract prepared above was diluted to obtain 7 different concentrations. Each concentration was repeated 3 times, resulting in a total of 21 mixed standards.

[0029] (3) Agilent 7890B-5977A GC-MS analysis of chrysanthemum samples: The chromatographic conditions for analyzing chrysanthemum samples using an Agilent 7890B-5977A GC-MS were as follows: a DB-5MS column (50 m × 0.25 mm, 0.25 mm). Helium (99.999%) was used as the carrier gas, with a split ratio of 30:1. The injection port temperature was set to 280℃. The temperature program was as follows: initial temperature 70℃, held for 3 min; then increased to 150℃ and 300℃ at constant rates of 5℃ / min and 3℃ / min respectively, maintaining the column temperature at 300℃ for 10 min; finally, the temperature was increased to 310℃ and run for 10 min.

[0030] The mass spectrometry conditions for analyzing chrysanthemum samples using an Agilent 7890B-5977A GC-MS were as follows: transfer line temperature set to 280℃. EI source parameters were: ion source temperature maintained at 230℃, collision energy set to 70 eV. The mass spectrometer acquisition range was 50-550 Da, with an optimal scan rate of 5 spectra / s. The solvent delay was 10.6 min. All other conditions were set to default. The final gas-mass spectrum obtained from the GC-MS analysis was then obtained.

[0031] (4) VOCs capture process during flaxseed roasting: Approximately 800 g of fresh flaxseed was weighed, crushed, and transferred to a preheated roaster. The roasting temperature was 170℃, and the roasting time was 30 min. VOCs released during roasting were guided by a circulating water vacuum pump at a flow rate of 0.2 m-3 / h. A Carboxen / PDMS SPME (Supelco, USA) was selected to capture VOCs by absorbing the VOCs in the vials. After roasting, the roasted flaxseed was removed from the roaster and divided into two parts for cooling: one part was cooled naturally, and the other part was cooled with liquid nitrogen. Approximately 30 mg of the naturally cooled and liquid nitrogen-cooled flaxseed were weighed into headspace vials. The SPME was placed in the headspace vials and the released VOCs were captured in an oven at 90℃. The capture time was set to 10 min. Each VOCs capture method was repeated 8 times, resulting in 24 samples.

[0032] (5) Agilent 7890B-5977A GC-MS analysis of flaxseed samples The chromatographic conditions for analyzing flaxseed samples using an Agilent 7890B-5977A GC-MS were as follows: an Agilent DB-WAXETR capillary column (60 m × 0.25 mm, 0.25 mm). Helium (99.999%) was used as the carrier gas at a flow rate of 1 mL / min. The injection port temperature was set to 250 °C. The temperature program was as follows: an initial temperature of 50 °C was maintained for 3 minutes; then the temperature was increased to 250 °C at a rate of 3 °C / min and maintained for 10 minutes; finally, the temperature was maintained at 250 °C for 15 minutes.

[0033] The mass spectrometry conditions for analyzing flaxseed samples using an Agilent 7890B-5977A GC-MS were as follows: an EI ion source with an ionization voltage of 70 eV was used. Data acquisition was performed in full scan mode, with a mass range of 50–500 Da and a scan rate of 1562 u / s. The ion source temperature was 230 °C, the quadrupole temperature was 150 °C, and the transfer line temperature was 250 °C. All other conditions were set to default. The final gas-mass spectra obtained from the GC-MS analysis were obtained.

[0034] Example 2 Its dynamic elimination of multivariate curve resolution-alternating least squares TIC peak resolution method, such as Figure 1 As shown.

[0035] Specifically, the steps include the following: S1. Determine the peak resolution range of a single TIC peak based on the data acquired by GC-MS; A Gaussian smoothing-based peak detection algorithm was applied to perform peak detection on the TIC and extracted ion chromatogram (EIC) curves, respectively. Based on the TIC peak detection, the start boundary (BL) and end boundary (BR) of the peak resolution range for each TIC peak were determined using the following formulas: BL = pce - wL * rL, BR = pce + wR * rR. Here, pce represents the vertex position of the TIC peak; wL and wR represent the widths of the left and right portions of the TIC peak, respectively, calculated as follows: wL = pce - pbe and wR = ped - pce, where pbe and ped are the start and end positions of the TIC peak, respectively; rL and rR represent the amplification factors of the left and right portions of the TIC peak, respectively, calculated as follows: rL = hce / (hce - hbe) and rR = hce / (hce - hed), where hbe, hed, and hce represent the responses at the start, end, and vertex of the TIC peak, respectively. To ensure the start and end boundaries fully cover the underlying components, the two boundaries are adjusted by shifting five data points forward and backward respectively: BL = BL-5 and BR = BR+5. Additionally, if a TIC at a given location has a high correlation coefficient (e.g., ≥0.95), and the nearest TIC mass spectrum is on the left, the start boundary remains unchanged: BL = pbe. Similarly, if the TIC mass spectrum has a high correlation coefficient with the nearest TIC mass spectrum on the right, the end boundary is determined to be: BR = ped.

[0036] S2. Based on the determined TIC peak resolution range, construct an initial chromatogram; To evaluate the impact of different initial values ​​on peak resolution performance, four strategies were developed based on EIC clustering, ICA, ITTFA, and model peaks for constructing initial chromatograms. The EIC clustering strategy is used as an example to illustrate the workflow below.

[0037] EIC clustering is based on the similarity of EIC profiles. First, a correlation coefficient matrix is ​​constructed using all EICs within a defined peak resolution range. Second, clusters are formed based on elution time and correlation coefficient thresholds: two EIC peaks are grouped if the elution time difference is <0.015 min and their correlation coefficient is >0.90. For each cluster, singular value decomposition (SVD) is used to generate chromatograms based on the bilinear structure of EICs derived from the same compound. Chromatograms with severe collinearity or poor peak shape are then removed using filters, and the remaining chromatograms are normalized to produce an initialization set.

[0038] S3. Based on the initial chromatogram, the DEMCR-ALS algorithm is used to perform peak analysis on each TIC peak to extract pure chromatograms and mass spectrometry data; The principle of this method is based on the bilinear structure of GC-MS data, which can be described as: X = CST + E. Here, X represents a matrix of all detected EIC peaks within the resolution range of a TIC peak. Each column corresponds to the elution curve of an EIC peak at a specific m / z value, and each row corresponds to the obtained mass spectrum; C represents the chromatogram matrix; S represents the mass spectrum matrix; and E represents instrument noise. The goal of DEMCR-ALS is to solve for the C and S matrices through an iterative optimization process. The iterative optimization process includes the following steps: Step 1: Perform ill-condition analysis on the chromatographic matrix C to delete chromatograms with a condition number < 30, and generate a new C chromatographic matrix; Step 2: Calculate the mass spectrometry matrix S. Using the chromatographic matrix C generated in Step 1, calculate the mass spectrometry matrix S: ST = C⁻¹X. Calculate the correlation coefficient RS between mass spectra using the Pearson correlation coefficient method. Eliminate mass spectra with RS greater than 0.90 to obtain the new S matrix; where X represents the matrix of all detected EIC peaks within the resolution range of a TIC peak, C represents the chromatographic matrix, and S represents the mass spectrum matrix. Step 3: Update the chromatographic matrix C. Update the chromatographic matrix C with the mass spectrometry matrix S, as shown below: Cnew = X(ST)-1. Calculate the mass spectrometry matrix S with a fixed chromatographic matrix C using the nonnegative least squares function in MATLAB. During the iteration process, nonnegativity and unimodality constraints are applied to the chromatographic matrix C to ensure physical relevance; X represents the matrix containing all detected EIC peaks within the resolution range of a TIC peak, and S represents the mass spectrometry matrix. Step 4: Convergence Test. Using the updated chromatographic matrix (Cnew) and mass spectrometry matrix (S), calculate the new X matrix: Xnew = CnewST. Evaluate the convergence criteria by calculating the correlation coefficient (RE) and residual (v) between X and Xnew. If the smallest element in the RE is > 0.99 and v < 1 × 10⁻⁶, the iterative optimization is considered converged. Step 5: Iteration process, repeat steps 1-4, until convergence is achieved or the number of iterations exceeds 300.

[0039] Example 3 The analytical data results of this invention.

[0040] Figure 1 Figure b shows a comparison between the TIC peak resolution workflow of AntDAS-CPR and traditional methods. Compared with the traditional method in Figure a, the TIC peak resolution workflow of AntDAS-CPR is more intelligent and does not require component estimation before iterative optimization.

[0041] Figure 2This is a diagram illustrating the TIC peak resolution in AntDAS-CPR. Figures a1-a3 show an example of peak resolution range determination: a1 illustrates an example of TIC peak detection within an elution time range of 20.8–21.4 min. Four TIC peaks (83, 84, 85, and 86) were detected, each meeting the peak extraction requirements. Based on this, a2 uses TIC peak 83 as an example to illustrate the determination of the peak resolution range; using the aforementioned peak resolution range determination strategy, the resolution range for peak 83 was determined to be 20.917–21.205 min. Within this determined peak resolution range, all detected EICs are shown in a3. Figures b1-b3) illustrate the initial chromatogram construction method. b1 shows the aggregation result of the 83rd TIC peak, with different colors representing different aggregations. For each cluster, singular value decomposition (SVD) is applied to generate chromatograms, utilizing the bilinear structure of EICs derived from the same compound. b2 shows eight chromatograms obtained from the aggregated EICs, then filters are used to remove chromatograms with severe collinearity or poor peak shape, and the remaining chromatograms are normalized to generate the initial set. b3 shows the final eight initial chromatograms of the 83rd TIC peak. Figures c1-c3) show component analysis using DEMCR-ALS. The initial chromatograms obtained through the EIC clustering strategy are used to illustrate the performance of DEMCR-ALS. c1 shows the number of components during the iterative optimization process. Initially, eight components were introduced into DEMCR-ALS, but after the iterative optimization process, only three components remained. c2 shows the three resolved chromatograms, and c3 shows the three retrieved components within the peak resolution range. Figures d1-d2) show the compound identification of the resolved components. The compounds of components "1#" and "2#" were identified as L-Methionine and L-5-Oxoproline, respectively.

[0042] The strategy described above was then used for practical application of TIC peak analysis. Figure 3Examples of TIC peak resolution performed for AntDAS-CPR under different resolution scenarios. Figure a) shows the TIC of a complex sample with approximately 273 detection peaks; the figure shows the expanded regions highlighting overlapping peaks (e.g., 81–84, 118–120, 138–142, 157–160). Due to the complexity of the sample, overlapping TIC peaks often indicate co-eluted compounds. Figure b) shows six representative resolution cases of selected TIC peaks. For each case (from top to bottom): TIC peak(s), EICs within the defined resolution range, and the resolution chromatograms of the retrieved components (labeled 1# and 2#). Figure b) highlights representative resolution schemes, including: (i) co-elution of two components with low or high mass spectrometric similarity; (ii) baseline separation of two components with low or high mass spectrometric similarity; and (iii) baseline separation of two components with low or high mass spectrometric similarity. Figure c) shows the corresponding retrieved mass spectra of components 1# and 2#, with spectral similarity reported at the top of each plot. This result demonstrates that AntDAS-CPR can resolve co-elution peaks under various conditions.

[0043] The performance of AntDAS-CPR was benchmarked against several advanced GC-MS data analysis tools, including AMDIS, ADAP-GC, MS-DIAL, and eRah. Figure 4 This figure compares AntDAS-CPR with advanced GC-MS data analysis tools (AMDIS, ADAP-GC, MS-DIAL, and eRah) on dataset 1, which contains data obtained by GC-MS analysis after adding 22 standards to a chrysanthemum matrix. Figure a) shows the number of target compounds detected at seven concentration levels (40–160 ng / mL). AntDAS-CPR successfully detected all 22 target compounds in the complex chrysanthemum matrix at each concentration level. AMDIS achieved complete detection at concentrations greater than 60 ng / mL. In contrast, ADAP-GC, MS-DIAL, and eRah failed to detect all targets at any concentration level, with eRah consistently detecting the fewest compounds, followed by MS-DIAL and ADAP-GC. Figure b) shows the distribution of MFs of the detected targets, further highlighting the superior performance of AntDAS-CPR. Figure c) evaluates the quantification capability by R² determined through calibration at different concentration levels. AntDAS-CPR yielded the highest median MF value, followed by AMDIS, MS-DIAL, and ADAP-GC, while eRah produced the lowest. The results indicate that AntDAS-CPR achieved a higher median R² compared to the other four tools, suggesting superior quantitative capabilities.

[0044] For dataset 2 (specifically, GC-MS data of volatile organic compounds from processed flaxseed collected under different capture methods), a total of 24 samples collected using the three VOC capture methods were processed in batch mode using each data analysis tool (AMDIS, ADAP-GC, MS-DIAL, and eRah). Each method generated a list of components for subsequent analysis. Figure 5 A comparison of the non-target identification performance of four GC-MS data analysis tools (AntDAS-CPR, ADAP-GC, MS-DIAL, and eRah) on Dataset 2 is presented. Figure a) shows the batch analysis output for each tool: total retrieved components (cyan), library matches meeting MF>700 and RI tolerance <30 (orange), and manually confirmed unique identifications after removing duplicates within similar retention times (red). For compound identification, mass spectra of each component were output in MSP format, and components meeting both MF>700 and RI tolerance <30 were selected based on a query of the NIST library. Under these criteria, ADAP-GC retrieved the most components, followed by eRah and AntDAS-CPR, with MS-DIAL returning the fewest. To further compare the methods, the identifier intersection between the different tools was examined. Figure b) shows the cross-analysis of identified compounds from the four different tools (AntDAS-CPR, ADAP-GC, MS-DIAL, and eRah). The bars represent the counts of compounds identified by one, two, three, or four methods; the connecting dots (bottom) represent specific intersections of AntDAS-CPR, ADAP-GC, MS-DIAL, and eRah. The largest intersection corresponds to compounds detected by all four methods (n=35). Considering intersections involving three of the four tools, combinations including AntDAS-CPR consistently produce higher counts. eRah ranks first in terms of uniquely identified compounds, followed by AntDAS-CPR, ADAP-GC, and MS-DIAL. These findings suggest that AntDAS-CPR outperforms the other three tools in identification.

[0045] In addition, cluster analysis was performed on the samples. (See attached image) Figure 6Sample discrimination and Monte Carlo evaluation were performed on PLS-DA models constructed for AntDAS-CPR, ADAP-GC, MS-DIAL, and eRah differential components on Dataset 2. Components with p-values ​​<0.01 (ANOVA) and fold differences >3 were selected for PLS-DA. The PLS-DA models for each method were further evaluated using the Monte Carlo method. In this case, 80% of the samples from each acquisition mode were selected as the training set to build the discriminative model, and the remaining 20% ​​of the samples were used for testing, followed by 10,000 random simulations. Figure (ad) shows the PLS-DA plot with 95% confidence ellipses in online (cyan) and two offline acquisition modes (orange, red), and a bar chart summarizing the model validation and prediction accuracy of 10,000 random simulations. AntDAS-CPR performed best in prediction, with the highest validation and prediction accuracies of 99.8% and 87.9%, respectively.

[0046] In summary, these results demonstrate that AntDAS-CPR offers better compound identification capabilities and better preserves sample-specific clustering patterns compared to the other three methods.

[0047] The above description is merely a preferred embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An automated enhancement and analysis method for GC-MS co-elution peaks in complex plant matrix samples, characterized in that, Includes the following steps: S1. GC-MS analysis was performed on complex plant matrix samples containing multiple chemical components to obtain data, and TIC plots were generated based on the data; S2. Based on the TIC plot of each peak, a Gaussian smoothing peak detection algorithm is used to determine the peak resolution range of a single TIC peak; S3. Based on the determined peak resolution range, construct an initial chromatogram; S4. Based on the initial chromatogram, the dynamic elimination multivariate curve resolution alternating partial least squares algorithm is used to automatically perform peak resolution for each TIC peak to obtain the mass spectrum of each component; S5. Based on the obtained mass spectra of each component, automatically export them to an MSP file, and then import them into the National Institute of Standards and Technology (NIST) library for component identification.

2. The automatic enhanced parsing method according to claim 1, characterized in that, In step S2, determining the peak resolution range of a single TIC peak using a Gaussian smoothing peak detection algorithm specifically involves: applying a Gaussian smoothing-based peak detection method to detect TIC peaks, and using the following equation to determine the starting boundary B of the peak resolution range for each TIC peak. L and the ending boundary B R B L =p ce -w L *r L B R =p ce +w R *r R Among them, p ce Indicates the position of the peak of the TIC peak; w L and w R Let w represent the widths of the left and right portions of the TIC peak, respectively, calculated as follows: L =p ce -p be and w R =p ed -p ce , where p be and p ed These are the start and end positions of the TIC peak, respectively; r L and r R The amplification factors for the left and right portions of the TIC peak are represented respectively, and are calculated as follows: r L =h ce / (h ce -h be ) and r R =h ce / (h ce -h ed ), where h be h ed and h ce These represent the responses at the start, end, and peak of the TIC peak, respectively.

3. The automatic enhanced parsing method according to claim 2, characterized in that, The starting boundary B L and the ending boundary B R Adjust the boundary by moving five data points forward and backward respectively: B L =B L -5 and B R =B R +5; If the TIC mass spectrum has a correlation coefficient ≥0.95 with the nearest TIC mass spectrum on the left, then BL=p be If the TIC mass spectrum has a correlation coefficient ≥0.95 with the nearest TIC mass spectrum on the right, then BR=p ed .

4. The automatic enhanced parsing method according to claim 1, characterized in that, The method for constructing the initial chromatogram in step S3 is as follows: the initial chromatogram is constructed using four strategies: EIC clustering, independent component analysis, iterative target transformation factor analysis, and model peak.

5. The automatic enhanced parsing method according to claim 1, characterized in that, The dynamic elimination of multivariate curve-resolved alternating partial least squares algorithm solves for the C and S matrices through an iterative optimization process; the algorithm is based on the bilinear structure of GC-MS data, and can be described as: X = CS T +E; where X represents the matrix of all detected EIC peaks within the resolution range of a TIC peak, C represents the matrix of chromatograms, S represents the matrix of mass spectra, and E represents instrument noise.

6. The automatic enhanced parsing method according to claim 5, characterized in that, The iterative optimization process of the dynamic elimination multivariate curve-resolved alternating partial least squares algorithm includes the following steps: Step 1: Perform ill-condition analysis on the chromatographic matrix C to delete chromatograms with a condition number < 30, and generate a new C chromatographic matrix; Step 2: Using the chromatographic matrix C generated in Step 1, calculate the mass spectrometry matrix S: S T =C -1 X; The correlation coefficient R between mass spectra was calculated using the Pearson correlation coefficient method. S Eliminate R S A new S matrix is ​​obtained by mass spectrometry with a resolution greater than 0.90; where X represents the matrix of all detected EIC peaks included in the resolution range of a TIC peak, C represents the matrix of the chromatogram, and S represents the matrix of the mass spectrum. Step 3: Update the chromatographic matrix C using the mass spectrometry matrix S, with the formula: C new =X(S T ) -1 ; where C new The non-negative least squares function calculation has a fixed chromatographic matrix, where X represents the matrix of all detected EIC peaks included within the resolution range of a TIC peak, and S represents the mass spectrometry matrix. Step 4: Use the updated chromatographic matrix C new Given the mass spectrum matrix S, calculate the new X matrix: X new =C new S T X and X' were calculated using the Pearson correlation coefficient method. new The convergence criterion is evaluated using the correlation coefficient RE and the residual v; the minimum value of the correlation coefficient RE is greater than 0.99 and the residual v is less than 1 × 10⁻⁶. -6 If so, then the iterative optimization has converged; Step 5: Repeat steps 1-4 until convergence is achieved or the number of iterations exceeds 300.