An automated quantitative analysis method targeting liquid-mass metabolomics data
Patent Information
- Application Number
- CN202310617207.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-29
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2043-05-29
AI Technical Summary
但整体来说,这些方法存在自动化程度低,费时、耗力,且受实际数据不同色谱峰形等影响较大,定量分析结果往往并不理想等多个方面的劣势
[0005] To address the aforementioned problems, the technical solution adopted in this invention is: an automated quantitative analysis method for targeted liquid-mass metabolomics data, comprising the following steps:
Smart Images

Figure CN116642989B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of analytical chemistry and relates to an automated quantitative analysis method for targeted liquid-mass metabolomics data. Background Technology
[0002] Peak identification, peak extraction, and peak resolution are important steps in metabolomics studies based on liquid chromatography-mass analysis. [1] Highly complex data contains tens of thousands, or even more, of metabolite ion signatures. It also inevitably includes metabolites with relatively low concentrations, as well as inaccurate or erroneous component measurements and co-elution, significantly increasing the difficulty of obtaining accurate qualitative and quantitative information on metabolites. However, inaccurate metabolite analysis or biomarker discovery results can reduce the acceptability of subsequent analytical results and even lead to erroneous conclusions. The high resolution and sensitivity of liquid chromatography-mass spectrometry (LC-MS) complicate peak shapes and overlap, making it difficult to obtain uncontaminated chromatographic and mass spectrometric information in metabolomics analysis. While software provided by instrument manufacturers is widely used for quantitative analysis, the time-consuming and tedious analytical procedures, including manual peak area correction, are a significant headache when processing hundreds of samples, each containing thousands, or even more compounds. In our experience, it often takes a month or even longer to validate the quantitative analysis results of hundreds of samples in metabolomics analysis and their accuracy.
[0003] In the field of chemometrics, multivariate curve resolution (MCR) mathematically extends the qualitative and quantitative analytical capabilities of coupled instruments based on Lambert-Beer's law. [2,3] To date, numerous algorithms have been proposed to simultaneously resolve and obtain chromatographic and spectroscopic analysis results of pure components. These can be broadly categorized into two types: one type utilizes chromatographic evolutionary elution characteristics to attempt to discover unique analytical results for overlapping systems, such as Evolutionary Factor Analysis (EFA), Heuristic Evolutionary Feature Projection (HELP), Sub-Window Factor Analysis (SFA), and Alternating Moving Window Factor Analysis (AMWFA). [2] Another type is to find reasonable solutions through iterative optimization, including Alternating Least Squares (ALS), Iterative Target Transformation Factor Analysis (ITTFA), Simple Interactive Self-Mode Mixture Analysis (SIMPLISMA)[3], etc. The former utilizes the characteristic elution regions of a single component or multiple components, such as "selective regions" and "zero concentration regions", while the latter uses fuzzy constraints, such as the nonnegativity and unimodality of chromatographic peak shapes, etc. [4,5] .
[0004] The "mathematical separation" strategy proposed through these methods [2]While MCR analysis can simultaneously obtain chromatographic and spectral profiles in complex component systems, it struggles to meet the analytical requirements of UPLC / TQMS data. This is primarily because this type of data tends to deviate from the resolving conditions of the MCR method, namely the singlet nature of chromatographic peaks and the uniqueness of mass spectrometry for the same compound. In particular, measuring some mass spectrometric peaks may deviate from the chromatographic separation model, producing erroneous peak quantification results. Therefore, considering the specific properties of UPLC / TQMS data, novel deconvolution methods for overlapping systems are needed. For example, by generating C / Br / Cl / S isotope patterns and using machine learning classifiers, a deconvolution strategy for HRMS isotope profiles has been developed to reduce false positive results. To quantify isomers without chromatographic separation, linear equation deconvolution analysis (LEDA) has been used to define the mixing matrix, and overdetermined linear equations are employed to perform deconvolution analysis on isomers. In ion mobility-based mass spectrometry, stoichiometric deconvolution is used to address the quantitative analysis of overlapping species, allowing for the calculation of precise collision cross-section (CCS) values for pure component ions, including isomers. [6] Quantitative analysis of asymmetric chromatographic peaks with low signal-to-noise ratios also presents significant challenges, which can be addressed through estimation processes using dual Gaussian mixture models and mixed processing. Furthermore, many platforms are widely used in both targeted and non-targeted metabolomics analyses, simplifying the processing of UPLC / TQMS data, including data preprocessing and deconvolution. These platforms include MS-DIAL, DecoMetDIA, and MZmine2. However, overall, these methods suffer from several disadvantages, including low automation, time-consuming and labor-intensive processes, significant susceptibility to variations in peak shapes in actual data, and often unsatisfactory quantitative analysis results. Summary of the Invention
[0005] To address the aforementioned problems, the technical solution adopted in this invention is: an automated quantitative analysis method for targeted liquid-mass metabolomics data, comprising the following steps:
[0006] S1: First, read and parse the targeted liquid-mass metabolomics analysis data. Based on the characteristics of the targeted liquid-mass analysis data, extract the ion chromatograms of the target ion pairs and perform data preprocessing on the extracted chromatograms, including using the adaptive iterative reweighted penalized least squares method to remove the data background and estimate its noise level, and then perform deconvolution quantitative analysis.
[0007] S2: Based on the data segments divided according to the noise level, the second derivative method is used for each data segment to identify the chromatographic peaks to be analyzed separately. The initial peak parameters of each component in the chromatographic peak family are estimated based on the chromatographic peak model. Similarly, multiple components contained in the flow chromatogram are gradually stripped away.
[0008] S3: Based on the initial chromatographic peak parameters of each component obtained above, the entire data is input into the fitting program for one-time fitting optimization and chromatographic peak deconvolution analysis. The fitting analysis results of the corresponding components for each chromatographic peak are then obtained as the final analysis results. The peak area integral results of the corresponding chromatographic peaks are calculated to achieve quantitative analysis.
[0009] Furthermore: the reading and parsing of the targeted liquid-mass metabolomics analysis data file includes the following steps:
[0010] c. Various attributes and parameters for file reading and parsing are included in the control dictionary, making it convenient and automated to add attributes and parameters;
[0011] d. Following the fixed format of the mzML file, parse and extract the retention time, first-order mass spectrometry, and peak intensity information required for data processing.
[0012] Furthermore, it also includes: the chromatographic peak pretreatment and multi-component stripping analysis method, comprising the following steps:
[0013] For the data corresponding to a certain ion pair that has been intercepted, the data is divided based on the noise level to obtain several data segments with noise levels greater than the noise level, and each data segment is processed separately.
[0014] For data points in a data segment that exceed the noise level, a threshold of 4 is set for data points that are greater than or equal to the preset noise level. If such data points are found to be of a certain value, they are considered to be of a certain value and will not be analyzed by deconvolution.
[0015] For the chromatographic peak region to be analyzed, point by point is compared from the left and right sides of the peak apex until the value of the next point is found to be greater than the value of the previous point.
[0016] For the points found above, ensure that the number of data points meets the requirements of chromatographic peak fitting analysis. Set a noise level threshold of 5 for the data points. For data points that are greater than or equal to the preset data point threshold of 5, perform subsequent deconvolution analysis. If the data point is less than the data point threshold of 5, skip the last point and continue to compare point by point until the value of the next point is found to be greater than the value of the previous point.
[0017] For the chromatographic peak family data obtained above for deconvolution analysis, calculate its second derivative to determine whether there is a co-elution region. If so, determine the pure chromatographic elution region based on its quantity.
[0018] Based on the above chromatographic peak elution regions, the chromatographic peak fitting method, namely the different methods contained in the lmfit function in Python, is used to gradually peel off and calculate the chromatographic peak parameters;
[0019] The fitting function is used to obtain the parameter results of the chromatographic peak, the peak area is calculated, and then the quantitative analysis of the chromatographic peak is achieved.
[0020] Furthermore, the multi-component deconvolution stripping analysis method for chromatographic peaks also includes:
[0021] By performing preliminary fitting analysis on each chromatographic peak, the preliminary analytical parameters of all chromatographic peaks are obtained. Then, these preliminary parameters are treated as a whole and input into the fitting function at once to obtain the results of simultaneous fitting analysis on each chromatographic peak, thereby improving the fitting accuracy of the chromatographic peak model.
[0022] Furthermore: the chromatographic peak deconvolution stripping analysis also includes
[0023] By comparing different chromatographic peak fitting functions and using the minimum sum of squared residuals as an indicator, the method with the best chromatographic peak fitting effect is selected, and the chromatographic peak area is calculated based on the fitting result to achieve quantitative analysis.
[0024] Based on multi-model optimization of chromatographic peaks, stepwise component stripping, and overall optimization based on initial parameters of chromatographic peaks, the peak area integral results of each component in the chromatographic peak family are obtained. The peak table is updated based on the quantitative analysis results, and then subsequent differential metabolite discovery analysis is carried out in metabolomics analysis.
[0025] Based on the data structure and characteristics of UPLC / TQMS, starting from the raw instrument data, this method first extracts the ion chromatography (EIC) data of the target measurement mass spectrometry from the data obtained from liquid chromatography-mass spectrometry. On the basis of data loading, transformation and extraction, the adaptive iterative reweighted penalized least squares (airPLS) method is used to remove the background data of chromatographic peak families. Then, the noise level is estimated, chromatographic peaks are extracted, and a strategy of component fitting, stepwise stripping and overall optimization is adopted to achieve peak clustering for deconvolution analysis. Based on the analysis results, the ability to analyze highly complex liquid chromatography-mass spectrometry data is improved, which can become an important quantitative analysis solution in metabolomics analysis.
[0026] The method proposed in this invention first reads a folder containing one or more raw measurement files from instruments. Various attributes and parameters of the files are included in a "control dictionary," facilitating and automating the addition and removal of attributes and parameters. Following the established format of the mzML file, information such as retention time (tR), first-order mass spectrometry (m / z), and peak intensity are extracted. Based on the characteristics of the data files used in targeted liquid chromatography-mass spectrometry, ion chromatograms of the target ion pairs are extracted.
[0027] For the data (EIC) corresponding to a specific ion pair, the data is divided based on the noise level, resulting in several data segments with values exceeding the noise level. Each data segment is processed individually. For data points in a segment that exceed the noise level, a preset threshold (e.g., 4) is set to indicate the presence of a chromatographic peak to be analyzed within that segment. For each chromatographic peak region, point-by-point comparisons are performed from both sides of the peak apex until a subsequent point's value is found to be greater than the preceding point. For each found point, the number of data points is ensured to meet the requirements of the fitting analysis. A preset threshold (e.g., 5) is used; if the value is less than this threshold, the last point is skipped, and point-by-point comparisons continue until a subsequent point's value is found to be greater than the preceding point. For the obtained data, the second derivative is calculated to determine if a co-elution region exists. If so, the value is used to determine a pure chromatographic elution region. Based on the data segments divided according to the noise level, the second derivative method is used for each segment to individually identify the chromatographic peaks to be analyzed. Component parameters are estimated based on the chromatographic model, and eluting chromatographic components are gradually separated.
[0028] Based on the aforementioned chromatographic peak elution regions, chromatographic peak fitting methods, specifically different methods of the `lmfit` function in Python, are employed to progressively extract and calculate chromatographic peak parameters, fit functions, and obtain the chromatographic peak results, calculating the peak area. Using the initial chromatographic peak parameters obtained above, comprehensive fitting optimization and peak analysis are performed on the initial chromatographic peaks to obtain the fitting analysis results for each chromatographic peak, and the corresponding peak integral area results are calculated. That is, after obtaining the preliminary fitting analysis results, the obtained component parameters are then treated as a whole, and the initial data is systematically fitted to obtain the results of simultaneous fitting analysis. Different chromatographic peak fitting functions are compared, and the method and result with the best fitting effect are selected based on the minimum sum of squared residuals.
[0029] Based on the chromatographic peak analysis results obtained from the above stepwise stripping, multi-model optimization and overall optimization, the peak table obtained from peak matching is updated, and subsequent differential metabolite analysis is performed. Attached Figure Description
[0030] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0031] Figure 1 This is a flowchart of the automated quantitative analysis method of this application;
[0032] Figure 2 This is a flowchart of peak deconvolution analysis in liquid chromatography-mass spectrometry.
[0033] Figure 3 This is a schematic diagram illustrating the principle of peak deconvolution analysis in liquid chromatography.
[0034] Figure 4 The figure shows the deconvolution analysis results for a typical real-world overlapping system. Detailed Implementation
[0035] It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the present invention or its application or use. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0037] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0038] Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps described in these embodiments do not limit the scope of the invention. It should also be understood that, for ease of description, the dimensions of the various parts shown in the drawings are not drawn to actual scale. Techniques, methods, and devices known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and devices should be considered part of the specification. In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values. It should be noted that similar reference numerals and letters in the following figures denote similar items; therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0039] In the description of this invention, it should be understood that the orientation or positional relationship indicated by directional terms such as "front, back, up, down, left, right", "horizontal, vertical, horizontal" and "top, bottom" is generally based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing this invention and simplifying the description. Unless otherwise stated, these directional terms do not indicate or imply that the device or element referred to must have a specific orientation or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on the scope of protection of this invention. The directional terms "inner" and "outer" refer to the inner and outer contours relative to the outline of each component itself.
[0040] Data processing in targeted liquid chromatography-mass spectrometry (LC-MS) metabolomics analysis, especially the automated extraction and quantitative analysis of overlapping chromatographic peaks under interference from complex noise, background, and spurious peaks, is crucial for improving efficiency and accuracy. It is also one of the most lacking functions in the software of currently available commercial instruments. Based on the analysis and improvement of UPLC / TQMS peak shape characteristics, the raw LC-MS data is converted to mzXML format, from which single-component or unresolved multi-component extracted ion chromatography (EIC) is extracted, corresponding to a specific or series of ion characteristics. By identifying the changes in chromatographic peak elution regions and the distribution trends of second derivatives, a chromatographic peak model optimization method is used for fitting analysis to determine the initial parameters for deconvolution of overlapping chromatographic peaks. A strategy for stepwise deconvolution to remove overlapping components is constructed to achieve automated analysis. Finally, using the obtained model parameters as the overall initial input, multi-component simultaneous analysis and modeling are performed to further optimize the deconvolution results. Based on the analysis results of complex overlapping chromatographic peaks, the peak tables obtained in targeted and pseudo-targeted metabolomics analyses are replaced, improving the accuracy and reliability of subsequent differential metabolite and biomarker discovery.
[0041] To achieve deconvolution analysis of multiple components under the aforementioned complex environments, the technical solution adopted in this invention is as follows: starting from the original data, the intelligence and convenience of the deconvolution analysis process are improved; after obtaining the data of the multi-component overlapping system, noise reduction and background removal are performed to improve data quality; chromatographic peak shape analysis and second derivative analysis are used to obtain the elution regions of the components; chromatographic peak model fitting, stepwise stripping analysis, and global optimization fitting are used to obtain the chromatographic quantitative analysis results of each component in the multi-component system, which can then be used for subsequent applications in metabolomics and other fields.
[0042] An automated quantitative analysis method for targeted liquid-mass metabolomics data includes the following steps:
[0043] S1: First, read and parse the targeted liquid-mass metabolomics analysis data. Based on the characteristics of the targeted liquid-mass analysis data, extract the ion chromatograms of the target ion pairs and perform data preprocessing on the extracted chromatograms, including using the adaptive iterative reweighted penalized least squares method to remove the data background and estimate its noise level, and then perform deconvolution quantitative analysis.
[0044] S2: Based on the data segments divided according to the noise level, the second derivative method is used for each data segment to identify the chromatographic peaks to be analyzed separately. The initial peak parameters of each component in the chromatographic peak family are estimated based on the chromatographic peak model. Similarly, multiple components contained in the flow chromatogram are gradually stripped away.
[0045] S3: Based on the initial chromatographic peak parameters of each component obtained above, the entire data is input into the fitting program for one-time fitting optimization and chromatographic peak deconvolution analysis. The fitting analysis results of the corresponding components for each chromatographic peak are then obtained as the final analysis results. The peak area integral results of the corresponding chromatographic peaks are calculated to achieve quantitative analysis.
[0046] Steps S1, S2, and S3 are executed sequentially.
[0047] Furthermore: the reading and parsing of the targeted liquid-mass metabolomics analysis data file includes the following steps:
[0048] a. Various attributes and parameters for file reading and parsing are included in the control dictionary, making it convenient and automated to add attributes and parameters;
[0049] b. Following the fixed format of the mzML file, parse and extract information such as retention time (tR), first-order mass spectrometry (m / z), and peak intensity required for data processing. Based on the characteristics of the data file for targeted liquid chromatography-mass spectrometry, extract the ion flow map under the target ion pair.
[0050] Furthermore, the chromatographic peak pretreatment and multi-component stripping analysis method includes the following steps:
[0051] For the EIC data corresponding to a certain ion pair that has been intercepted, the data is divided based on the noise level to obtain several data segments with noise levels greater than the noise level, and each data segment is processed separately.
[0052] For data points in a data segment that exceed the noise level, a threshold of 4 is set for data points that are greater than or equal to the preset noise level. If such data points are found to be of a certain value, they are considered to be of a certain value and will not be analyzed by deconvolution.
[0053] For the chromatographic peak region to be analyzed, point by point is compared from the left and right sides of the peak apex until the value of the next point is found to be greater than the value of the previous point.
[0054] For the points found above, ensure that the number of data points meets the requirements of chromatographic peak fitting analysis. Set a noise level threshold of 5 for the data points. For data points that are greater than or equal to the preset data point threshold of 5, perform subsequent deconvolution analysis. If the data point is less than the data point threshold of 5, skip the last point and continue to compare point by point until the value of the next point is found to be greater than the value of the previous point.
[0055] For the chromatographic peak family data obtained above for deconvolution analysis, calculate its second derivative to determine whether there is a co-elution region. If so, determine the pure chromatographic elution region based on its quantity.
[0056] Based on the above chromatographic peak elution regions, the chromatographic peak fitting method, namely the different methods contained in the lmfit function in Python, is used to gradually peel off and calculate the chromatographic peak parameters;
[0057] The fitting function is used to obtain the parameter results of the chromatographic peak, the peak area is calculated, and then the quantitative analysis of the chromatographic peak is achieved.
[0058] Furthermore, the multi-component deconvolution stripping analysis method for chromatographic peaks also includes:
[0059] By performing preliminary fitting analysis on each chromatographic peak, the preliminary analytical parameters of all chromatographic peaks are obtained. Then, these preliminary parameters are treated as a whole and input into the fitting function at once to obtain the results of simultaneous fitting analysis on each chromatographic peak, thereby improving the fitting accuracy of the chromatographic peak model.
[0060] Furthermore, the chromatographic peak deconvolution and stripping analysis also includes:
[0061] By comparing different chromatographic peak fitting functions and using the minimum sum of squared residuals as an indicator, the method with the best chromatographic peak fitting effect is selected, and the chromatographic peak area is calculated based on the fitting result to achieve quantitative analysis.
[0062] Based on multi-model optimization of chromatographic peaks, stepwise component stripping, and overall optimization based on initial parameters of chromatographic peaks, the peak area integral results of each component in the chromatographic peak family are obtained. The peak table is updated based on the quantitative analysis results, and then subsequent differential metabolite discovery analysis is carried out in metabolomics analysis.
[0063] Taking actual pseudo-targeted metabolomics data analysis as an example, this paper introduces the method proposed in this invention for deconvolution analysis of actual overlapping multi-component systems.
[0064] The specific implementation data is obtained from liver cancer targeted metabolic analysis, and the deconvolution analysis process is illustrated by way of example.
[0065] In UPLC / TQMS-based metabolomics analysis, data signal processing is crucial for identifying and further extracting metabolite features from each individual sample. The same analytical steps are performed on each EIC dataset from the original data set. The results are then summarized, and the most suitable peak fitting function is selected based on the fitted residuals. By comparing peaks across different data files, the minimum value from the residuals is selected as the optimal peak and its area corresponding to the target ion.
[0066] In deconvolution analysis, each mass spectrometric feature in the peak table is extracted stepwise to obtain the EIC of the entire ion within a fixed-size window, which serves as the input for deconvolution. The precise quantification results from deconvolution analysis are then used to replace the quantitative intensities in the peak table for subsequent processing. In this study, we implement this strategy using an example of pseudo-targeted metabolomics, which requires deconvolution of overlapping and spurious peaks to achieve precise metabolite quantification. First, peak matching is performed using methods such as XCMS to obtain the peak table, and overlapping deconvolution results are replaced in the peak table. Automatic and stepwise stripping is used for multi-model deconvolution, such as... Figure 1 As shown, the EIC extraction and parameter threshold settings after moving the m / z window are determined by understanding the data characteristics. This helps generate appropriate EIC profiles and ensures adaptive handling of deconvolution analysis with different chromatographic profiles. A sufficient number of EIC data points need to be extracted during the analysis to ensure that the target analyte can be included in the analysis.
[0067] Figure 2 This is a flowchart of peak deconvolution analysis in liquid chromatography-mass spectrometry.
[0068] Figure 3 This is a schematic diagram illustrating the principle of peak deconvolution analysis in liquid chromatography. Figure 3The deconvolution principle of overlapping peaks is presented, showing four typical cases with different degrees of overlap. First, the airPLS method is used to remove the data background. Then, the noise level of the chromatogram is determined by calculating the division frequency of different intensity regions, which should have a higher frequency than the chromatographic peak signal. If a predefined number of data points, such as threshold_numberofpoints, exist consecutively and their intensity is greater than the noise level, each peak cluster will be deconvolved independently. The maximum value of the peak with the highest intensity is then identified from the peak cluster, and the entire peak is "scanned" along both sides of the peak apex. Fitting analysis is performed using all data points with intensities lower than the previous data point. If the number of data points is insufficient for model fitting, the scanning process continues even if the above principle is not met. The Levenburg-Marquardt algorithm is used to obtain the second derivative of the data to be analyzed to estimate the peak elution region of the chromatographic peak. Next, the residuals are calculated and analyzed, and the calculation is iterated again to gradually peel away overlapping peak clusters from the EIC profile, achieving deconvolution analysis.
[0069] Figure 4 The following is a diagram showing the deconvolution analysis results for a typical real-world overlapping system, as shown below. Figure 4 As shown. Due to the complexity of UPLC / TQMS peak clusters and the local component stripping characteristics, deviations from the global optimal solution may still exist. Therefore, by utilizing the parameters obtained from all deconvolution deconvolution peaks in each EIC profile, a global multi-model comprehensive optimization is performed. That is, the parameters of each peak are used as initial inputs to perform global optimization on the entire original data, which helps to find the overall optimal result for all components.
[0070] As mentioned above, if multiple clusters exist in the EIC profile, the peak clusters will be deconvolved independently. This method has significant advantages in deconvolving UPLC / TQMS peak clusters, providing an effective quantitative analysis method for metabolomics studies.
[0071] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
[0072] 1.J.Zhang,E.Gonzalez,T.Hestilow,W.Haskins,Y.Huang,Review of PeakDetection Algorithms in Liquid-Chromatography-Mass Spectrometry,CurrentGenomics 10(6)(2009)388-401.
[0073] 2.A.de Juan,R.Tauler,Multivariate Curve Resolution:50years addressingthe mixture analysis problem-Areview,Analytica Chimica Acta 1145(2021)59-78.3.C.Ruckebusch,L.Blanchet,Multivariate curve resolution:Areview ofadvanced and tailored applications and challenges,Analytica Chimica Acta 765(2013)28-36.4.Y.Z.Liang,O.M.Kvalheim,H.R.Keller,D.L.Massart,P.Kiechle,F.Erni,Heuristic evolving latent projections:resolving two-way multicomponentdata.2.Detection and resolution of minor constituents,Analytical Chemistry 64(8)(1992)946-953.
[0074] 5.O.M.Kvalheim,Y.Z.Liang,Heuristic evolving latent projections:resolving two-way multicomponent data.1.Selectivity,latent-projective graph,datascope,local rank,and unique resolution,Analytical Chemistry 64(8)(1992)
[0075] 10.1021 / ac00032a00019.
[0076] 6.H.-M.Mayke,L.B.Bruno,M.Fabrice,G.-C.A.M.,D.-P.Gaud,Collision CrossSection(CCS)Database:An Additional Measure to Characterize Steroids,Analytical Chemistry 90(7)(2018)4616-4625.
Claims
1. An automated quantitative analysis method for targeted liquid-mass metabolomics data, characterized in that: Includes the following steps: S1: First, read and parse the targeted liquid-mass metabolomics analysis data file. Based on the characteristics of the targeted liquid-mass metabolomics analysis data, extract the ion chromatograms of the target ion pairs and perform data preprocessing on the extracted chromatograms, including using the adaptive iterative reweighted penalized least squares method to remove the data background and then estimate its noise level. S2: Based on the data segments divided according to the noise level, the second derivative method is used for each data segment to identify the chromatographic peaks to be analyzed separately. The initial chromatographic peak parameters of each component in the chromatographic peak family are estimated based on the chromatographic peak model. Similarly, multiple components contained in the chromatographic peak family are gradually stripped away. S3: Based on the initial chromatographic peak parameters of each component, the entire sample is input into the fitting program for one-time fitting optimization and chromatographic peak deconvolution analysis. The fitting analysis results of the corresponding components for each chromatographic peak are then obtained as the final analysis results. The peak area integral results of the corresponding chromatographic peaks are calculated to achieve quantitative analysis. The reading and parsing of the targeted liquid-mass metabolomics analysis data files includes the following steps: a. Various attributes and parameters for file reading and parsing are included in the control dictionary, making it convenient and automated to add attributes and parameters; b. According to the fixed format of the mzML file, parse and extract the retention time, first-order mass spectrometry and peak intensity information required for the data processing; It also includes the following steps: For the data corresponding to a certain ion pair that has been intercepted, the data is divided based on the noise level to obtain several data segments with noise levels greater than the noise level, and each data segment is processed separately. For data points in a data segment that are greater than the noise level, a threshold is set for data points that are greater than or equal to the preset noise level. If such data points are found to be chromatographic peaks, then the data segment is considered to contain chromatographic peaks to be analyzed. For data points that are less than the threshold, they are considered not to constitute chromatographic peaks and are not deconvolved for analysis. For the chromatographic peak region to be analyzed, point by point is compared from the left and right sides of the peak apex until the value of the next point is found to be greater than the value of the previous point. For the found points, ensure that the number of data points meets the requirements of chromatographic peak fitting analysis, preset the threshold of the noise level of the data points, and perform subsequent deconvolution analysis for data points that are greater than or equal to the preset threshold. If the data point value is less than the threshold, skip the last point and continue to compare point by point until the value of the next point is found to be greater than the value of the previous point. For the chromatographic peak family data obtained above for deconvolution analysis, calculate its second derivative to determine whether there is a co-elution region. If so, determine the pure chromatographic elution region based on its quantity. Based on the pure chromatographic elution region, the chromatographic peak fitting method, namely the different methods contained in the lmfit function in Python, is used to gradually peel off and calculate the chromatographic peak parameters; The fitting function is used to obtain the parameter results of the chromatographic peak, the peak area is calculated, and then the quantitative analysis of the chromatographic peak is achieved.
2. The method of claim 1, wherein the method is characterized by, It also includes the following steps: By performing preliminary fitting analysis on each chromatographic peak, the preliminary analytical parameters of all chromatographic peaks are obtained. Then, these preliminary analytical parameters are used as a whole and input into the fitting function at once to obtain the results of simultaneous fitting analysis on each chromatographic peak, thereby improving the fitting accuracy of the chromatographic peak model.
3. The automated quantitative analysis method for targeted liquid-mass metabolomics data according to claim 1, characterized in that, Its features are: It also includes the following steps: By comparing different chromatographic peak fitting functions and using the minimum sum of squared residuals as an indicator, the method with the best chromatographic peak fitting effect is selected, and the chromatographic peak area is calculated based on the fitting result to achieve quantitative analysis. Based on multi-model optimization of chromatographic peaks, stepwise component stripping, and overall optimization based on initial parameters of chromatographic peaks, the peak area integral results of each component in the chromatographic peak family are obtained. The peak table is updated based on the quantitative analysis results, and then subsequent differential metabolite discovery analysis is carried out in metabolomics analysis.