Metabonomics chromatographic peak area correction method based on model cluster analysis
By using model cluster analysis method in metabolomics, a large number of SVR submodels were established and verified, and the subdata sets final used for chromatographic peak area correction were determined, which solved the problem of inaccurate chromatographic peak area correction in metabolomics, and achieved higher correction accuracy and robustness.
Patent Information
- Application Number
- CN202510614522.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-14
AI Technical Summary
In metabolomics analysis, the accurate correction of the chromatographic peak area is affected by factors such as the early experimental conditions of the sample, noise during the detection process, and instrument performance decay, resulting in inaccurate peak area and affecting the reliability of metabolic substance content.
Using a model cluster analysis method, a large number of SVR sub-models are established to determine the sub-data sets final for SVR modeling, and a final SVR model is constructed for correction of chromatographic peak area. This method uses the QC samples from metabolomics to find the subdataset that contributes the greatest contribution to chromatographic peak area correction, ensuring the robustness of the final model.
Through the model cluster analysis method, the accuracy and robustness of chromatographic peak area correction are improved, ensuring that the peak area can accurately reflect the linear relationship between metabolic substances in different samples, and improving the comparability of the same substances in different samples.
Smart Images

Figure CN120142548A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data science and technology, and in particular to a metabolomics chromatographic peak area correction method based on model cluster analysis. Background Art
[0002] Metabolomics uses high-resolution detection instruments to detect the content of metabolites in different types of samples (such as early, middle and late stages of a disease), and explores the changing trends of metabolites in different types of samples by analyzing their changing patterns. It is an important part of systems biology and is widely used in disease diagnosis, toxicology, botany, nutrition and food science, environmental science and other fields, with very broad application prospects.
[0003] Because the basis of metabolomics analysis is the detection data of different types of samples by high-resolution detection instruments, it is affected by many factors such as the preliminary experimental conditions of the samples, the noise introduced by the samples and the experimental environment during the detection process, and the data drift caused by the performance attenuation of the experimental instruments. Therefore, it is necessary to perform peak area correction on the chromatographic peak table extracted in metabolomics to ensure that the peak area of the chromatographic peak table can accurately represent the content of the metabolites in the sample. That is, the peak area correction of the chromatographic peak table is an important part of the metabolomics analysis process and is also the basis for ensuring the success of metabolomics analysis.
[0004] Based on this, the present invention applies for a metabolomics chromatographic peak area correction method based on model cluster analysis. This method uses the QC samples of metabolomics, adopts the model cluster method to establish a large number of SVR sub-models, determines the sub-datasets ultimately used for SVR modeling, and constructs the final SVR model for chromatographic peak area correction. The present invention adopts the model cluster method to find the sub-dataset with the greatest contribution to the chromatographic peak area correction SVR model, ensures that the final SVR model constructed has better robustness, obtains the peak area correction effect that best meets the actual situation, and improves the comparability of the same substance in different samples. Summary of the invention
[0005] The purpose of this application is to provide a metabolomics chromatographic peak area correction method based on model cluster analysis.
[0006] The present application provides a metabolomics chromatographic peak area correction method based on model cluster analysis, the method comprising: Step 1: For the chromatographic peak to be detected, a SVR sub-model set is established by using a model clustering method; Step 2: Perform statistical analysis on the variables selected by the SVR sub-model set to determine the screening variable set; Step 3: constructing a final SVR model of the chromatographic peak to be detected according to the screening variable set; Step 4: Calibrate the chromatographic peak areas of the chromatographic peaks to be detected in all samples according to the SVR model; Step 5: Loop through Steps 1 - 4 until the calibration of the peak areas of all chromatographic peaks is completed.
[0007] One or more technical solutions provided in this application have at least the following technical effects or advantages: For the chromatographic peaks to be detected, establish a set of SVR sub - models in the way of model cluster; conduct statistical analysis on the variables selected by the set of SVR sub - models to determine the screening variable set; construct the final SVR model of the chromatographic peaks to be detected according to the screening variable set; calibrate the chromatographic peak areas of the chromatographic peaks to be detected in all samples according to the SVR model; loop through the foregoing steps until the calibration of the peak areas of all chromatographic peaks is completed.
[0008] First, with the help of the QC sample data with better data stability (generally speaking, the QC samples in metabolomics experiments are mixed samples of all other samples, and the QC sample data are the instrument data of the same QC sample detected at different times), perform chromatographic peak area calibration to ensure that the calibrated peak areas can reflect the linear relationship of metabolites among different samples; second, adopt the model cluster method for variable selection to avoid the problem of model over - fitting caused by the loss of important variables and the problem of model under - fitting caused by the application of unimportant variables, improve the robustness of the peak area calibration model, obtain the peak area calibration effect that most conforms to the actual situation, and enhance the comparability of the same substances in different samples.
[0009] The above description is only an overview of the technical solution of this application. In order to be able to more clearly clarify the technical means of this application, it can be implemented according to the content of the specification. And in order to make the above - mentioned and other purposes, features and advantages of this application more obvious and understandable, the following specifically illustrates the specific embodiments of this application. Brief Description of the Drawings
[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments of the present disclosure will be briefly introduced below. Flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of this application. It should be understood that the operations in the front or below do not necessarily need to be executed precisely in sequence. On the contrary, according to the need, they can be executed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or several operations can be removed from these processes.
[0011] Figure 1 It is the basic flowchart of a method for calibrating chromatographic peak areas in metabolomics based on model cluster analysis of the present invention; Figure 2Histogram of the QC RSD values of the chromatographic peak areas before chromatographic peak correction for a method for correcting chromatographic peak areas in metabolomics based on model cluster analysis according to the present invention; Figure 3 Histogram of the QC RSD values of the chromatographic peak areas after chromatographic peak correction for a method for correcting chromatographic peak areas in metabolomics based on model cluster analysis according to the present invention. Detailed implementation manner
[0012] The present application provides a method for correcting chromatographic peak areas in metabolomics based on model cluster analysis.
[0013] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0014] It should be noted that the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those clearly listed steps or units, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0015] The present application provides a method for correcting chromatographic peak areas in metabolomics based on model cluster analysis, wherein the method includes: For the chromatographic peaks to be detected, an SVR sub-model set is established in a model cluster manner; further, the step 1 includes: When the number of QC samples in the chromatographic peak table is , the number of non-QC samples is , and the number of chromatographic peaks is , the chromatographic peak areas of the QC samples in the chromatographic peak table form a matrix; For the chromatographic peak to be corrected for peak area, a data set is constructed, where is the matrix formed by the peak areas of all chromatographic peaks other than the chromatographic peak to be corrected for peak area in the QC samples, and its size is , is the matrix formed by the peak areas of the chromatographic peak in all QC samples, and its size is ; For the data set , in a non - replacement manner, randomly select variables to obtain a sub - dataset , whose size is . After extractions through Monte Carlo sampling, a total of sub - datasets are obtained, denoted as respectively. And according to datasets , , adopt the - fold cross - validation method to train, construct and validate SVR sub - models respectively, denoted as . The coefficient of determination of their verification results is denoted as respectively.
[0016] Conduct statistical analysis on the variables selected by the SVR sub - model set to determine the screened variable set; Further, step 2 includes: Based on the definitions of the dataset , sub - dataset , sub - dataset and sub - model in step 1, for the variable to be selected in , , divide the SVR sub - models into and two categories, where represents the set of sub - models in which the sub - dataset contains the variable to be selected when constructing the SVR sub - model, and represents the set of sub - models in which the sub - dataset does not contain the variable to be selected when constructing the SVR sub - model; Naturally, according to the mean values and of the coefficients of determination ( ) of the two sub - model sets and , calculate the mean difference in the calibration ability of the two sub - model sets for the chromatographic peak according to Equation (1); ; (1) If , it indicates that when selecting the th variable , the prediction ability of the SVR model will be enhanced. Therefore, the th variable Add it to the final data subset; If , it indicates that when selecting the th variable , the prediction ability of the SVR model will be weakened. Therefore, the th variable is not added to the final data subset; By setting respectively and performing the above steps respectively, the screened variable set is finally obtained.
[0017] Construct the final SVR model of the chromatographic peak to be detected according to the screened variable set; Furthermore, step 3 includes: Through the above three steps, the screened variable set for the chromatographic peak to be corrected is obtained, and finally the data set is constructed, where is the data set composed of the peak areas of the variables in in QC samples, and its size is , is the size of the screened variable set ; According to the constructed data set , use the Gaussian kernel shown in Equation (2) to train the SVR model; ; (2) Map the data set to an infinite-dimensional feature space through the Gaussian kernel, and construct the SVR model in the feature space , denoted as .
[0018] Correct the chromatographic peak areas of the chromatographic peaks to be detected in all samples according to the SVR model; Loop through steps 1 - 4 until the chromatographic peak areas of all chromatographic peaks are corrected.
[0019] Furthermore, step 4 includes: For the chromatographic peak to be corrected , where represents the peak area of the chromatographic peak in the th sample (including QC samples and non-QC samples), and the SVR model constructed for the chromatographic peak is denoted as , and training The selected variable set for is denoted as For the chromatographic peak area to be corrected , first find the selected variable set The peak areas of each chromatographic peak in the th sample are used to construct a vector , and according to the SVR model in the th sample (including QC samples and non-QC samples), the peak area of the chromatographic peak is predicted, and its calculation formula is shown in Equation (3); ; (3) Finally, according to Equation (4), the corrected peak area of the chromatographic peak in the th sample (including QC samples and non-QC samples) is obtained ; ; (4) Let , and perform Step 4 respectively to complete the correction of the chromatographic peak area of the chromatographic peak in all samples (including QC samples and non-QC samples); Let , and perform Steps 1 - 4 respectively to complete the correction of the peak areas of all chromatographic peaks in the peak table.
[0020] Specifically, 160 mouse urine samples collected from 40 mice divided into 6 groups: normal group (5 mice), model group (8 mice), positive group (6 mice), low-dose group (7 mice), medium-dose group (7 mice), and high-dose group (7 mice) at 4 stages: day 0, day 10, day 20, and day 30. Plus 16 QC samples mixed from the above 160 samples, labeled as QC1, QC2, QC3,..., QC17, a total of 160 + 16 = 176 experimental samples are obtained. The SCIEX 7600 instrument is used for the detection experiment in the positive ion ionization mode. After peak extraction by XCMS, 22,165 chromatographic peaks in the experimental data are obtained, labeled as p1, p2, p3,..., p22165.
[0021] Based on the 22,165 chromatographic peaks, for the peak areas of the chromatographic peaks in 16 QC samples, the statistical distribution of the RSD values before chromatographic peak area calibration is shown as Figure 2 follows.
[0022] RSD is the Relative Standard Deviation, and its calculation formula is shown as follows.
[0023] Among them, , represents the QC RSD value of the chromatographic peak, represents the average value of the QC RSD values of 22,165 chromatographic peaks.
[0024] Using a metabolomics chromatographic peak area correction method based on model cluster analysis proposed by the present invention, the specific implementation steps for correcting the peak areas of 22,165 chromatographic peaks in 176 samples are as follows. Without loss of generality, in the following steps, the process of correcting the peak area of chromatographic peak p1 is mainly introduced. The implementation process of correcting the peak areas of chromatographic peaks p2~p22165 is completely similar to the implemented process of correcting the peak area of chromatographic peak p1: Step 1: In the example data, the number of QC samples is 16, the number of non-QC samples is 160, and the number of chromatographic peaks is 22,165. That is, the chromatographic peak areas of the QC samples in the chromatographic peak table form a 16×22,165 matrix, and the unit value in the matrix is the chromatographic peak area in each QC sample. Without loss of generality, when correcting the peak area of chromatographic peak p1, a data set (X1, y1) is constructed, where X1 is composed of the peak areas of chromatographic peaks p2~p22165 in 16 QC samples, and is a matrix with a size of 16×22,164; y1 is composed of the peak areas of chromatographic peak P1 in 16 samples, and is a matrix with a size of 16×1. For X1, Monte Carlo method sampling is performed. Each time sampling, 11083 chromatographic peaks are randomly drawn from the data set X1 without replacement, and a 16×11083 data matrix is obtained. After independently performing the above Monte Carlo sampling 1000 times, a total of 1000 data matrices are obtained, which are respectively denoted as X1sub1, X1sub2,...., X1sub1000. According to these 1000 data matrices and y1 respectively, using the 3-fold cross-validation method, 1000 SVR sub-models are trained, constructed and verified, which are respectively denoted as SVR1sub1, SVR1sub2,..., SVR1sub100. Thus, the first step operation of the present invention is completed.
[0025] Step 2: For the dataset and SVR sub-models in Step 1, divide the 1000 SVR sub-models into two categories according to whether they contain (or do not contain) chromatographic peaks other than chromatographic peak p1. For example, in the dataset where 1000 SVR models are respectively constructed, there are 506 SVR sub-models that contain chromatographic peak p2 and 494 SVR sub-models that do not contain chromatographic peak p2. The mean difference in the coefficient of determination of the two SVR model sets is 0.137, indicating that adding the information of chromatographic peak p2 can better benchmark the peak area of chromatographic peak p1. Repeat Step 2 and continue to judge whether considering chromatographic peaks p3 to p22165 is beneficial to the correction of chromatographic peak p1. Finally, 694 chromatographic peaks that are beneficial to the correction of chromatographic peak p1 are obtained, and thus the second operation of the present invention is completed.
[0026] Step 3: According to the 694 chromatographic peaks screened out in the above Step 2, the peak areas in 16 QC samples and the peak area of chromatographic peak p1 in 16 QCs are used to form a dataset (X, y) for training and constructing the final SVR model, and thus the third operation of the present invention is completed.
[0027] Step 4: According to the SVR model constructed in the above Step 3, use the peak areas of the 694 chromatographic peaks screened out in Step 2 in 176 samples to correct the chromatographic peak area of chromatographic peak p1 in 176 samples, and thus the fourth operation is completed, that is, the correction of the chromatographic peak area of chromatographic peak p1 is completed.
[0028] Step 5: Repeat the above Steps 1 to 4 to respectively correct the peak areas of chromatographic peaks p2 to p22165 in 176 samples, and thus the correction of all chromatographic peak areas in all samples is completed.
[0029] After using a method for correcting the chromatographic peak area in metabolomics based on model cluster analysis proposed by the present invention to correct the peak areas of 22,165 chromatographic peaks in 176 samples respectively, the statistical distribution of the RSD values of the corrected peak areas in 16 QC samples is recalculated as Figure 3 shown.
[0030] The comparison of the histograms of the QC RSD values before and after the peak area correction of 22,165 chromatographic peaks using a method for correcting the chromatographic peak area in metabolomics based on model cluster analysis proposed by the present invention is shown in the following table.
[0031] Table 1 Comparison of the distribution of chromatographic peak QC RSD values before and after correction using the method of the present invention As shown in Table 1, after using a method for correcting the chromatographic peak area of metabolomics based on model cluster analysis proposed by the present invention to correct the data of the examples, the proportion of samples with a QC RSD value ≤ 0.1 increased from 66.27% to 95.41%, and the peak area consistency of the entire peak table was significantly improved.
[0032] It should be noted that the above-mentioned order of the embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above description of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0033] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
[0034] This specification and the drawings are only exemplary illustrations of the present application and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalent technologies, the present application is intended to include these changes and modifications therein.
Claims
1. A metabolomics chromatographic peak area correction method based on model cluster analysis, characterized in that: A model cluster method is used to determine a sub-data set for chromatographic peak area correction, and a regression model based on a support vector machine is constructed based on the sub-data set to perform chromatographic peak area correction. The method includes: Step 1: For the chromatographic peak to be detected, a SVR sub-model set is established by using a model clustering method; Step 2: Perform statistical analysis on the variables selected by the SVR sub-model set to determine the screening variable set; Step 3: constructing a final SVR model of the chromatographic peak to be detected according to the screening variable set; Step 4: Correcting the chromatographic peak areas of the chromatographic peaks to be detected in all samples according to the SVR model; Step 5: Repeat steps 1 to 4 until the peak area correction of all chromatographic peaks is completed.
2. The metabolomics chromatographic peak area correction method based on model cluster analysis according to claim 1, characterized in that: The step 1 comprises: When the number of QC samples in the chromatogram peak table is , the number of non-QC samples is , the number of chromatographic peaks is When the chromatographic peak area of the QC sample in the chromatographic peak table constitutes a Matrix of For the chromatographic peak to be corrected for peak area , build a dataset ,in, The chromatographic peak to be corrected for peak area is The matrix of the peak areas of all the other chromatographic peaks in the QC sample is , Chromatographic peak The matrix of peak areas in all QC samples is ; For the dataset , using a no-replacement approach, select variables, and obtain a sub-dataset , whose size is , through Monte Carlo sampling After extraction, the total sub-datasets, respectively , and according to Datasets , ,use -Fold cross-validation method, training, construction and verification respectively SVR sub-models, respectively , the determination coefficient of its verification result , respectively .
3. The metabolomics chromatographic peak area correction method based on model cluster analysis according to claim 1, characterized in that: The step 2 comprises: Based on the dataset , sub-dataset , sub-dataset and submodels The definition of Variables to be selected in , ,Will The SVR sub-model is divided into and Two categories, including Indicates that the sub-dataset contains the candidate variables when building the SVR sub-model A collection of sub-models, Indicates that the sub-dataset does not contain the candidate variable when building the SVR sub-model A collection of sub-models; Natural, according to and The coefficient of determination of the two sub-model ensembles ( ) Mean and According to formula (1), the two sub-model sets are calculated for the chromatographic peaks Mean difference in corrected ability; ; (1) like , it indicates that the selection variables The prediction ability of the SVR model will be enhanced when variables Added to the final data subset; like , it indicates that the selection variables The prediction ability of the SVR model will be weakened when variables Added to the final data subset; By setting , perform the above steps respectively, and finally obtain the screening variable set .
4. The metabolomics chromatographic peak area correction method based on model cluster analysis according to claim 1, characterized in that: The step 3 comprises: Obtain the chromatographic peak to be corrected The set of screening variables , and finally construct the dataset ,in, for The variables in The data set consists of the peak areas in QC samples, and its size is , To filter the variable set size; According to the constructed data set , the Gaussian kernel shown in formula (2) is used to train the SVR model; ; (2) The data set is transformed into Mapping to an infinite-dimensional feature space and in the feature space The SVR model is constructed in .
5. The metabolomics chromatographic peak area correction method based on model cluster analysis according to claim 1, characterized in that: The step 4 comprises: For the chromatographic peak to be corrected ,in Indicates chromatographic peak In the Samples (including QC samples and non-QC samples), for chromatographic peaks The constructed SVR model is denoted as , and training The set of screening variables is denoted as ; For the chromatographic peak area to be corrected , first find the set of screening variables The chromatographic peaks in Samples The peak area vector is constructed , and according to the SVR model , predict chromatographic peaks In the Samples Peak area in (including QC samples and non-QC samples) , and its calculation formula is shown in formula (3); ; (3) Finally, the chromatographic peak is obtained according to formula (4): In the Samples Corrected peak areas in (including QC samples and non-QC samples) ; ; (4) make , perform step 4 respectively to complete the chromatographic peak In all Chromatographic peak area calibration in samples (including QC samples and non-QC samples); make , execute steps 1 to 4 respectively to complete the peak area correction of all chromatographic peaks in the peak table.
Citation Information
Patent Citations
Model-cluster-analysis-based laser-induced breakdown spectroscopy variable selection method
CN103487410A
Metabolism group method for distinguishing false positive mass spectra peak signals and quantificationally correcting mass spectra peak area
CN106018600A
Method for establishing forsythia suspensa fingerprint spectrum quality control system
CN110954617A
Method for identifying litchi varieties by using wide targeted metabonomics technology
CN111721857A
Metabolome biomarkers of milk of different processing technologies and screening method and application thereof
CN113671079A