A metabolomics chromatographic peak area correction method based on model cluster analysis
Through the model cluster analysis method, the SVR sub-model set was established and variable screening was performed, the final SVR model was constructed, and the peak area of metabolomics was corrected, which solved the data drift problem in the peak area correction of chromatographic, and achieved a more accurate comparison of metabolic substance content between samples.
Patent Information
- Application Number
- CN202510614522.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-05-14
AI Technical Summary
In the existing metabolomics analysis, the correction of chromatographic peak area is affected by factors such as the sample's early experimental conditions and the attenuation of the instrument's performance, resulting in data drift and it is difficult to accurately reflect the changes in the content of metabolic substances in different samples.
The model cluster analysis method is used to establish the SVR sub-model set, determine the filter variable set through statistical analysis, build the final SVR model, correct the chromatographic peak area, and perform the cycle until all peak area correction is completed.
It improves the robustness of chromatographic peak area correction, ensures that the corrected peak area can reflect the linear relationship between metabolic substances in different samples, and improves the comparability of the same substances in different samples.
Smart Images

Figure CN120142548B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data science and technology, and in particular to a metabolomics chromatographic peak area correction method based on model cluster analysis. Background Art
[0002] Metabolomics uses high-resolution detection instruments to detect the content of metabolites in different types of samples (such as the early, middle and late stages of the disease). By analyzing the patterns of their changes, it explores the changing trends of metabolites in different types of samples. It is an important part of systems biology and is widely used in disease diagnosis, toxicology, botany, nutrition and food science, environmental science and other fields. It has very broad application prospects.
[0003] Because the basis of metabolomics analysis is the detection data of different types of samples by high-resolution detection instruments, which is affected by many factors such as the preliminary experimental conditions of the samples, the noise introduced by the samples and the experimental environment during the detection process, and the data drift caused by the performance degradation of the experimental instruments, it is necessary to perform peak area correction on the chromatographic peak table extracted from metabolomics to ensure that the peak area of the chromatographic peak table can accurately represent the content of the metabolites in the sample. That is, the peak area correction of the chromatographic peak table is an important part of the metabolomics analysis process and is also the basis for ensuring the success of metabolomics analysis.
[0004] Based on this, the present invention applies for a metabolomics chromatographic peak area correction method based on model cluster analysis. This method uses the QC samples of metabolomics to establish a large number of SVR sub-models by using model clustering to determine the sub-datasets ultimately used for SVR modeling, and construct a final SVR model to correct the chromatographic peak area. The present invention adopts the method of model clustering to find the sub-dataset with the greatest contribution to the chromatographic peak area correction SVR model, ensure that the final SVR model constructed has better robustness, obtain the peak area correction effect that best meets the actual situation, and improve the comparability of the same substance in different samples. Summary of the Invention
[0005] The purpose of this application is to provide a metabolomics chromatographic peak area correction method based on model cluster analysis.
[0006] The present application provides a metabolomics chromatographic peak area correction method based on model cluster analysis, the method comprising:
[0007] Step 1: For the chromatographic peak to be detected, a set of SVR sub-models is established by using the model clustering method;
[0008] Step 2: Perform statistical analysis on the variables selected by the SVR sub-model set to determine the screening variable set;
[0009] Step 3: constructing a final SVR model of the chromatographic peak to be detected based on the screening variable set;
[0010] Step 4: Correcting the chromatographic peak areas of the chromatographic peak to be detected in all samples according to the SVR model;
[0011] Step 5: Repeat steps 1 to 4 until the peak area correction of all chromatographic peaks is completed.
[0012] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0013] For the chromatographic peak to be detected, an SVR sub-model set is established by using a model clustering method; the variables selected by the SVR sub-model set are statistically analyzed to determine a screening variable set; a final SVR model of the chromatographic peak to be detected is constructed based on the screening variable set; the chromatographic peak area of the chromatographic peak to be detected in all samples is corrected based on the SVR model; and the aforementioned steps are repeated until the peak area correction of all chromatographic peaks is completed.
[0014] First, with the help of QC sample data with better stability and qualitative properties (generally speaking, the QC sample in a metabolomics experiment is a mixed sample of all other samples, and the QC sample data is the instrument data of the same QC sample detected at different times), chromatographic peak area correction is performed to ensure that the corrected peak area can reflect the linear relationship between metabolites in different samples; secondly, the model clustering method is used for variable selection to avoid the problem of model overfitting caused by the loss of important variables; and the problem of model underfitting caused by the application of unimportant variables, thereby improving the robustness of the peak area correction model, obtaining the peak area correction effect that best conforms to the actual situation, and improving the comparability of the same substance in different samples.
[0015] The above description is only an overview of the technical solution of the present application. In order to more clearly illustrate the technical means of the present application and to implement it in accordance with the contents of the specification, and to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments of the present disclosure are briefly introduced below. Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in precise order. Instead, various steps may be processed in reverse order or simultaneously as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0017] Figure 1This is a basic flow chart of a metabolomics chromatographic peak area correction method based on model cluster analysis of the present invention;
[0018] Figure 2 A distribution histogram of QC RSD values of chromatographic peak areas before chromatographic peak correction in a metabolomics chromatographic peak area correction method based on model cluster analysis of the present invention;
[0019] Figure 3 The present invention provides a metabolomics chromatographic peak area correction method based on model cluster analysis, and a distribution histogram of the QC RSD values of the chromatographic peak area after chromatographic peak correction. DETAILED DESCRIPTION
[0020] The present application provides a metabolomics chromatographic peak area correction method based on model cluster analysis.
[0021] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0022] It should be noted that the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products or devices.
[0023] The present application provides a metabolomics chromatographic peak area correction method based on model cluster analysis, wherein the method comprises:
[0024] For the chromatographic peak to be detected, an SVR sub-model set is established by using a model clustering method; further, the step 1 includes:
[0025] When the number of QC samples in the chromatogram peak table is , the number of non-QC samples is , the number of chromatographic peaks is When the chromatographic peak area of the QC sample in the chromatographic peak table constitutes a Matrix of
[0026] For the chromatographic peak to be corrected for peak area , build a dataset ,in, The chromatographic peak to be corrected for peak area is The matrix of peak areas of all other chromatographic peaks except , Chromatographic peak The matrix of peak areas in all QC samples is ;
[0027] For the dataset , using a no-replacement approach, select variables, and obtain sub-datasets , whose size is , through Monte Carlo sampling After extraction, the total sub-datasets, respectively , and according to Datasets , ,use -Fold cross validation method, training, construction and verification respectively SVR sub-models, respectively denoted as , the coefficient of determination of the verification result , respectively .
[0028] Perform statistical analysis on the variables selected by the SVR sub-model set to determine the screening variable set;
[0029] Furthermore, the step 2 includes:
[0030] Based on step 1, about the dataset , subdataset , subdataset and submodels The definition of Variables to be selected , ,Will The SVR sub-model is divided into and Two categories, including Indicates that the sub-dataset contains candidate variables when building the SVR sub-model A collection of sub-models, Indicates that the sub-dataset does not contain the candidate variables when building the SVR sub-model A collection of sub-models;
[0031] Natural, according to and The coefficient of determination of the two sub-model ensembles ( ) mean and , calculate the two sub-model sets for chromatographic peaks according to formula (1) mean difference in corrected ability;
[0032] ; (1)
[0033] like , it indicates that the variables , it will enhance the predictive ability of the SVR model. variables Added to the final data subset;
[0034] like , it indicates that the variables When the SVR model is used, the prediction ability of the SVR model will be weakened. variables Added to the final data subset;
[0035] By setting , perform the above steps respectively, and finally obtain the screening variable set .
[0036] Constructing a final SVR model for the chromatographic peak to be detected based on the screening variable set;
[0037] Furthermore, the step 3 includes:
[0038] Through the above three steps, the chromatographic peak to be corrected is obtained. The set of screening variables , and finally build the dataset ,in, for The variables in The data set consists of the peak areas in QC samples, and its size is , To filter variable sets size;
[0039] According to the constructed data set , the Gaussian kernel as shown in formula (2) is used to train the SVR model;
[0040] ; (2)
[0041] The data set is transformed into Mapping to an infinite-dimensional feature space and in the feature space Construct the SVR model in .
[0042] Correcting the chromatographic peak areas of the chromatographic peak to be detected in all samples according to the SVR model;
[0043] Repeat steps 1 to 4 until the peak area calibration of all chromatographic peaks is completed.
[0044] Furthermore, the step 4 includes:
[0045] For the chromatographic peak to be corrected ,in Indicates chromatographic peak In the samples Peak area in chromatographic peaks (including QC samples and non-QC samples) The constructed SVR model is denoted as , and training The set of screening variables is denoted as ;
[0046] For the chromatographic peak area to be corrected , first find the set of screening variables The chromatographic peaks in samples Peak area vector , and according to the SVR model , predict chromatographic peaks In the samples Peak areas in (including QC samples and non-QC samples) , its calculation formula is shown in formula (3);
[0047] ; (3)
[0048] Finally, the chromatographic peak is obtained according to formula (4): In the samples Corrected peak area in (including QC samples and non-QC samples) ;
[0049] ; (4)
[0050] make , perform step 4 respectively to complete the chromatographic peak In all Chromatographic peak area calibration in samples (including QC samples and non-QC samples);
[0051] make , execute steps 1 to 4 respectively to complete the peak area correction of all chromatographic peaks in the peak table.
[0052] Specifically, 40 x 4 = 160 urine samples were collected from 40 mice divided into six groups: normal (5 mice), model (8 mice), positive (6 mice), low-dose (7 mice), medium-dose (7 mice), and high-dose (7 mice) groups. These samples were combined with 16 quality control (QC) samples, labeled QC1, QC2, QC3, ..., QC17, resulting in a total of 160 + 16 = 176 experimental samples. Detection was performed using a SCIEX 7600 instrument in positive ionization mode. After peak extraction using XCMS, 22,165 chromatographic peaks were obtained from the experimental data, labeled p1, p2, p3, ..., p22165.
[0053] Based on the peak areas of 22,165 chromatographic peaks in 16 QC samples, the statistical distribution of the RSD values before chromatographic peak area calibration was calculated as follows: Figure 2 shown.
[0054] RSD stands for relative standard deviation, and its calculation formula is shown below.
[0055]
[0056] in, , Indicates the QC RSD value of the chromatographic peak, The average QC RSD values of 22,165 chromatographic peaks are shown.
[0057] The specific implementation steps for correcting the peak areas of 22,165 chromatographic peaks in 176 samples using a metabolomics chromatographic peak area correction method based on model cluster analysis proposed in the present invention are as follows. Without loss of generality, the following steps mainly introduce the peak area correction process of chromatographic peak p1. The peak area correction implementation process of chromatographic peaks p2 to p22165 is completely similar to the peak area correction implementation process of chromatographic peak p1:
[0058] Step 1: In the example data, there are 16 QC samples, 160 non-QC samples, and 22,165 chromatographic peaks. This means that the chromatographic peak areas of the QC samples in the chromatographic peak table form a 16×22,165 matrix, with the cell values representing the chromatographic peak areas of each QC sample. Without loss of generality, when performing peak area correction for chromatographic peak p1, a dataset (X1, y1) is constructed, where X1 is a 16×22,164 matrix consisting of the peak areas of chromatographic peaks p2–p22165 from the 16 QC samples; and y1 is a 16×1 matrix consisting of the peak areas of chromatographic peak P1 from the 16 samples. Monte Carlo sampling is performed on X1, with each sampling period randomly selecting 11,083 chromatographic peaks from dataset X1 without replacement, resulting in a 16×11,083 data matrix. After independently performing the Monte Carlo sampling 1000 times, a total of 1000 data matrices are obtained, denoted as X1sub1, X1sub2, ..., X1sub1000. Based on these 1000 data matrices and y1, a 3-fold cross-validation method is used to train, construct, and validate 1000 SVR sub-models, denoted as SVR1sub1, SVR1sub2, ..., SVR1sub100, respectively. This completes the first step of the present invention.
[0059] Step 2: For the dataset and SVR submodels from step 1, the 1000 SVR submodels were divided into two categories based on whether they included (or excluded) other chromatographic peaks besides peak p1. For example, in the dataset from which 1000 SVR models were constructed, 506 SVR submodels included peak p2, while 494 SVR submodels did not. The difference in the mean coefficient of determination between the two SVR model sets was 0.137, indicating that adding information about peak p2 can better benchmark the peak area of peak p1. Step 2 was repeated to determine whether considering peaks p3-p22165 was beneficial for correcting peak p1. Ultimately, 694 chromatographic peaks were found that were beneficial for correcting peak p1, completing step 2 of the present invention.
[0060] Step 3: Based on the peak areas of the 694 chromatographic peaks screened in step 2 above in the 16 QC samples and the peak area of chromatographic peak p1 in the 16 QC samples, a data set (X, y) is formed to train and construct the final SVR model, thus completing step 3 of the present invention.
[0061] Step 4: Based on the SVR model constructed in step 3 above, the peak areas of the 694 chromatographic peaks screened out in step 2 in 176 samples are used to correct the chromatographic peak area of chromatographic peak p1 in 176 samples. This completes step 4, i.e., the chromatographic peak area correction of chromatographic peak p1.
[0062] Step 5: Repeat steps 1 to 4 above to calibrate the peak areas of chromatographic peaks p2 to p22165 in 176 samples respectively. At this point, the calibration of all chromatographic peak areas in all samples is completed.
[0063] The metabolomics chromatographic peak area correction method based on model cluster analysis proposed in this invention was used to correct the peak areas of 22,165 chromatographic peaks in 176 samples. The statistical distribution of the RSD values of the corrected peak areas in 16 QC samples was recalculated. Figure 3 shown.
[0064] The following table shows the comparison of the QC RSD values of 22,165 chromatographic peaks before and after peak area correction using the metabolomics chromatographic peak area correction method based on model cluster analysis proposed in the present invention.
[0065] Table 1 Comparison of the distribution of chromatographic peak QC RSD values before and after correction using the method of the present invention
[0066]
[0067] As shown in Table 1, after the data of the embodiment were corrected using the metabolomics chromatographic peak area correction method based on model cluster analysis proposed in the present invention, the proportion of samples with QC RSD values ≤ 0.1 increased from 66.27% to 95.41%, and the peak area consistency of the entire peak table was significantly improved.
[0068] It should be noted that the above-mentioned order of the embodiments of the present application is for descriptive purposes only and does not represent the superiority or inferiority of the embodiments. The above description is of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0069] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
[0070] This specification and drawings are merely illustrative of the present application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Obviously, those skilled in the art may make various modifications and variations to this application without departing from the scope of this application. Thus, this application is intended to include such modifications and variations as fall within the scope of this application and its equivalents.
Claims
1. A metabolomics chromatographic peak area correction method based on model cluster analysis, characterized in that: A model clustering method is used to determine a sub-dataset for chromatographic peak area correction, and a support vector machine-based regression model is constructed based on the sub-dataset to perform chromatographic peak area correction. The method includes: Step 1: For the chromatographic peak to be detected, a set of SVR sub-models is established by using the model clustering method; Step 2: Perform statistical analysis on the variables selected by the SVR sub-model set to determine the screening variable set, wherein step 2 includes: Based on the dataset (X i ,y i ), sub-dataset Subdataset and submodels Definition of X i The variable x to be selected i,j , j∈{1,2,…,i-1,i+1,…,m}, divide the N SVR sub-models into and Two categories, including Indicates that the sub-dataset contains the candidate variable x when building the SVR sub-model j A collection of sub-models, Indicates that the sub-dataset does not contain the candidate variable x when building the SVR sub-model j A collection of sub-models; Natural, according to and The coefficient of determination (R) of the two sub-models 2 ) mean i,j,A and mean i,j,B , calculate the two sub-model sets for the chromatographic peak p according to formula (1) i mean difference in corrected ability; Dmean i,j =mean i,j,A -mean i,j,B ; (1) If Dmean i,j ≥0, it indicates that the jth variable p is selected j The predictive ability of the SVR model will be enhanced when the jth variable p j Added to the final data subset; If Dmean i,j <0, it indicates that the jth variable p is selected j The predictive ability of the SVR model will be weakened when the j-th variable p is j Added to the final data subset; By setting j = 1, 2, ..., i-1, i+1, ..., m respectively, and executing the above steps respectively, we can finally obtain the screening variable set. Step 3: constructing a final SVR model for the chromatographic peak to be detected based on the screening variable set; Step 4: Correcting the chromatographic peak areas of the chromatographic peak to be detected in all samples according to the SVR model; Step 5: Repeat steps 1 to 4 until the peak area correction of all chromatographic peaks is completed.
2. The metabolomics chromatographic peak area correction method based on model cluster analysis according to claim 1, characterized in that: The step 1 comprises: When the number of QC samples in the chromatographic peak table is n1, the number of non-QC samples is n2, and the number of chromatographic peaks is m, the chromatographic peak areas of the QC samples in the chromatographic peak table form an n1×m matrix; For the chromatographic peak p to be corrected for peak area i , build a data set (X i ,y i ), where X i is the chromatographic peak p except the peak area correction to be performed i The matrix of the peak areas of all the other chromatographic peaks in the QC sample is n1×(m-1), y i is the chromatographic peak p i The matrix of peak areas in all QC samples is n1×1 in size; For dataset X i , using a no-replacement method, Q variables are selected from it to obtain the sub-dataset X i,sub1 , whose size is n1×Q. After N extractions through Monte Carlo sampling, a total of N sub-datasets are obtained, which are recorded as And based on N data sets The k-fold cross validation method is used to train, build and verify N SVR sub-models, which are respectively denoted as The determination coefficient R of the verification result 2 , respectively 3. The metabolomics chromatographic peak area correction method based on model cluster analysis according to claim 1, characterized in that: The step 3 comprises: Get the chromatographic peak p to be corrected i The screening variable set D i , and finally construct the data set (X' i ,y), where X' i D i The data set consists of the peak areas of the variables in n1 QC samples, and its size is n1×||D i ||,||D i || is the screening variable set D i size; According to the constructed data set (X' i ,y), the Gaussian kernel as shown in formula (2) is used to train the SVR model; The data set (X' i ,y) is mapped into an infinite-dimensional feature space Ω, and an SVR model is constructed in the feature space Ω, denoted as SVR i .
4. The metabolomics chromatographic peak area correction method based on model cluster analysis according to claim 1, characterized in that: The step 4 comprises: For the chromatographic peak to be corrected where p i,j (j=1,2,…,n1+n2) represents the chromatographic peak p i In the jth sample Spec j (including QC samples and non-QC samples), for the peak p i The constructed SVR model is denoted as SVR i , and training SVR i The set of screening variables is denoted as D i ={p i,1 ,p i,2 ,…,p i,t }(t≤m-1); For the chromatographic peak area to be corrected (j=1,2,…,n1+n2), first find the screening variable set D i Each chromatographic peak in the jth sample Spec j The peak area constructs the vector P = {p i,1,j ,p i,2,j ,…,p i,t,j }, and according to the SVR model SVR i , predict chromatographic peak p i In the jth sample Spec j Peak area in (including QC samples and non-QC samples) The calculation formula is shown in formula (3); Finally, the chromatographic peak p is obtained according to formula (4): i In the jth sample Spec j The corrected peak area p in (including QC samples and non-QC samples) i ',j; Let j = 1, 2, ..., n1 + n2, and execute step 4 respectively to complete the chromatographic peak p i Correction of chromatographic peak areas in all n1+n2 samples (including QC samples and non-QC samples); Let i = 1, 2, ..., m, and execute steps 1 to 4 respectively to complete the peak area correction of all chromatographic peaks in the peak table.
Citation Information
Patent Citations
Metabolome biomarkers of milk of different processing technologies and screening method and application thereof
CN113671079A
Biomarker screening method based on wide-target metabonomics and machine learning, chronic kidney disease biomarker group selected by biomarker screening method and application of biomarker group
CN117825480A