Metabonomics chromatographic peak area correction method based on model cluster analysis

Through the model cluster analysis method, an SVR sub-model screening variable set was established, and the final model was constructed for chromatographic peak area correction, which solved the inaccuracy of peak area correction in metabolomics and achieved the comparability of material content between samples.

CN120369871APending Publication Date: 2025-07-25DALIAN CHEM DATA SOLUTION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410107143.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-25
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, the peak area correction method of metabolomics chromatography is affected by factors such as the sample's early experimental conditions and the attenuation of the instrument's performance, resulting in data drift, making it difficult to ensure that the peak area accurately reflects the content of metabolic substances in different samples and affects the comparability between samples.

Method used

The model cluster analysis method is adopted to establish a large number of SVR submodels, filter the variable sets, and build the final SVR model for chromatographic peak area correction to ensure that the correction effect is in line with the actual situation and avoid overfitting or underfitting the model.

Benefits of technology

It improves the robustness of chromatographic peak area correction, improves the comparability of the same substances in different samples, and ensures that the corrected peak area can reflect the linear relationship between metabolic substances in different samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120369871A_ABST
    Figure CN120369871A_ABST
Patent Text Reader

Abstract

The invention discloses a metabonomics chromatographic peak area correction method based on model cluster analysis, and the metabonomics chromatographic peak area correction method is an MPA / MN (Model Pollution Analysis / MN) (Model Pollution Analysis-Metabolomics Normalization) method. According to the method, by means of a QC (Quality Control) sample of metabonomics, a large number of SVR (Support Vector Regression) sub-models are established in a model clustering mode, a sub-data set finally used for SVR modeling is determined, and a final SVR model is constructed to correct the chromatographic peak area. According to the method, a model clustering method is adopted, the sub-data set with the maximum contribution degree for the chromatographic peak area correction SVR model is searched, it is ensured that the constructed final SVR model has better robustness, the peak area correction effect most conforming to the actual situation is obtained, and the comparability of the same substance in different samples is improved. Figure 1 in the attached drawing of the specification is a flow chart of the invention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses a metabolomics chromatographic peak area correction method based on model cluster analysis, which belongs to the field of data science and is used for correcting the metabolomics chromatographic peak area and improving the comparability of the same substance in different samples. Background Art

[0002] Metabolomics uses high-resolution detection instruments to detect the content of metabolites in different types of samples (such as the early, middle and late stages of a disease), and explores the changing trends of metabolites in different types of samples by analyzing their changing patterns. It is an important part of systems biology and is widely used in fields including disease diagnosis, toxicology, botany, nutrition and food science, and environmental science, and has very broad application prospects.

[0003] Because the basis of metabolomics analysis is the detection data of different types of samples by high-resolution detection instruments, it is affected by many factors such as the preliminary experimental conditions of the samples, the noise introduced by the samples and the experimental environment during the detection process, and the data drift caused by the performance attenuation of the experimental instruments. Therefore, it is necessary to perform peak area correction on the chromatographic peak table extracted in metabolomics to ensure that the peak area of the chromatographic peak table can accurately represent the content of the metabolites in the sample. That is, the peak area correction of the chromatographic peak table is an important part of the metabolomics analysis process and is also the basis for ensuring the success of metabolomics analysis.

[0004] Based on this, the present invention develops a metabolomics chromatographic peak area correction method based on model cluster analysis, namely MPA / MN (Model Population Analysis-Metabolomics Normalization) method. This method uses the QC (Quality Control) samples of metabolomics, and adopts the model clustering method to establish a large number of SVR (Support Vector Regression) sub-models to determine the sub-datasets ultimately used for SVR modeling, and construct the final SVR model for chromatographic peak area correction. The present invention adopts the method of model clustering to find the sub-dataset with the greatest contribution to the chromatographic peak area correction SVR model, ensure that the final SVR model constructed has better robustness, obtain the peak area correction effect that best meets the actual situation, and improve the comparability of the same substance in different samples. Summary of the invention

[0005] The problem to be solved by the present invention is to provide a metabolomics chromatographic peak area correction method based on model cluster analysis.

[0006] To achieve the above object, the technical solution adopted by the present invention is:

[0007] A method for correcting the chromatographic peak area in metabolomics based on model cluster analysis, comprising the following steps:

[0008] (1) Step 1: For the chromatographic peaks to be detected, a large number of SVR sub-models are established in the way of model cluster.

[0009] (2) Step 2: Statistical analysis is carried out on the variables selected by all SVR sub-models to determine the screened variable set.

[0010] (3) Step 3: According to the screened variable set, the final SVR model of the chromatographic peaks to be detected is constructed.

[0011] (4) Step 4: Standardize the chromatographic peak areas of the chromatographic peaks to be detected in all samples according to the SVR model.

[0012] (5) Step 5: Loop through Steps 1 to 4 until the peak area correction of all chromatographic peaks is completed.

[0013] Compared with the prior art, the present invention develops a method for correcting the chromatographic peak area in metabolomics based on model cluster analysis, which has the following excellent effects: ① By means of the QC sample data with better data stability (generally speaking, the QC samples in metabolomics experiments are mixed samples of all other samples, and the QC sample data are instrument data detected for the same QC sample at different times), the chromatographic peak table peak area is corrected to ensure that the corrected peak area can reflect the linear relationship of metabolites among different samples; ② The model cluster method is used for variable selection to avoid the problem of model overfitting caused by the loss of important variables and the problem of model underfitting caused by the application of unimportant variables, improve the robustness of the peak area correction model, obtain the peak area correction effect that most conforms to the actual situation, and enhance the comparability of the same substances in different samples. Description of the Drawings

[0014] Figure 1 It is the basic flowchart of the present invention.

[0015] Figure 2 It is the histogram of the QC RSD value of the chromatographic peak table before correction.

[0016] Figure 3 It is the histogram of the QC RSD value of the chromatographic peak table after correction using a method for correcting the chromatographic peak area in metabolomics based on model cluster analysis proposed by the present invention. Detailed Embodiments

[0017] The following further describes the present invention in detail with reference to the specific embodiments of the present invention in conjunction with the drawings:

[0018] The specific embodiment data of the present invention are as follows: 160 mouse urine samples collected from 40 mice divided into 6 groups, namely the normal group (5 mice), the model group (8 mice), the positive group (6 mice), the low-dose group (7 mice), the medium-dose group (7 mice), and the high-dose group (7 mice), at 4 stages of day 0, day 10, day 20, and day 30. Additionally, 16 QC samples (a total of 160 + 16 = 176 experimental samples) were mixed from the above 160 samples. The detection experiment was carried out using the positive ion ionization mode of the SCIEX 7600 instrument, and after peak extraction by XCMS, 22,165 chromatographic peaks were obtained.

[0019] Among the 22,165 chromatographic peaks, the statistical distribution of the RSD values calculated based on the peak areas of the chromatographic peaks in the 16 QC samples is as Figure 2 shown.

[0020] RSD is the Relative Standard Deviation, and its calculation formula is shown as follows.

[0021] Where n = 22165, x i represents the QC RSD value of the i-th chromatographic peak, represents the average value of the QC RSD values of the 22,165 chromatographic peaks.

[0022] After correcting the peak areas of the 22,165 chromatographic peaks in the 176 samples respectively using a metabolomics chromatographic peak area correction method based on model cluster analysis proposed by the present invention, the statistical distribution of the RSD values of the corrected peak areas in the 16 QC samples is recalculated as Figure 3 shown.

[0023] Before and after the peak area correction of the 22,165 chromatographic peaks using a metabolomics chromatographic peak area correction method based on model cluster analysis proposed by the present invention, the histogram comparison of the QC RSD values is shown in the following table. Table 1 Distribution comparison of chromatographic peak QC RSD values before and after correction using the method of the present invention RSD value distribution range of QC samples Before peak area correction After peak area correction ≤0.1 66.27% 95.41% (0.1,0.2] 31.06% 3.74% (0.2,0.3] 2.3% 0.71% >0.3 0.37% 0.14%

[0024] As shown in Table 1, after correcting the example data using a metabolomics chromatographic peak area correction method based on model cluster analysis proposed by the present invention, the proportion of samples with QC RSD values ≤ 0.1 increased from 66.27% to 95.41%, and the peak area consistency of the entire peak table was significantly improved.

Claims

1. A method for correcting the chromatographic peak area in metabolomics based on model cluster analysis, characterized in that Using the method of a model cluster, determine a sub-dataset for chromatographic peak area correction, and based on this, construct a robust support vector regression (SVR) model for chromatographic peak area correction to enhance the comparability of chromatographic peak areas. The method includes the following steps: Step 1: For the chromatographic peaks to be detected, establish a large number of SVR sub-models by using the model cluster method. Step 2: Conduct a statistical analysis on the variables selected by all SVR sub-models to determine a set of screened variables. Step 3: Construct a final SVR model for the chromatographic peaks to be detected according to the set of screened variables. Step 4: Standardize the chromatographic peak areas of the chromatographic peaks to be detected in all samples according to the SVR model. Step 5: Repeat Steps 1 to 4 until the peak area correction of all chromatographic peaks is completed.

2. The SVR sub-model method according to claim 1, wherein The specific method of Step 1 is as follows: Assume that the number of QC samples in the chromatographic peak table is n1, the number of non-QC samples is n2, and the number of chromatographic peaks is m. The chromatographic peak areas of the QC samples in the chromatographic peak table form an n1×m matrix. For the chromatographic peak p to be corrected for peak area i , construct a data set (X i , y i ). Among them, X i is the matrix formed by the peak areas of all chromatographic peaks other than chromatographic peak p i in the QC samples, and its size is n1×(m - 1); y i is the matrix formed by the peak areas of chromatographic peak p i in all QC samples, and its size is n1×1. For the dataset X i , Q variables are randomly selected from it without replacement to obtain the sub-dataset X i,sub1 , whose size is n1×Q. After N extractions by Monte Carlo sampling, a total of N sub-datasets are obtained, which are respectively denoted as And according to the N datasets k = 1, 2, …, N, N SVR sub-models are respectively constructed, which are respectively denoted as 3. The method for determining a sub-data set according to claim 1, wherein The specific method of Step 2 is as follows: Based on the definitions of the dataset (X i , y i ), sub-datasets sub-dataset and sub-models in Step 1, for the variable x i to be selected in X i,j , where j ∈ {1, 2, …, i - 1, i + 1, …, m}, the N SVR sub-models are divided into and two categories, where represents the set of sub-models that contain the variable x j to be selected in the sub-dataset when constructing the SVR sub-model, and represents the set of sub-models that do not contain the variable x j to be selected in the sub-dataset when constructing the SVR sub-model. Naturally, according to and the corresponding number of SVR intervals are respectively denoted as N i,j,A and N i,j,B , respectively based on these two sets of interval data, two corresponding distributions are obtained, and the mean difference between the two distributions is calculated according to Equation (1). Dmean i,j = mean i,j,A - mean i,j,B (1) If Dmean i,j ≥ 0, it indicates that when selecting the j-th variable p j , there is a high probability of the margin of the SVR model. Therefore, the j-th variable p j is added to the final data subset. If Dmean i,j < 0, it indicates that when selecting the j-th variable p j there is a high possibility of increasing the margin of the SVR model, thereby reducing the prediction ability of the model. Therefore, the j-th variable p j is not added to the final data subset. By separately setting \(j = 1, 2, \ldots, i - 1, i + 1, \ldots, m\) and respectively executing the above steps, a screened variable set is finally obtained.

4. The SVR model construction method according to claim 1, wherein The specific method of Step 3 is as follows: Through the above three steps, the screening variable set D for the chromatographic peak p to be corrected is obtained. i Finally, the data set (X′ i , y) is constructed, where X′ i is the data set composed of the peak areas of the variables in D i in n1 QC samples respectively, with a size of n1×||D i ||, and ||D i || is the size of the screening variable set D i . i ​ Based on the constructed dataset (X′ i , y), the Gaussian kernel shown in Equation (2) is used to train the SVR model. The data set (X′ i , y) is mapped into an infinite-dimensional feature space Ω through a Gaussian kernel, and an SVR model is constructed in the feature space Ω, denoted as SVR i .

5. The chromatographic peak correction method according to claim 1, wherein The specific method of Step 4 is as follows: For the chromatographic peak to be corrected where p i,j (j = 1, 2, …, n1 + n2) represents the peak area of the chromatographic peak p i in the j-th sample Spec j (including QC samples and non-QC samples), for the chromatographic peak p i The constructed SVR model is denoted as SVR i , and the screening variable set for training SVR i is denoted as D i ={p i,1 ,p i,2 ,…,p i,t}(t ≤ m - 1). For the chromatographic peak areas to be corrected (\(j = 1, 2, \ldots, n1 + n2\)), first find the screening variable set \(D\). i Construct a vector \(P=\{p\) j , \(p\) i,1,j , \(\ldots\), \(p\) i,2,j \} from the peak areas of each chromatographic peak in the \(j\)-th sample Spec i,t,j in the screening variable set \(D\), and predict the peak area of chromatographic peak \(p\) i in the \(j\)-th sample Spec i (including QC samples and non-QC samples) according to the SVR model SVR j . Its calculation formula is shown in Equation (3). ​ Finally, the chromatographic peak p is obtained according to Equation (4). i In the jth sample Spec j (including QC samples and non-QC samples), the corrected peak area p' i,j . Let \(j = 1, 2, \ldots, n1 + n2\), and perform Step 4 respectively to complete the chromatographic peak \(p\). i Correction of chromatographic peak areas in all \(n1 + n2\) samples (including QC samples and non-QC samples). Let i = 1, 2, …, m, and execute Steps 1 to 4 respectively to complete the peak area correction of all chromatographic peaks in the peak table.