Method for Identifying VOCs Biomarkers in Human Exhaled Breath Based on Machine Learning
Through an integrated algorithm based on machine learning, combined with alveolar gradient method and multi-dimensional statistical method, biomarkers of human expiratory VOCs were screened out, solving the accuracy and comparability of marker screening in the prior art, and achieving efficient and accurate biomarker screening.
Patent Information
- Application Number
- CN202210558092.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-05-19
AI Technical Summary
The prior art has problems such as uncertain sampling technology, high cost, insufficient number of samples, inadequate markers, and difficulty in accurately determining low-content VOCs in the screening of human expiratory VOCs, resulting in poor accuracy and comparability of marker screening.
The integrated algorithm based on machine learning is adopted to obtain and process high-dimensional expiratory VOCs data, combine with the alveolar gradient method to identify the internal and exogenous properties of VOCs, and use single-dimensional and multi-dimensional statistical methods to screen primary biomarkers, and use correlation analysis and Lasso logistic regression model to screen secondary biomarkers, and use random forest algorithms to evaluate the classification and prediction performance of secondary biomarkers.
It realizes efficient screening of biomarkers in massive expiratory VOCs data, improves the accuracy and universality of markers, reduces classification costs, simplifies the analysis process, and improves the sensitivity and specificity of the analysis.
Smart Images

Figure CN115064219B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of biomarker detection, and particularly relates to a method for identifying VOCs biomarkers in human exhaled breath. Background Art
[0002] The components of human exhaled breath can directly reflect the health information of the human body. During the metabolic process of human cells, volatile organic compounds (VOCs) are produced. Under the action of the circulatory system, these substances will pass through various systems and organs and finally reach the lungs, and diffuse into the respiratory tract through blood-gas exchange. Each exhalation contains abundant VOCs extracted from the blood. Therefore, there is a certain correlation between the organic substances enriched in the blood and the components of exhaled breath. External stimuli or internal physiological reactions can cause an increase or decrease in the content of some VOCs, or new compounds may be produced. Therefore, exhalation can directly reflect the current physiological state of cells, tissues, and microorganisms in the body, and thus provide information about personal health. The analysis and research of the components of respiratory gas have also become an emerging means for environmental health assessment.
[0003] At present, respiratory detection has been preliminarily studied from different perspectives. More than 3,000 volatile compounds have been detected in human exhaled breath at present. These compounds include aldehyde and ketone compounds, alkanes, alkenes, nitrogen-containing compounds, sulfur-containing compounds, etc. These compounds mainly have three sources. One is the VOCs produced by human body's own metabolism, called endogenous VOCs; the second is the VOCs that enter the human body through inhalation, diet, drugs, skin contact and other channels, and they will be metabolized and consumed in the body, called exogenous VOCs; the third is the VOCs released by the parasitic microorganisms in the host body, and these compounds can also reflect the health status of the host. Abroad, there have been early screening and diagnosis tests for exhaled breath VOCs. By comparing the exhaled breath VOCs of cancer patients, inflammatory patients, and chronic disease patients, attempts have been made to find VOCs biomarkers indicating the corresponding diseases. However, due to problems such as sampling methods, analysis methods, and biomarker screening processes, the progress has been very slow. Since there is no standardized sampling and analysis method for exhaled breath VOCs at home and abroad, the differences between the results are large and not comparable, which also severely restricts the development of respiratory detection. As a non-invasive method, respiratory detection has broad prospects in disease diagnosis. In the future, combined with the research of new technologies such as artificial intelligence, it will surely play an important role.
[0004] In addition, there are still many problems in the screening of exhaled VOC biomarkers, resulting in poor consistency of the biomarkers found in various studies. First, the sampling technology severely limits the screening of exhaled VOC biomarkers. Currently, most studies use Tedlar bags as the main sampling tool, but the sampling volume cannot be determined, and the bags themselves will release pollutants, affecting the target substances in the bags. Second, the cost of exhaled gas sampling is relatively high, and the sample size of most studies is insufficient, resulting in the lack of universality of the screened biomarkers. At the same time, the content of human exhaled VOCs fluctuates greatly, and the sampling conditions are not unified and standardized, seriously affecting the stability of the biomarkers. In addition, the content of exhaled VOCs is extremely low, posing a great challenge to the analysis technology. Therefore, a suitable enrichment technology is needed to accurately measure the content of exhaled VOCs. Moreover, the collected VOCs have two sources, endogenous and exogenous. Most studies do not distinguish them and directly screen biomarkers, ignoring the influence of environmental factors, resulting in a decrease in the accuracy of the biomarkers. Finally, how to screen high-specificity, sensitivity, and accuracy biomarkers from a large amount of exhaled VOC data is a major problem. Currently, most studies only perform simple statistical analysis or a single screening, and the performance of the obtained biomarkers is not good.
[0005] The present invention mainly focuses on solving the difficult problem of screening exhaled VOC biomarkers, and creates an integrated algorithm based on machine learning to efficiently screen biomarkers from a large amount of exhaled VOCs. Summary of the Invention
[0006] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a method for identifying VOC biomarkers in human exhaled gas based on machine learning, which has a suitable sample size, a simple process, and high precision.
[0007] The present invention provides a method for identifying (screening) VOC biomarkers in human exhaled gas based on machine learning, including, for the obtained high-dimensional exhaled VOC data, finding a set of simplified biomarker parameters of exhaled VOCs through statistical and machine learning methods, representing the main information in the entire dataset without losing important information, and generating a classification and recognition ability basically the same as that of the full exhaled VOC parameters.
[0008] The method for identifying VOC biomarkers in human exhaled gas based on machine learning provided by the present invention specifically comprises the following steps:
[0009] S1: Acquisition and processing of high-dimensional exhaled VOC data; including acquiring high-dimensional exhaled VOC data of a certain number (such as 3 liters) of specific diseases (or pathological reactions) and healthy people, and constructing an experimental group and a control group;
[0010] S2: Identifying the endogenous and exogenous properties of VOCs by using the alveolar gradient method;
[0011] S3: Screen the primary biomarkers using univariate statistics and multivariate statistics methods;
[0012] S4: Construct a combined biomarker through correlation analysis, and screen the secondary biomarkers using the Lasso logistic regression model;
[0013] S5: Use the random forest algorithm to evaluate the classification and prediction performance of the secondary biomarkers.
[0014] The specific content of each step is further introduced below.
[0015] (1) The acquisition and processing of the high-dimensional exhaled VOCs data described in step S1, the specific process is as follows:
[0016] Volunteers are divided into an experimental group and a control group. After sitting quietly and calming down, a certain amount (for example, 3L) of breath samples are accurately collected in an adsorption tube (Tenax TA + Carbograph-5TD) using ReCIVA (Owlstone); the samples are analyzed by thermal desorption-two-dimensional gas chromatography-time-of-flight mass spectrometry (TD-GC×GC-TOF MS) to obtain high-dimensional VOCs data; the mass spectrometry peak data is compared with the NIST spectral library (v.2.3) to confirm the chemical components of each VOC; the data is corrected by internal standards to generate the ratio of the peak area of each VOC to the peak area of the internal standard as the relative content of each VOC; outlier tests and missing value imputations are performed on the mass spectrometry data of all samples.
[0017] (2) The specific process of identifying the endogenous and exogenous properties of VOCs using the alveolar gradient described in step S2 is as follows:
[0018] Simultaneously during the sample data collection stage, a hand-held gas sampler (Advantest) is used to collect the same amount (for example, 3L) of ambient air samples in an adsorption tube. The analysis method of the ambient air samples is the same as that of the breath samples. The alveolar gradient AG of a certain VOC is calculated by the following formula (1):
[0019]
[0020] In the formula, A Sample,VOCi refers to the peak area of a certain VOC in the sample, A Sample,IS refers to the peak area of the internal standard in the sample, A Air,VOCi refers to the peak area of a certain VOC in the ambient air, A Air,ISRefers to the peak area of the internal standard in ambient air. If AG > 0, the VOCs are considered endogenous; if AG < 0, the VOCs are considered exogenous. The main basis of this theory is that the content of endogenous VOCs in exhaled breath must be higher than that of the VOCs in ambient air. However, for some VOCs with very low content, if the enrichment effect is not obvious, AG may be very close to 0, and there may be large errors and uncertainties in the judgment of the endogenous and exogenous properties of VOCs. Therefore, the alveolar gradient of each VOC in each sample is further calculated, and the probability P that each VOC belongs to the endogenous source among all people is statistically analyzed. Endo,i (Formula 2).
[0021]
[0022] In the formula, N i,G>0 represents the number of times that the alveolar gradient of compound i is greater than 0 among all subjects; N represents the number of subjects.
[0023] If P Endo,i is greater than 50%, the compound is listed as an endogenous compound; otherwise, it is regarded as an exogenous compound.
[0024] (3) In step S3, the one-dimensional statistical method and multi-dimensional statistical method are used to screen the primary biomarkers. The specific process is as follows:
[0025] Use the one-dimensional statistical test method to identify the primary biomarkers. By performing a normality test on the kurtosis and skewness of the data distribution of each variable, it is determined whether to use parametric tests (data is normally distributed) or non-parametric tests (data is not normally distributed). If the kurtosis and skewness of the data distribution of each variable are close to 0, it is a normal distribution; otherwise, it is a non-normal distribution. Use SPSS Statistics statistical analysis software to calculate the significance p-value. And calculate the fold change FC according to the mean values of each group of variables (Formula 3):
[0026]
[0027] In the formula, is the average value of the i-th VOC in the diseased group, is the average value of the i-th VOC in the control group.
[0028] And draw a volcano plot. Screening the primary biomarkers with the conditions of p < 0.05, FC < 0.5 or FC > 2. Although the one-dimensional statistical method can only give a significance p-value and the change of the mean value between groups, it cannot quantitatively evaluate the significant differences and cannot construct the relationship between various exhaled VOCs, but this method can simply and intuitively confirm whether there are significant differences between the disease group and the control group.
[0029] Using multivariate statistical methods, potential primary biomarkers were further explored. Supervised orthogonal partial least squares discriminant analysis (OPLS-DA) was performed on RStudio. Compared with univariate statistical analysis, OPLS-DA can eliminate the influence of uncontrollable variables on the data through top-down classification and further quantify the differences caused by characteristic variables between different groups. To verify the reliability of the OPLS-DA model, the sample sequences were randomly permuted, and after 200 random permutation tests, the differences between the two groups were evaluated using the OPLS-DA scatter plot. After the model was fitted, the output parameters R 2 (y) and Q 2 values can be used to evaluate the applicability and predictive ability of the model. R 2 (y) represents the percentage of the matrix information in the y-axis direction explained by the OPLS-DA model; Q 2 represents the prediction rate of the model. If both of these values are greater than 0.5, it indicates that the model has good classification performance, and the closer to 1, the better the model fits the data. The variable importance in the projection (VIP) is the variable weight value of the variables in the OPLS-DA model and can be used to measure the influence intensity and explanatory ability of the cumulative differences of each variable on the classification discrimination of each group of samples. In the present invention, the importance of exhaled VOCs was ranked based on the VIP value, and variables with a VIP value greater than 1 (p < 0.05) could be selected into the list of primary biomarkers. This method will supplement the primary biomarkers whose fold change does not meet the conditions of FC < 0.5 or FC > 2. Considering the results of both univariate and multivariate statistical methods, the FC screening criteria and VIP screening criteria can be appropriately adjusted to determine suitable primary biomarkers.
[0030] (4) In step S4, constructing a combined biomarker through correlation analysis and screening secondary biomarkers using the Lasso logistic regression model, the specific process is as follows:
[0031] Reconstructing biomarkers; for different exhaled VOCs, they may be produced by the same enzyme or similar metabolic pathways. According to their ratios, the connection between variables can be strengthened. The Pearson correlation coefficient is used for data with a normal distribution, and the Spearman rank correlation coefficient is used for data with a non-normal distribution. The SPSS Statistic statistical analysis software is used to obtain a list of correlation coefficients between variables. Taking the correlation coefficient greater than or equal to 0.6 as the standard, a group of compounds meeting this condition are reconstructed by ratio, and all the reconstructed ratios become new independent biomarkers.
[0032] Simplify the biomarkers and extract key features. Considering that the number of the above - screened first - level biomarkers and ratio biomarkers may be large, the Lasso (Least Absolute Shrinkage and Selection Operator) - logistic regression (LLR) model can be used to refine the input variables in the subsequent diagnostic model to be generated. The present invention uses the LLR model to perform L1 regularization on all single biomarkers and combined biomarkers, shrinking the regression coefficients of each variable in the model through maximum likelihood estimation. The coefficients of relatively unimportant variables will be shrunk to 0, and these variables can be excluded based on this criterion. This penalized estimation method prevents any overfitting that may occur due to the collinearity or high - dimensionality of the independent variables. The number of covariates in the model decreases as the tuning parameter λ increases, but the fitting error does not decrease unidirectionally with the change of λ. According to the results given by the model, when λ is between λ min and λ 1se , the model error is within an acceptable range. λ min represents the value of λ when the mean squared error is the smallest; λ 1se represents the value of λ at one standard error away from the smallest mean squared error. One can determine to select λ min to generate a model with relatively more variables or λ 1se to generate the most simplified model according to the research needs. Substitute the selected λ into the model to reconstruct the LLR model, and determine the variables with non - zero regression coefficients as the second - level biomarkers. These biomarkers contribute the most to distinguishing the differences between the two groups, and the number of second - level biomarkers is also much smaller than the sum of the number of first - level biomarkers and combined biomarkers. Based on these biomarkers, the model prediction score (PS) formula can be obtained:
[0033]
[0034] In the formula, C i is the regression coefficient of each non - zero variable; Var i is the non - zero variable i; n is the number of non - zero variables.
[0035] (5) The classification and prediction performance of the second - level biomarkers are evaluated using the random forest algorithm, and the specific process is as follows:
[0036] Use the random forest (RF) model to test the sensitivity, specificity, and accuracy of the second - level biomarkers. RF is a hierarchical non - parametric modeling method and is robust to outliers and the correlations between compounds. RF is superior to some other multivariate classification algorithms for pattern recognition because it can detect the non - linear relationship between compounds and outcomes. Since it is less susceptible to overfitting, RF provides stronger prediction and better error measurement. In addition, RF has the advantage of incorporating randomness into its prediction by repeated bootstrap sampling and random variable selection when generating a single decision tree. Perform biomarker validation in RStudio and optimize the hyperparameter m tryConstruct an RF classification model after and ntrees. Determine the sensitivity, specificity, and accuracy of the secondary biomarker through the Receiver Operating Characteristic (ROC) curve.
[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0038] (1) The present invention can achieve accurate classification of different groups by selecting only a small number of key biomarkers, with good reliability, high sensitivity, and strong specificity;
[0039] (2) The method of the present invention can greatly reduce the classification cost;
[0040] (3) It is beneficial to develop a faster heating program for key compounds, reducing the analysis time and pretreatment time;
[0041] (4) The streamlined key variables can simplify the originally cumbersome and even potentially contradictory chemical explanations, facilitating the focused exploration of important metabolic processes and mechanism principles. Brief Description of the Drawings
[0042] Figure 1 is the screening process of exhaled breath VOC biomarkers.
[0043] Figure 2 is to identify the endogenous and exogenous properties of VOCs in the alveolar gradient. Among them, (a) control group; (b) experimental group.
[0044] Figure 3 is the screening of primary biomarkers. Among them, (a) univariate statistical method; (b) multivariate statistical method.
[0045] Figure 4 is to screen secondary biomarkers by the LLR model. Among them, (a) binomial distribution deviation; (b) regression coefficient.
[0046] Figure 5 is the ROC curve of the classification and prediction ability of secondary biomarkers. Detailed Embodiments
[0047] In order to enable those skilled in the art to better understand the technical solutions in this application, the following provides a more comprehensive description of each step and further elaborates with examples. It should be noted that the described embodiments are only a part of the embodiments of this application, not all of them, and the present invention is not limited by the following embodiments.
[0048] Example 1
[0049] The present invention is applicable to the screening of biomarkers for high-dimensional exhaled breath VOC data, and its screening process is as Figure 1 shown.
[0050] (1) Obtain high-dimensional exhaled breath VOC data
[0051] Vaccination is one of the effective ways to prevent the spread of the COVID-19 pandemic. Inactivated vaccine inoculation can trigger an immune response, which may affect the body's metabolic processes and subsequently lead to changes in exhaled VOCs. In this case study, the vaccinated population was selected as the experimental group, and the unvaccinated population as the control group. Fifty subjects who had not received the COVID-19 vaccine were recruited, including 23 males and 27 females, aged between 22 - 58 years (average age 34 years), with a BMI index ranging from 16.87 - 32.49 kg / m 2 (average BMI index 23.28 kg / m 2 ). Fifty-four subjects who had received the second dose of the COVID-19 vaccine were recruited, including 24 males and 30 females, aged between 22 - 30 years (average age 24 years), with a BMI index ranging from 16.42 - 29.58 kg / m 2 (average BMI index 21.09 kg / m 2 ). 3L of exhaled breath samples were collected from each subject and analyzed using TD-GC×GC-TOF MS. Hundreds of exhaled VOCs were detected. 100 compounds with an occurrence rate higher than 60% were selected as target compounds for further analysis, including 23 alkanes, 7 alkenes, 4 alcohols, 4 aldehydes, 14 ketones, 5 acids, 7 esters, 8 aromatic compounds, and other nitrogen-, oxygen-, and sulfur-containing compounds (Table 1). The Chromspace software (SepSolve Analytical) was used to calculate the peak areas of various VOCs, and internal standards were used for normalization correction. Missing values were filled with the mean value.
[0052] Table 1 Qualitative analysis list of exhaled VOCs
[0053]
[0054]
[0055]
[0056]
[0057] (2) Alveolar gradient to identify the endogenous and exogenous properties of VOCs
[0058] The alveolar gradient can roughly determine whether a VOC is of alveolar origin or inhaled from the environment. The alveolar gradient of each VOC is not fixed and will vary depending on the differences in its production and consumption rates in the body. In this case study, the alveolar gradients of various VOCs were calculated for both the experimental group and the control group, and their endogenous rate P Endo,i ( Figure 2 ). Among the 100 target compounds, 72 VOCs had a PEndo,i If it is greater than 50%, these compounds are endogenous VOCs (the first 72 compounds in Table 1). The remaining 28 compounds are exogenous VOCs (the last 28 compounds in Table 1). The subsequent screening of vaccine-induced biomarkers will be selected from these 72 VOCs. Currently, few studies distinguish between endogenous and exogenous compounds in exhaled breath, which may lead to misidentifying some exogenous compounds as potential biomarkers and obtaining false results. In this case, the alveolar gradient is used to distinguish between endogenous and exogenous sources, simply and quickly eliminating the influence of environmental background interference on the accuracy of exhaled breath VOC biomarkers.
[0059] (3) Screening of primary biomarkers using univariate statistics and multivariate statistics
[0060] The VOCs exhaled by the vaccine group and the control group of unvaccinated subjects were basically non-normally distributed. Therefore, non-parametric tests were used for univariate statistical analysis. To eliminate the influence of factors such as age and gender, paired homogeneous experimental and control group data were used for differential analysis. If there are no paired samples, it is necessary to ensure that factors such as age and gender in the two groups are as close as possible, or use statistical methods to correct the influence of these factors. The results of the Wilcoxon signed-rank test showed that the p-values of 43 endogenous VOCs were less than 0.05, showing significant differences. Based on the fold change ( Figure 3 a), 26 VOCs showed downregulation, of which 11 VOCs showed significant downregulation (FC < 0.5); 17 VOCs showed upregulation, of which 1 VOC showed significant upregulation (FC > 2).
[0061] The results of OPLS-DA analysis showed that ( Figure 3 b), there were obvious different clusters between the vaccine group and the control group, and the R 2 intercept was 0.154, and the Q 2 intercept was -0.876. The model showed good applicability (R 2 Y = 0.832, p < 0.01) and predictive ability (Q 2 = 0.756, p < 0.01). The VIP values of 24 compounds were greater than 1, which basically included the differential compounds screened by the univariate statistical method, indicating that there was good consistency between the two methods. Considering both univariate and multivariate statistics, as well as the occurrence rate of compounds (> 90%), a total of 21 VOCs were finally selected as primary biomarkers (Table 2).
[0062] Table 2 List of primary biomarkers of exhaled breath VOCs related to COVID-19 vaccines
[0063]
[0064] (4) Screening of secondary biomarkers
[0065] Analyze the correlation of 21 single biomarkers and strengthen the connection between biomarkers. The Spearman rank correlation coefficient indicates that the correlation coefficients of 26 pairs of compounds are greater than 0.6. Therefore, these 26 pairs of compounds are reconstructed into ratio biomarkers in the form of ratios. The 26 ratio biomarkers are independent of the 21 single biomarkers.
[0066] Simplify the biomarker composition through the LLR model. Set 70% of the data in the experimental group and the control group as the training set, and the remaining 30% of the data as the test set. The LLR model uses 5-fold cross-validation to perform L1 regularization on the data, and the range of λ is 0.000167 - 0.313. As λ increases, the number of variables in the model continuously decreases. The binomial distribution deviation indicates that between the λ min and λ 1se , the model error is small ( Figure 4 a). Within this range, the classification and prediction capabilities of the model are good. In this case, select λ 1se to generate a more simplified biomarker combination, which includes a total of 12 biomarkers, including 9 single biomarkers and 3 ratio biomarkers. These 12 biomarkers are secondary biomarkers ( Figure 4 b). In addition, the calculation formula for the model prediction score can also be obtained:
[0067]
[0068] (5) Validation and evaluation of secondary biomarkers
[0069] Perform biomarker validation in RStudio. When optimizing the model, set 70% of the data in the two groups as the training set, and the remaining 30% of the data as the test set. First, optimize the hyperparameter m try . In the case of the default 500 decision trees, the range of m try is from 1 to 10, and the model error is the smallest when m try = 3. On this basis, optimize the number of decision trees ntrees. Set the range of the number of decision trees for the optimized model to 1 - 1000 (the upper limit can be increased if necessary), and select the number of decision trees that first reaches a stable and low out-of-bag error within the range. Too many decision trees may lead to overfitting and an increase in the model error. In this case, finally determine ntrees = 200. Under this condition, the out-of-bag error of the model is less than 5%. Reconstruct the RF classification model based on the two optimized hyperparameters. 70% of the original data is used as the training set, and the remaining 30% is used as the test set. Through loop operations, test the results of 1000 random trainings. The ROC curve results show that the average sensitivity of the secondary biomarkers is 98.33%, the average specificity is 95.98%, the average area under the curve (AUC) is 0.89, and the overall accuracy is 97.21% ( Figure 5)。Compared with the results of other studies of the same type, the sensitivity, specificity, and accuracy of the biomarkers screened here are better. Therefore, through the invented integrated exhaled VOC biomarker identification method based on machine learning, we successfully screened out 12 biomarkers induced by inactivated COVID-19 vaccines.
Claims
1. A method for identifying VOCs biomarkers in human exhaled breath based on machine learning, characterized in that, It includes obtaining high-dimensional exhaled VOCs data, and by means of statistics and machine learning methods, finding a set of simplified biomarker parameters of breath VOCs to characterize the main information in the entire dataset without losing important information, and generating a classification and recognition ability basically the same as that of the full exhaled VOCs parameters; the specific steps are as follows: S1: Acquisition and processing of high-dimensional exhaled VOCs data; including obtaining high-dimensional exhaled VOCs data of a certain number of people with specific diseases or pathological reactions and healthy people, and constructing an experimental group and a control group; S2: Using the alveolar gradient method to identify the endogenous and exogenous properties of VOCs; S3: Screening primary biomarkers based on endogenous property VOCs using univariate statistics and multivariate statistics methods; S4: Constructing combined biomarkers through correlation analysis and screening secondary biomarkers using the Lasso logistic regression model; S5: Using the random forest algorithm to evaluate the classification and prediction performance of secondary biomarkers.
2. The method for identifying VOCs biomarkers in human exhaled breath based on machine learning according to claim 1, characterized in that, In step S1, the acquisition and processing of the high-dimensional exhaled VOCs data has the following specific process: Volunteers are divided into an experimental group and a control group, and a certain amount of breath samples are collected in an adsorption tube using ReCIVA; the samples are analyzed by thermal desorption-two-dimensional gas chromatography-time-of-flight mass spectrometry to obtain high-dimensional VOCs data; the mass spectrometry peak data is compared with the NIST spectral library to confirm the chemical components of each VOC; the data is corrected by internal standard to generate the ratio of the peak area of VOCs to the peak area of the internal standard as the relative content of each VOC; outlier tests and missing value imputation are performed on the mass spectrometry data of all samples.
3. The method for identifying VOCs biomarkers in human exhaled breath based on machine learning according to claim 2, characterized in that, In step S2, the process of using the alveolar gradient to identify the endogenous and exogenous properties of VOCs is as follows: Simultaneously during the sample data collection stage, a handheld gas sampler is used to collect the same amount of ambient air samples as above in an adsorption tube; the analysis method of the ambient air samples is the same as that of the breath samples; the alveolar gradient AG of a certain VOC is calculated by the following formula (1): Where A Sample,VOCi refers to the peak area of a certain VOC in the sample, A Sample,IS refers to the peak area of the internal standard in the sample, A Air,VOCi refers to the peak area of a certain VOC in ambient air, A Air,IS refers to the peak area of the internal standard in ambient air; if AG > 0, it is considered that the VOCs are endogenous, and if AG < 0, it is considered that the VOCs are exogenous; the main basis is that the content of endogenous VOCs in exhaled breath must be higher than that of the VOCs in ambient air; however, for some VOCs with very low contents, if the enrichment effect is not obvious, AG may be very close to 0, and there may be large errors and uncertainties in the judgment of the endogenous and exogenous properties of VOCs at this time; therefore, the alveolar gradient of each VOC in each sample is further calculated, and the probability P Endo,i : Wherein, N i,G>0 represents the number of times that the alveolar gradient of compound i is greater than 0 among all subjects; N represents the number of subjects; If P Endo,i is greater than 50%, then list this kind of compound as an endogenous compound; otherwise, consider it as an exogenous compound.
4. The method for identifying VOCs biomarkers in human exhaled breath based on machine learning according to claim 3, characterized in that, In step S3, the process of screening primary biomarkers based on endogenous property VOCs using univariate statistics and multivariate statistics methods is as follows: Using the univariate statistical test method to identify primary biomarkers; Performing a normality test on the kurtosis and skewness of the data distribution of each variable to determine whether to use parametric tests or non-parametric tests; If the kurtosis and skewness of the data distribution of each variable are close to 0, it is a normal distribution, otherwise it is a non-normal distribution; the SPSS Statistics statistical analysis software is used to calculate the significance p value, and the fold change FC is calculated according to the mean values of each group of variables; wherein, is the average value of the i-th VOC in the diseased group, is the average value of the i-th VOC in the control group; Drawing a volcano plot; screening primary biomarkers with the conditions of p < 0.05, FC < 0.5 or FC > 2; Use the multivariate statistical test method to further explore potential primary biomarkers; perform OPLS-DA on RStudio; after 200 random permutation tests, evaluate the differences between the two groups using the OPLS-DA scatter plot; R 2 (y) and Q 2 values are used to evaluate the applicability and predictive ability of the model; rank the importance of variables according to the VIP values, and select variables with VIP values greater than 1 and p < 0.05 into the list of primary biomarkers; thus, primary biomarkers with fold change not meeting the conditions of FC < 0.5 or FC > 2 will be supplemented; comprehensively consider the results of univariate and multivariate statistical methods, adjust the FC screening criteria and VIP screening criteria, and determine appropriate primary biomarkers; here, R 2 (y) represents the percentage of the OPLS-DA model explaining the matrix information in the y-axis direction; Q 2 represents the prediction rate of the model; if both of these values are greater than 0.5, it indicates that the model has good classification performance, and the closer to 1, the better the model fits the data; the VIP value is the variable weight value of the OPLS-DA model variables, which is used to measure the influence intensity and explanatory ability of the cumulative differences of each variable on the classification and discrimination of each group of samples.
5. The method for identifying VOCs biomarkers in human exhaled breath based on machine learning according to claim 4, characterized in that, In step S4, the process of constructing combined biomarkers through correlation analysis and screening secondary biomarkers using the Lasso logistic regression model is as follows: Refactored markers: For different exhalations, VOCs may be produced by the same enzyme or similar metabolic pathways. According to their ratios, the connections between variables are strengthened. For data with a normal distribution, the Pearson correlation coefficient is used, and for data with a non-normal distribution, the Spearman rank correlation coefficient is used. The SPSS Statistic statistical analysis software is applied to obtain a list of correlation coefficients between variables. Taking the correlation coefficient greater than or equal to 0.6 as the standard, a group of compounds meeting this condition are subjected to ratio refactoring, and all refactored ratios become new independent markers; Simplified markers, extracting key features; The L1 regularization is performed on all single markers and combined markers using the LLR regression model, and the regression coefficients of each variable in the model are shrunk through maximum likelihood estimation; the variables in the model decrease as the tuning parameter λ increases, but the fitting error does not decrease unidirectionally with the change of λ; according to the results given by the model, the acceptable range of λ is between λ min and λ 1se , and λ min or λ 1se can be determined according to the research needs; Substitute the selected λ into the model to reconstruct the LLR model, and determine the variables with non-zero regression coefficients as secondary biomarkers, which contribute the most to distinguishing the differences between the two groups. The model prediction score PS can be obtained based on these biomarkers: where C i is the regression coefficient of each non-zero variable; Var i is the non-zero variable i; n is the number of non-zero variables.
6. The method for identifying VOCs biomarkers in human exhaled breath based on machine learning according to claim 5, characterized in that, In step S5, the random forest algorithm is used to evaluate the classification and prediction performance of secondary biomarkers. The specific process is as follows: Use the RF model to test the sensitivity, specificity, and accuracy of secondary biomarkers; perform biomarker validation in RStudio and optimize the hyperparameters m try and ntrees to construct an RF classification model; determine the sensitivity, specificity, and accuracy of secondary biomarkers through the receiver operating characteristic curve.
Citation Information
Patent Citations
Expiratory air detection device and establishment method of expiratory air marker of expiratory air detection device
CN111710372A
Lung cancer staging detection system based on VOC metabolic profile
CN114384160A