Pulmonary nodule malignant risk prediction system based on exhaled gas VOCs

Through the pulmonary nodule malignant risk prediction system based on exhaled gas VOCs, a random forest model combined with sample data processing is used to solve the problem of inaccurate prediction caused by doctors' experience dependence, and a more accurate malignant risk assessment is achieved.

CN120148841APending Publication Date: 2025-06-13JIANGSU RUIZHI BIOTECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510164641.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the prior art, doctors rely on experience to predict malignant risks of pulmonary nodules, and the prediction results are largely different and inaccurate.

Method used

A malignant risk prediction system for pulmonary nodules based on exhaled gas VOCs is adopted to automatically predict patients' malignant risk levels through the combination of sample data collection, processing and random forest models.

Benefits of technology

It improves the accuracy of predicting malignant risks of pulmonary nodules, reduces physician experience dependence, and provides more reliable risk assessment results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148841A_ABST
    Figure CN120148841A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of medical treatment, and particularly relates to a pulmonary nodule malignant risk prediction system based on exhaled gas VOCs, which comprises a sample data collection system, a sample processing system and a risk prediction system, the sample data collection system is used for collecting and analyzing patient exhaled gas VOCs, epidemiological data, iconography signs and laboratory indexes to form a sample data set, the sample processing system is used for preprocessing all features in the sample data set, the preprocessed sample data set is divided into a test set and a training set, and the test set and the training set are connected with the sample data collection system. All features of the patients in the divided test set and training set are screened, and the risk prediction system comprises a random forest model. On the basis of the random forest, on the basis of the VOCs characteristics of the exhaled gas and in combination with epidemiological data, imaging signs and laboratory index characteristics, the pulmonary nodule malignant risk prediction system based on VOCs volatilized by the exhaled gas is provided, and finally the risk level of a patient can be automatically predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical technologies, and particularly relates to a lung nodule malignant risk prediction system based on exhaled gas VOCs. Background Art

[0002] Lung cancer is one of the cancers with the highest incidence and mortality rates globally. The prevention and treatment of lung cancer pose a major challenge to the prevention and control of malignant tumors in China. Due to the lack of obvious symptoms in the early stage, most lung cancer patients (about 75%) are diagnosed at an advanced stage (stage III / IV), and the 5-year survival rate is less than 15%; if detected in the early stage (stage I), the 5-year survival rate can be as high as 70%-90%. Therefore, the screening and early diagnosis of lung cancer have important public health and clinical significance.

[0003] Lung nodules are the earliest lesion forms in the vast majority of lung cancer patients. According to expert estimates, there are currently 100 million lung nodule patients in China, among which 80 million are small nodules less than 8 mm, and only 2%-3.6% of them are malignant lung nodules (i.e., lung cancer). If malignant lung nodules are not intervened in the early stage, they have a high degree of malignancy, rapid disease progression, and poor prognosis. If the lesion can be surgically removed in the early stage, it will significantly improve the prognosis of lung cancer patients.

[0004] Commonly used lung cancer screening methods are mainly divided into three categories: imaging examinations, pathological examinations, and circulating tumor markers. They all have their own unique advantages and also have undeniable drawbacks. Imaging technologies (such as LDCT) have higher resolution, are more sensitive in showing small nodules in the lungs, and have accurate positioning, but they have deficiencies such as a high false positive rate and overdiagnosis. In the long-term follow-up of lung nodules, imaging technologies bring risks of radiation exposure and cause high physical and mental stress to patients, while increasing the economic burden on patients and society. Pathological diagnosis by tissue biopsy is currently the gold standard for lung cancer diagnosis. Whether it is percutaneous biopsy through the respiratory tract or the skin, it is an invasive operation with risks such as bleeding, infection, and pneumothorax, and is mainly used for the diagnosis of advanced lung cancer. And now, more and more small lung nodules are not suitable for endobronchial ultrasound biopsy. For small nodules in the periphery, the difficulty of percutaneous lung biopsy is high, the positive rate is difficult to guarantee, and there is also a risk of distant metastasis. There have been many advancements in the use of circulating tumor markers for lung cancer diagnosis, such as lung cancer autoantibodies, complement fragments, micro-nucleotides, circulating tumor nucleic acids, etc. Although there have been many advancements in recent years and received a lot of attention, the currently discovered lung cancer markers generally cannot have both high sensitivity and specificity, and most lung cancer markers have not been externally verified. Repeated blood sampling is required during sampling, causing repeated harm to the subjects.

[0005] Therefore, in view of the above technical problems, it is necessary to provide a lung nodule malignant risk prediction system based on exhaled gas VOCs.

[0006] The information disclosed in this background section is only intended to enhance the overall understanding of the present invention and should not be regarded as an admission or any form of implication that this information constitutes prior art already known to those of ordinary skill in the art. Summary of the Invention

[0007] An object of the present invention is to provide a lung nodule malignant risk prediction system based on exhaled gas VOCs, which can solve the problems that the current prediction of lung nodule malignant risk relies on the ability and experience of doctors, the prediction results show large differences, and the prediction is inaccurate.

[0008] To achieve the above object, the technical solution provided by a specific embodiment of the present invention is as follows: A lung nodule malignant risk prediction system based on exhaled gas VOCs includes: a sample data collection system, a sample processing system, and a risk prediction system. The sample data collection system is used to collect and analyze the exhaled gas VOCs, epidemiological data, imaging signs, and laboratory indicators of patients to form a sample data set. The sample processing system is used to preprocess all features in the sample data set, divide the preprocessed sample data set into a test set and a training set, and screen all features of the patients in the divided test set and training set. The risk prediction system includes a random forest model. The training set after feature screening is used to perform prediction training on the malignant risk of lung nodules through the random forest model, and a lung nodule malignant risk prediction result is obtained.

[0009] In one or more embodiments of the present invention, the collection method of the sample data set is as follows: S1. Preparation before exhaled gas sampling: (1) Fast for 8 - 12 hours; (2) Prohibit smoking for 12 hours; (3) Avoid using antioxidants (β-carotene, Vit E, etc.) within 24 hours before gas collection, and avoid consuming a large amount of fruits, garlic, and high-fat foods; (4) Stop using long-acting bronchodilators 24 hours before sampling or stop inhaling short-acting bronchodilators and glucocorticoids 12 hours before sampling; (5) Avoid strenuous exercise, sit during collection, and breathe calmly; (6) Gargle with pure water before gas collection; (7) Control the collection time basically between 7:30 and 9:00 in the morning; S2. Sample collection: Collect 1L of exhaled gas through a Tedlar gas sampling bag and transfer all the samples to a dedicated adsorption tube for exhaled gas. S3. Sample processing; S4. Obtain all exhaled gas VOCs in the sampling sample, and combine the epidemiological data, imaging signs, and laboratory indicators of all patients in the sampling sample to form a sample data set.

[0010] In one or more embodiments of the present invention, the S3 specifically includes the following steps: S301. Thermally desorb the sample in the adsorption tube using an autosampler (equipment model: TD 100-xr™). The desorption conditions for the adsorption tube are: 300 °C (10 min); the desorption flow rate is: 1 mL / min; the enrichment focusing cold trap: 'Materialemissions' (product number: U-T12ME-2S), and the injection split ratio is: 5 mL / min; S302. Analyze the exhaled gas sample using a comprehensive two-dimensional gas chromatography-mass spectrometer (equipment model: Bench TOF-Select™), with a mass range of: m / z 29 - 350; S303. Data processing and analysis: Perform compound identification and data processing through ChromSpace software.

[0011] In one or more embodiments of the present invention, in the sample processing system, the method for preprocessing all features in the sample data set is to perform multiple imputations of missing values, centering, standardization, and equalization processing.

[0012] In one or more embodiments of the present invention, in the sample processing system, the NIST Library is used for compound identification and screening, and univariate feature selection algorithms, LASSO, and logistic regression are used for feature screening of epidemiological data, imaging signs, and laboratory index features.

[0013] In one or more embodiments of the present invention, the specific operating steps of the risk prediction system are as follows: First, select the exhaled gas VOCs features after screening and use the LASSO algorithm to build a model and perform cross-validation; then, select the epidemiological data, imaging signs, and laboratory index features after feature screening, combine them with the exhaled gas VOCs features, establish a random forest model, calculate the predicted values and draw an ROC curve, and use the optimal critical value of this ROC curve as the threshold for high and low risk classification; finally, draw a calibration curve based on the established random forest model.

[0014] In one or more embodiments of the present invention, the discrimination method for high and low risk classification is: if the predicted value is greater than the optimal critical value of the ROC curve, it is determined to be high risk; if the predicted value is less than or equal to the optimal critical value of the ROC curve, it is determined to be low risk.

[0015] In one or more embodiments of the present invention, the risk prediction system further includes bringing the test set after feature screening into the random forest model for evaluating the lung nodule malignancy risk prediction model.

[0016] In one or more embodiments of the present invention, the method for evaluating the lung nodule malignant risk prediction model using the test set through a random forest model is as follows: bringing the test set into the random forest model to calculate the prediction results, drawing an ROC curve according to the prediction results and the high and low risk classifications marked in the original test set, and calculating the AUC values of the training set and the test set respectively. Finally, draw a clinical decision curve according to the established random forest model.

[0017] In one or more embodiments of the present invention, if the AUC values of both the training set and the test set are greater than 0.9, it proves that the prediction results of the model are ideal and the model can be used.

[0018] Compared with the prior art, a lung nodule malignant risk prediction system based on exhaled gas VOCs of the present invention is based on a random forest, based on the characteristics of exhaled gas VOCs, and combines epidemiological data, imaging signs and laboratory index characteristics to propose a lung nodule malignant risk prediction system based on exhaled gas VOCs, and finally can automatically predict the risk level of patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 It is a flowchart of the lung nodule malignant risk prediction system in an embodiment of the present invention; Figure 2 It is the variable importance ranking of the lung nodule malignant risk prediction system in an embodiment of the present invention; Figure 3 It is an ROC curve graph of the lung nodule malignant risk prediction system in an embodiment of the present invention; Figure 4 It is a calibration curve of the lung nodule malignant risk prediction system in an embodiment of the present invention; Figure 5 It is a clinical decision curve of the lung nodule malignant risk prediction system in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] To enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0022] As Figure 1 shown, a lung nodule malignant risk prediction system based on exhaled gas VOCs in an embodiment of the present invention includes a sample data collection system, a sample processing system, and a risk prediction system. It mainly utilizes the characteristics of exhaled gas VOCs of patients, combines epidemiological data, imaging signs, and laboratory index characteristics, and uses the random forest algorithm to construct a lung nodule malignant risk prediction model to achieve automatic risk prediction for auxiliary diagnosis.

[0023] Among them, the sample data collection system is used to collect and analyze the exhaled gas VOCs, epidemiological data, imaging signs, and laboratory indicators of patients to form a sample data set.

[0024] As Figure 1 shown, the collection method of the sample data set is as follows: S1. Preparation before exhaled gas sampling: (1) Fast for 8 - 12 hours; (2) Prohibit smoking for 12 hours; (3) Avoid using antioxidants (β-carotene, Vit E, etc.) within 24 hours before gas collection, and avoid consuming a large amount of fruits, garlic, and high-fat foods; (4) Stop using long-acting bronchodilators 24 hours before sampling or stop inhaling short-acting bronchodilators and glucocorticoids 12 hours before sampling; (5) Avoid strenuous exercise, take a sitting position during collection, and breathe calmly; (6) Gargle with pure water before gas collection; (7) The collection time is basically controlled between 7:30 and 9:00 in the morning; S2. Sample collection: Collect 1L of exhaled gas through a Tedlar gas sampling bag and transfer all the samples to a dedicated adsorption tube for exhaled gas; S3. Sample processing; S4. Obtain all exhaled gas VOCs in the sampling sample, and combine the epidemiological data, imaging signs, and laboratory indicators of all patients in the sampling sample to form a sample data set.

[0025] In addition, S3 specifically includes the following steps: S301. Thermally desorb the sample in the adsorption tube using an autosampler (equipment model: TD 100-xr™). The desorption conditions for the adsorption tube are: 300 °C (10 min); desorption flow rate: 1 mL / min; enrichment and focusing cold trap: 'Materialemissions' (product number: U-T12ME-2S), injection split: 5 mL / min; S302. Analyze the exhaled gas sample using a comprehensive two-dimensional gas chromatography-mass spectrometer (equipment model: Bench TOF-Select™), with a mass range of: m / z 29 - 350; S303. Data processing and analysis: Perform compound identification and data processing using ChromSpace software.

[0026] In the sample processing system, the method for preprocessing all features in the sample data set is to perform multiple imputations of missing values, centering, standardization, and equalization.

[0027] For example, in the sample data set, if there is a large difference in the number of samples between high-risk and low-risk patients with lung nodules, equalization processing is required. In machine learning algorithms, if the proportion of data samples of a certain class is too small, it may lead to low model training efficiency and poor prediction performance. Therefore, it is necessary to expand the sample data set to equalize the data samples of different classes.

[0028] The sample processing system is used to preprocess all features in the sample data set, divide the preprocessed sample data set into a test set and a training set, and screen all features of the patients in the divided test set and training set. After dividing the data set, it should be ensured that both the training set and the test set contain all types of lung nodule malignant risk level samples.

[0029] In addition, when selecting features, relevant features are screened, irrelevant and redundant features are removed to achieve the purpose of alleviating the curse of dimensionality, reducing the difficulty of the learning task, and improving the efficiency of the model. According to the feature data type, feature selection is performed on exhaled gas VOCs, imaging signs, laboratory index features (mainly continuous values), and epidemiological data features (mainly discrete values) respectively.

[0030] Specifically, there are 196 characteristics of exhaled breath VOCs, including: (-)-beta-Bourbonene, (+)-3-Carene, 1(3H)-Isobenzofuranone, 10-Undecen-1-ol, 1-Butanol, 1-Decanol, 1-Dodecanol, 1-Dodecene, 1-Heptanol, 1-Hexadecanol, 1-Naphthalenol, 1-Nonanol, 1-Nonen-3-ol, 1-Nonene, 1-Octanol, 1-Octen-3-ol, 1-Octene, 1-Pentanol, 1-Propanol, 1-Tetradecanol, 1-Tetradecene, 1-Tridecene, 1-Undecanol, 1-Undecene, 2(5H)-Furanone, 2-Butanone, 2-Butenal, 2-Naphthalenol, 2-n-Butyl furan, 2-Nonadecanol, 2-Pentanol, 2-Pentanone, 2-Propenenitrile, 2-Propenoic acid, 2-Thiophenecarboxaldehyde, 2-Undecanol, 3-Carene, 3-Octanone, 6-Tridecene, 9-Phenanthrenol, Acetaldehyde, Acetamide, Acetic acid, Acetone, Acetone cyanohydrin, Acetonitrile, Acetophenone, Alpha-Methylstyrene, alpha-Phellandrene, alpha-Terpineol, Anethole, Anisole, Argon, Benzaldehyde, Benzene, Benzeneacetaldehyde, Benzofuran, Benzoic acid, Benzonitrile, Benzophenone, Benzothiazole, Benzoylformic acid, Benzyl alcohol, Benzyl chloride, beta-Myrcene, beta-Ocimene, beta-Phellandrene, Beta-Pinene, Biphenyl, Butanal, Butanoic acid, ButylatedHydroxytoluene, Camphene, Carbondioxide, Carbon disulfide, Carbon Tetrachloride, Carveol, Carvone, Caryophyllene, Caryophyllene oxide, Cedrene, cis-3-Decene, cis-3-Hexenyl isovalerate, cis-3-Methylcyclohexanol, cis-Hept-4-enol, Citronellal, Citronellol, Copaene, Crotonicacid, Cyclohexane, Cyclotetradecane, D-Carvone, Decanal, Decane, Dibutyl phthalate, Diethyl Phthalate, Dihydrocarvyl acetate, Dimethyl sulfide, Dimethyl sulfone, Dimethyl trisulfide, Diphenyl ether, D-Limonene, Dodecanal, Dodecane, Dodecanoicacid, Estragole, Ethanol, Ethyl Acetate, Ethylbenzene, Furan, Furfural, gamma-Terpinene, Glutaraldehyde, Heptadecane, Heptanal, Heptane, Heptanoic acid, Hexadecane, Hexanal, Hexanoic acid, Indane, Isocrotonic acid, Isoflurane, IsopropylAlcohol, Isopulegol acetate, Isoquinoline, Limonene, Linalyl acetate, l-Menthone, Mesitylene, Methacrolein, Methenamine, Methyl 2-furoate, Methyl Alcohol, MethylIsobutyl Ketone, Methyl methacrylate, Methyl salicylate, Methyl vinyl ketone, Methyl-2-thiophene carboxylate, Methylene chloride, Myristoleicacid、Myrtenylacetate、Naphthalene、n-Decanoic acid、Nerolidyl acetate、n-Hexadecanoic acid、n-Hexane、Nonanal、Nonane、Nonanoic acid、n-Propyl acetate、Octadecanoic acid、Octanal、Octane、Octanoic acid、o-Cymene、Oleic Acid、Oxygen、o-Xylene、p-Cresol、p-Cymene、Pentadecane、Pentanal、Pentane、Pentanoic acid、Phenanthrene、Phenol、Phenylglyoxal、Phytol、Piperonal、Propanal、Propanoic acid、Propylene Glycol、Pulegone、p-Xylene、Pyridine、Pyrrole、Quinoline、Styrene、Sulfur dioxide、Tetrachloroethylene、Tetradecane、Tetraethylene glycol、Tetrahydrofuran、Thiophene、Thujone、Thymol、Toluene、Trichloroethylene、Trichloromethane、Trichloronitromethane、Tridecane、Tridecanoic acid、Undecanal、Undecane、Undecanoic acid。

[0031] In addition, there are 103 characteristics in epidemiological data, imaging signs, and laboratory indicators, including gender, age, race, marital status, education level, whether smoking, smoking index, which part the smoke is usually inhaled when smoking, whether there is exposure to second-hand smoke of cohabitants, total person-years of exposure to second-hand smoke of cohabitants in low-risk / medium-risk pulmonary nodule individuals, whether there is exposure to second-hand smoke of colleagues at the workplace, total person-years of exposure to second-hand smoke at the workplace in low-risk / medium-risk pulmonary nodule individuals, cooking frequency, total person-years of cooking in low-risk / medium-risk pulmonary nodule individuals, whether the house is heated in winter, previous housing decoration and occupancy situation, whether drinking alcohol, total person-years of alcohol consumption in low-risk / medium-risk pulmonary nodule individuals, whether drinking tea, total person-years of tea consumption in low-risk / medium-risk pulmonary nodule individuals, total number of tea cups consumed per day per person in low-risk / medium-risk pulmonary nodule individuals, whether drinking coffee, total person-years of coffee consumption in low-risk / medium-risk nodule individuals, type of occupation, type of work in the past 3 months, whether exposed to toxic and harmful substances during the career, whether there are any recent symptoms, whether suffering from any chronic diseases, whether suffering from any lung diseases, whether received radioactive examination in the past year, time of detection of pulmonary nodules (year), nodule diameter, cancer history, family history of tumors, physical activity at work, number of days of attendance per week, daily working hours, commuting mode, whether exercising, whether regularly taking a certain drug or health product, average daily sleep time (including naps) in the past 1 month, whether snoring, what time to go to bed at night, whether insomnia, napping situation, whether three meals are regular, whether having midnight snacks, dietary collocation, frequency of cereal consumption, frequency of fresh vegetable and fruit consumption, frequency of meat consumption, frequency of aquatic product consumption, frequency of milk and dairy product consumption, frequency of egg and its product consumption, frequency of bean and bean product consumption, frequency of nut consumption, frequency of sweet food consumption, frequency of fried food consumption, frequency of pickled and smoked food consumption, indirect bilirubin, GOT / GPT, albumin / globulin ratio, globulin, albumin, total protein, gamma-glutamyl transpeptidase, direct bilirubin, total bilirubin, aspartate aminotransferase, alanine aminotransferase, platelet volume, percentage of monocytes, red blood cell distribution width, percentage of basophils, percentage of eosinophils, percentage of neutrophils, percentage of lymphocytes, basophils, eosinophils, monocytes, lymphocytes, platelet distribution width, mean corpuscular volume, mean corpuscular hemoglobin, hematocrit, red blood cells, hemoglobin, white blood cells, platelets, mean platelet volume, neutrophils, mean corpuscular hemoglobin concentration, fasting blood glucose, AFP, CEA, low-density lipoprotein cholesterol, high-density lipoprotein cholesterol, triglyceride, total cholesterol, CA199, uric acid, creatinine, urea.

[0032] In the sample processing system, the NIST Library is used for compound identification and screening, and univariate feature selection algorithms, LASSO, and logistic regression are used to screen features from epidemiological data, imaging signs, and laboratory indicators. That is, for all exhaled gas VOCs, epidemiological data, imaging signs, and laboratory indicators, the univariate feature selection algorithm is first used for feature screening. 48 features (P < 0.1) are screened out, including 1-Decanol, 2-Butenal, 2-Naphthalenol, Acetaldehyde, Benzoic acid, Dimethyl sulfide, Furan, l-Menthone, Methyl-2-thiophene carboxylate, Naphthalene, Octanal, gender, smoking index, cigarette smoke inhalation site, time of exposure to second-hand smoke of cohabitants (person-years), whether there is exposure to second-hand smoke of colleagues in the workplace, previous housing decoration and occupancy, drinking situation, number of years of drinking (person-years), whether drinking tea, number of years of drinking tea (person-years), daily tea consumption, number of years of drinking coffee (person-years), nodule diameter, family history of cancer, daily working hours, napping situation, cereal consumption frequency, fresh vegetable and fruit consumption frequency, milk and dairy product consumption frequency, egg and its product consumption frequency, sweet food consumption frequency, glutamyl transferase, aspartate aminotransferase, alanine aminotransferase, eosinophils, monocytes, hematocrit, red blood cells, hemoglobin, white blood cells, neutrophils, CEA, high-density lipoprotein cholesterol, triglycerides, total cholesterol, uric acid, creatinine.

[0033] Then, the variables after univariate analysis feature screening are selected to build a model using the LASSO algorithm and perform cross-validation. LASSO can determine the best model and coefficients, and then can further screen the features that are meaningful in the above univariate analysis according to the coefficients of the established model. The screened features include (14): 1-Decanol, 2-Butenal, 2-Naphthalenol, Acetaldehyde, Benzoic acid, Dimethyl sulfide, Furan, l-Menthone, Methyl-2-thiophene carboxylate, Naphthalene, Octanal, smoking index, cigarette smoke inhalation site, nodule diameter.

[0034] The risk prediction system includes a random forest model. The training set after feature screening is used to perform prediction training on the malignant risk of pulmonary nodules through the random forest model, and the prediction result of the malignant risk of pulmonary nodules is obtained.

[0035] The random forest model is used for the prediction training of the malignant risk of pulmonary nodules. The random forest algorithm is a relatively flexible machine learning algorithm that combines the idea of ensemble learning and uses the method of self-sampling to sample the training data, with a certain degree of randomness.

[0036] According to the prediction results of the random forest model, the ROC curve is drawn, and the malignant probability of pulmonary nodules is evaluated. The variable importance ranking in this embodiment is shown in Figure 2 , the ROC curve is shown in Figure 3 , the calibration curve is shown in Figure 4 , the clinical decision curve is shown in Figure 5 .

[0037] The specific operation steps of the risk prediction system are as follows: First, the selected exhaled gas VOCs features are modeled using the LASSO algorithm and cross-validated; then, the selected epidemiological data, imaging signs, and laboratory index features after feature screening are combined with the exhaled gas VOCs features to establish a random forest model, calculate the predicted value, draw the ROC curve, and use the optimal critical value of this ROC curve as the threshold for high and low risk classification; finally, the calibration curve is drawn according to the established random forest model.

[0038] The discrimination method for high and low risk classification is: if the predicted value is greater than the optimal critical value of the ROC curve, it is determined as high risk; if the predicted value is less than or equal to the optimal critical value of the ROC curve, it is determined as low risk.

[0039] The risk prediction system also includes bringing the test set after feature screening into the random forest model to evaluate the malignant risk prediction model of pulmonary nodules.

[0040] The method for evaluating the malignant risk prediction model of pulmonary nodules by the test set through the random forest model is: bringing the test set into the random forest model to calculate the prediction results, drawing the ROC curve according to the prediction results and the high and low risk classifications marked in the original test set, and calculating the AUC values (i.e., the area size under the ROC curve) of the training set and the test set respectively. Finally, the clinical decision curve is drawn according to the established random forest model.

[0041] If the AUC values of both the training set and the test set are greater than 0.9, it proves that the prediction results of the model are ideal and the model can be used.

[0042] In addition, the pulmonary nodule malignant risk prediction system also includes a feedback system. When the risk prediction system predicts that the patient has a relatively high cancer risk, the patient can be notified in a timely manner through the feedback system.

[0043] Specifically, the feedback system includes a diagnosis and treatment system and a recommendation system. The diagnosis and treatment system can carry out systematic treatment according to the predicted conditions of patients, thereby avoiding further increase in the cancer risk degree of patients. The recommendation system can recommend relevant common sense knowledge of life to patients, enabling patients to pay timely attention to the dynamic physical information in life and maintain good health.

[0044] In addition, the feedback system also includes a sharing system. Through the analysis system, patients can share interesting life stories with each other and share their own experiences and feelings, so as to increase the communication among patients.

[0045] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.

[0046] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A pulmonary nodule malignancy risk prediction system based on exhaled gas VOCs, characterized in that: include: A sample data collection system, a sample processing system and a risk prediction system. The sample data collection system is used to collect and analyze the patient's exhaled gas VOCs, epidemiological data, imaging signs and laboratory indicators to form a sample data set. The sample processing system is used to preprocess all features in the sample data set, divide the preprocessed sample data set into a test set and a training set, and screen all features of the patients in the divided test set and training set. The risk prediction system includes a random forest model. The training set after feature screening is trained through the random forest model to predict the risk of malignant lung nodules and obtain the prediction result of the risk of malignant lung nodules.

2. A pulmonary nodule malignancy risk prediction system based on exhaled gas VOCs according to claim 1, characterized in that: The sample data set is collected by: S1. Preparation before sampling exhaled gas: (1) Fast for 8-12 hours; (2) Do not smoke for 12 hours; (3) Avoid the use of antioxidants (β-carotene, Vit E, etc.) and avoid consuming large amounts of fruit, garlic, and high-fat foods within 24 hours before sampling; (4) Stop using long-acting bronchodilators 24 hours before sampling or stop inhaling short-acting bronchodilators and glucocorticoids 12 hours before sampling; (5) Avoid strenuous exercise, sit down during sampling, and breathe calmly; (6) Rinse your mouth with pure water before sampling; (7) The sampling time is basically controlled between 7:30 and 9:00 in the morning; S2. Sample collection: Collect 1L of exhaled gas through a Tedlar gas sampling bag and transfer all samples to a dedicated adsorption tube for exhaled gas; S3, sample processing; S4. Obtain all exhaled gas VOCs in the sampling samples, and combine the epidemiological data, imaging signs, and laboratory indicators of all patients in the sampling samples to form a sample data set.

3. A pulmonary nodule malignancy risk prediction system based on exhaled gas VOCs according to claim 2, characterized in that: The S3 specifically includes the following steps: S301. Perform thermal analysis on the sample in the adsorption tube using an automatic sampler. The analysis conditions of the adsorption tube are: 300°C (10 min); the analysis flow rate is: 1 mL / min; the enrichment focusing cold trap is: 'Material emissions', and the injection split flow is: 5 mL / min; S302, analyzing the exhaled gas sample using a comprehensive two-dimensional gas chromatography-mass spectrometer, with a mass range of m / z 29-350; S303, data processing and analysis: Compound identification and data processing are performed using ChromSpace software.

4. The system for predicting the malignant risk of pulmonary nodules based on exhaled gas VOCs according to claim 1, characterized in that: In the sample processing system, the method for preprocessing all features in the sample data set is to fill in missing values, center, standardize and balance the data multiple times.

5. The system for predicting the malignant risk of pulmonary nodules based on exhaled gas VOCs according to claim 4, characterized in that: In the sample processing system, the NIST Library is used for compound identification and screening, and the univariate feature selection algorithm, LASSO and logistic regression are used for feature screening of epidemiological data, imaging signs, and laboratory index characteristics.

6. The system for predicting the malignant risk of pulmonary nodules based on exhaled gas VOCs according to claim 5, characterized in that: The specific operation steps of the risk prediction system are as follows: first, select the screened exhaled gas VOCs features, use the LASSO algorithm to model, and perform cross-validation; then, select the epidemiological data, imaging signs and laboratory index features after feature screening, combine the exhaled gas VOCs features, establish a random forest model, calculate the prediction value and draw the ROC curve, and use the optimal critical value of the ROC curve as the threshold for high and low risk classification; finally, draw a calibration curve based on the established random forest model.

7. The system for predicting the malignant risk of pulmonary nodules based on exhaled gas VOCs according to claim 6, characterized in that: The method for distinguishing the high and low risk classification is: if the predicted value is greater than the optimal critical value of the ROC curve, it is judged as high risk; if the predicted value is less than or equal to the optimal critical value of the ROC curve, it is judged as low risk.

8. The system for predicting the malignant risk of pulmonary nodules based on exhaled gas VOCs according to claim 1, characterized in that: The risk prediction system also includes bringing the test set after feature screening into the random forest model to evaluate the pulmonary nodule malignancy risk prediction model.

9. The system for predicting the malignant risk of pulmonary nodules based on exhaled gas VOCs according to claim 8, characterized in that: The test set is evaluated by the random forest model for predicting the malignant risk of pulmonary nodules: the test set is brought into the random forest model to calculate the prediction results, the ROC curve is drawn according to the prediction results and the high and low risk classifications marked in the original test set, and the AUC values ​​of the training set and the test set are calculated respectively. Finally, the clinical decision curve is drawn according to the established random forest model.

10. The system for predicting the malignant risk of pulmonary nodules based on exhaled gas VOCs according to claim 9, characterized in that: If the AUC values ​​of the training set and the test set are both greater than 0.9, it proves that the prediction results of the model are ideal and the model can be used.