Metabolic marker for distinguishing lung cancer and non-lung cancer diseases, diagnosis system and application
Through metabolomics, multiple metabolic markers were screened out and machine learning models were constructed, which solved the problem of accurate distinction between lung cancer and non-lung cancer diseases, achieved efficient early diagnosis and individualized treatment, and reduced the risk of invasive examinations.
Patent Information
- Application Number
- CN202510438989.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-18
Smart Images

Figure CN120334543A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of disease diagnosis, and particularly relates to a metabolic marker, a diagnostic system and an application for differentiating lung cancer and non-lung cancer diseases. Background Art
[0002] Lung cancer (LC) is one of the malignant tumors with the highest incidence and mortality rates globally. Currently, clinical diagnosis of lung cancer mainly relies on low-dose computed tomography (LDCT) for screening, but it has problems such as a high false positive rate and overdiagnosis. Common symptoms of lung cancer include cough, blood in sputum, wheezing, chest pain, etc. Since there are no specific symptoms in the early stage of lung cancer, most patients are diagnosed at an advanced stage, resulting in a significant decline in the cure rate and five-year survival rate. The occurrence and development of lung cancer is a complex and long-term process, usually including stages such as chronic inflammation and hyperplasia in the lungs to lung cancer. Non-lung cancer diseases (LNCD) mainly include pulmonary tuberculosis, pneumonia, pulmonary nodules, and pulmonary cysts, etc., and belong to a certain stage in the occurrence process of lung cancer. In the early stage of lung cancer, there are no obvious symptoms, or non-specific symptoms such as cough and chest pain appear, which are often confused with the symptoms of lung diseases such as chronic bronchitis and pneumonia, and are easily overlooked, resulting in a low early diagnosis rate of lung cancer and often being misdiagnosed or missed diagnosed. Existing differential diagnoses mainly rely on invasive examinations (such as percutaneous biopsy), but there are risks of complications and difficulties in sampling micro-nodules (<1 cm).
[0003] As an emerging omics technology, metabolomics focuses on the comprehensive analysis of small molecule metabolites in biological samples and can reflect the physiological and pathological states of organisms in real time. By analyzing the changes in metabolite profiles, metabolomics can reveal unique metabolic characteristics related to diseases. Combining algorithms of artificial intelligence and machine deep learning can screen metabolic markers related to diseases for disease diagnosis and typing. However, there is currently a lack of metabolic markers that can accurately differentiate lung cancer and non-lung cancer diseases, which undoubtedly hinders the timely treatment of diseases. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a metabolic marker for differentiating lung cancer and non-lung cancer diseases, which can achieve the diagnostic purpose of high sensitivity and high specificity, and provide effective help for the early diagnosis and individualized treatment of lung cancer.
[0005] The present invention provides a metabolic marker for differentiating lung cancer from non-lung cancer diseases, including at least the following 5: 4-pyridoxic acid, 5-aminolevulinic acid, choline, hippuric acid, L-carnosine, O-acetyl-L-homoserine, L-alanyl-L-phenylalanine, pyrrole-2-carboxylic acid, tartaric acid, carbamoyl-aspartic acid, 1-methyladenosine, anthranilic acid, phenylalanine, quinolinic acid, gluconic acid, normetanephrine, ribitol, D-glyceric acid, guanidinosuccinic acid, pantothenic acid, and 3-methoxybenzene-1,2-diol.
[0006] Preferably, it includes hippuric acid, tartaric acid, carbamoyl-aspartic acid, anthranilic acid, and phenylalanine.
[0007] Preferably, it includes choline, hippuric acid, L-carnosine, O-acetyl-L-homoserine, tartaric acid, carbamoyl-aspartic acid, 1-methyladenosine, anthranilic acid, phenylalanine, and quinolinic acid.
[0008] Preferably, it includes 4-pyridoxic acid, 5-aminolevulinic acid, choline, hippuric acid, L-carnosine, O-acetyl-L-homoserine, L-alanyl-L-phenylalanine, pyrrole-2-carboxylic acid, tartaric acid, carbamoyl-aspartic acid, 1-methyladenosine, anthranilic acid, phenylalanine, quinolinic acid, and gluconic acid.
[0009] The present invention provides the application of the metabolic marker for differentiating lung cancer from non-lung cancer diseases in the preparation of a diagnostic prediction model for differentiating lung cancer from non-lung cancer diseases.
[0010] Preferably, the non-lung cancer diseases include at least one of the following: pneumonia, pulmonary tuberculosis, and pulmonary infection.
[0011] Preferably, the diagnostic prediction model is constructed based on the machine learning support vector machine algorithm.
[0012] The present invention provides a diagnostic system for differentiating lung cancer from non-lung cancer diseases, which connects the following modules in a communicable manner:
[0013] A data acquisition module, which is used to acquire the sample information of the sample to be tested and the measurement data of the metabolic marker of the sample to be tested;
[0014] A data diagnosis module, which is used to input the measurement data of the data acquisition module into the diagnostic prediction model constructed by the application for analysis to obtain a diagnostic result;
[0015] A data output module, which is used to output the diagnostic result obtained by the data diagnosis module to the terminal for display.
[0016] The present invention provides a diagnostic device for differentiating lung cancer from non-lung cancer diseases, which is equipped with the diagnostic system.
[0017] Preferably, it further includes a metabolic marker determination device for the sample to be tested electrically connected to the data acquisition module in the diagnostic system and / or a terminal display device electrically connected to the data output module in the diagnostic system.
[0018] The present invention provides a metabolic marker for differentiating lung cancer and non-lung cancer diseases, including at least 5 of the following: 4-pyridoxic acid, 5-aminolevulinic acid, choline, hippuric acid, L-carnosine, O-acetyl-L-homoserine, L-alanyl-L-phenylalanine, pyrrole-2-carboxylic acid, tartaric acid, carbamoyl-aspartic acid, 1-methyladenosine, anthranilic acid, phenylalanine, quinolinic acid, gluconic acid, normetanephrine, ribitol, D-glyceric acid, guanidinosuccinic acid, pantothenic acid, and 3-methoxybenzene-1,2-diol. By selecting 5 to 21 metabolic markers from the above metabolic markers and constructing a diagnostic prediction model through machine algorithm learning, the present invention can accurately diagnose early-stage lung cancer, providing effective assistance for the prevention and reduction of the incidence of lung cancer. At the same time, the collection of samples (such as urine, etc.) of the metabolic markers is convenient, with little invasiveness to patients, suitable for long-term monitoring and large-scale population screening, and is widely applicable and popularized. Description of the Drawings
[0019] Figure 1 Results of the multivariate ROC curve analysis of 21 important markers for differentiating lung cancer and non-lung cancer diseases in the modeling group;
[0020] Figure 2 Results of the multivariate ROC curve analysis of 21 important markers for differentiating lung cancer and non-lung cancer diseases in the validation group. Detailed Embodiments
[0021] The present invention provides a metabolic marker for differentiating lung cancer and non-lung cancer diseases, including at least 5 of the following: 4-pyridoxic acid, 5-aminolevulinic acid, choline, hippuric acid, L-carnosine, O-acetyl-L-homoserine, L-alanyl-L-phenylalanine, pyrrole-2-carboxylic acid, tartaric acid, carbamoyl-aspartic acid, 1-methyladenosine, anthranilic acid, phenylalanine, quinolinic acid, gluconic acid, normetanephrine, ribitol, D-glyceric acid, guanidinosuccinic acid, pantothenic acid, and 3-methoxybenzene-1,2-diol.
[0022] In the present invention, when 5 metabolites are selected, it preferably includes a combination of hippuric acid, tartaric acid, carbamoyl-aspartic acid, anthranilic acid, and phenylalanine, or a combination of quinolinic acid, gluconic acid, normetanephrine, ribitol, and D-glyceric acid, or a combination of 4-pyridoxic acid, 5-aminolevulinic acid, choline, hippuric acid, and L-carnosine.
[0023] In the present invention, when 10 metabolites are selected, it is preferably a combination including choline, hippuric acid, L-carnosine, O-acetyl-L-homoserine, tartaric acid, carbamoyl-aspartic acid, 1-methyladenosine, anthranilic acid, phenylalanine, and quinolinic acid, or a combination including 4-pyridoxic acid, 5-aminolevulinic acid, choline, O-acetyl-L-homoserine, L-alanyl-L-phenylalanine, pyrrole-2-carboxylic acid, 1-methyladenosine, anthranilic acid, phenylalanine, and quinolinic acid, or a combination including L-carnosine, O-acetyl-L-homoserine, L-alanyl-L-phenylalanine, pyrrole-2-carboxylic acid, tartaric acid, 1-methyladenosine, anthranilic acid, phenylalanine, quinolinic acid, and gluconic acid.
[0024] In the present invention, when 15 metabolites are selected, it is preferably a combination including 4-pyridoxic acid, 5-aminolevulinic acid, choline, hippuric acid, L-carnosine, O-acetyl-L-homoserine, L-alanyl-L-phenylalanine, pyrrole-2-carboxylic acid, tartaric acid, carbamoyl-aspartic acid, 1-methyladenosine, anthranilic acid, phenylalanine, quinolinic acid, and gluconic acid, or a combination including 4-pyridoxic acid, 5-aminolevulinic acid, choline, hippuric acid, L-carnosine, pyrrole-2-carboxylic acid, tartaric acid, carbamoyl-aspartic acid, 1-methyladenosine, anthranilic acid, gluconic acid, normetanephrine, ribitol, D-glyceric acid, and guanidinosuccinic acid, or a combination including L-alanyl-L-phenylalanine, pyrrole-2-carboxylic acid, tartaric acid, carbamoyl-aspartic acid, 1-methyladenosine, anthranilic acid, phenylalanine, quinolinic acid, gluconic acid, normetanephrine, ribitol, D-glyceric acid, guanidinosuccinic acid, pantothenic acid, and 3-methoxybenzene-1,2-diol.
[0025] In the present invention, the method for screening the metabolic markers preferably includes the following steps:
[0026] Collect the urine of lung cancer patients and non-lung cancer disease patients, detect the metabolites in the urine, and obtain metabolomics data;
[0027] Screen the differential metabolites in the urine samples of lung cancer patients and non-lung cancer disease patients from the metabolomics data, and obtain metabolic markers through screening.
[0028] In the present invention, the non-lung cancer diseases preferably include at least one of the following: pneumonia, pulmonary tuberculosis, and pulmonary infection. The method for detecting metabolites in the urine preferably performs pretreatment on the urine and then uses an aqueous phase polar substance liquid chromatography-mass spectrometry instrument for detection. The metabolomics data further includes a data cleaning and noise reduction process. The metabolomics data further includes the identification and verification of metabolites. The screening criteria for the differential metabolites preferably perform LASSO (Least Absolute Shrinkage and Selection Operator) regression analysis on the data of the modeling group, obtain the optimal regularization parameter alpha = 0.0296 through minimizing the cross-validation error, and screen the metabolites according to non-zero coefficients.
[0029] The present invention provides the use of the metabolic markers for differentiating lung cancer and non-lung cancer diseases in the preparation of a diagnostic prediction model for differentiating lung cancer and non-lung cancer diseases.
[0030] In the present invention, the non-lung cancer diseases preferably include at least one of the following: pneumonia, pulmonary tuberculosis, and pulmonary infection.
[0031] In the present invention, the diagnostic prediction model is preferably constructed based on the machine learning support vector machine algorithm. The diagnostic prediction model preferably uses the machine learning support vector machine algorithm to randomly iterate 1000 times.
[0032] In the present invention, when constructing a diagnostic prediction model using the above 21 metabolic markers, the diagnostic prediction model has good sensitivity and specificity, with an AUC value of 0.895, a sensitivity of 0.841, and a specificity of 0.821. When constructing a diagnostic prediction model using the above 15 metabolic markers, the diagnostic prediction model has good sensitivity and specificity, with an AUC value of 0.887, a sensitivity of 0.864, and a specificity of 0.718. When constructing a diagnostic prediction model using the above 10 metabolic markers, the diagnostic prediction model has good sensitivity and specificity, with an AUC value of 0.818, a sensitivity of 0.750, and a specificity of 0.825. When constructing a diagnostic prediction model using the above 5 metabolic markers, the diagnostic prediction model has good sensitivity and specificity, with an AUC value of 0.814, a sensitivity of 0.773, and a specificity of 0.718.
[0033] The present invention provides a diagnostic system for differentiating lung cancer and non-lung cancer diseases, which connects the following modules in a communicable manner:
[0034] A data acquisition module, which is used to acquire the sample information of the sample to be tested and the measurement data of the metabolic markers of the sample to be tested;
[0035] A data diagnosis module, configured to input the measurement data of the data acquisition module into a diagnosis prediction model constructed by the application for analysis to obtain a diagnosis result;
[0036] A data output module, configured to output the diagnosis result obtained by the data diagnosis module to a terminal for display.
[0037] The present invention provides a diagnosis device for distinguishing between lung cancer and non-lung cancer diseases, equipped with the diagnosis system.
[0038] In the present invention, preferably further included are a metabolic marker determination device for a test sample electrically connected to the data acquisition module in the diagnosis system and / or a terminal display device for the data output module in the diagnosis system. The metabolic marker determination device preferably includes a liquid chromatography-mass spectrometry instrument. The terminal display device preferably includes a computer.
[0039] The following describes in detail a metabolic marker, a diagnosis system, and an application for distinguishing between lung cancer and non-lung cancer diseases provided by the present invention in conjunction with embodiments, but they should not be construed as limiting the protection scope of the present invention.
[0040] Example 1
[0041] Screening of a metabolic marker for distinguishing between lung cancer and non-lung cancer diseases and construction of a diagnosis prediction model
[0042] 1. Subject conditions
[0043] The inclusion criteria and exclusion criteria for non-lung cancer disease patients and lung cancer patients are as follows:
[0044] 1) Inclusion criteria:
[0045] Subjects must meet all of the following inclusion criteria to be eligible to participate in this example:
[0046] (1) Males or females aged ≥ 18 years old; (2) Read and fully understand, sign the informed consent form, and be able to provide urine samples for metabolomics detection; (3) Lung cancer patients are all diagnosed by the gold standard method of tissue pathological examination; (4) Confirmed by biopsy / postoperative pathology or clinically diagnosed by a clinician through comprehensive evaluation, including patients with pulmonary benign diseases such as pneumonia, pulmonary tuberculosis, and pulmonary infections.
[0047] 2) Exclusion criteria:
[0048] Subjects who meet any of the following exclusion criteria are ineligible to participate in this example:
[0049] (1) Pregnancy or lactation; (2) Emergency or need for rescue; (3) History of malignant tumors or any anti-tumor treatment before sampling; (4) Patients with multiple primary malignant tumors at the same time.
[0050] 3) Subject Information
[0051] In this example, urine samples from 660 subjects were collected through two medical centers, including 200 samples of Lung non-cancer disease (LNCD) and 460 samples of Lung cancer (LC). Specifically, for the modeling group, there were 157 people in the LNCD group and 352 people in the LC group; for the validation group, there were 43 people in the LNCD group and 108 people in the LC group (Table 1).
[0052] Table 1 Subject Information
[0053] Non-lung cancer disease group (LNCD) Lung cancer group (LC) Number of people in the modeling group 157 352 Number of people in the validation group 43 108 Total 200 460
[0054] 2. Urine Metabolite Detection
[0055] 1) Detection Reagents:
[0056] Methanol, acetonitrile, water, acetic acid, and methyl tert-butyl ether of mass spectrometry grade purity, and formic acid of high-performance liquid chromatography (HPLC) grade purity were all purchased from Sigma-Aldrich, USA.
[0057] 2) Sample Preparation:
[0058] After thawing the urine samples taken out from the -80°C refrigerator, 40 μL of urine samples were taken into extraction tubes, and 400 μL of pre-cooled extraction solution of methyl tert-butyl ether and methanol (methyl tert-butyl ether: methanol, volume ratio 3:1) was added. After vortexing and ultrasonic mixing, 360 μL of methanol-water mixed solution was added, and then vortexed and centrifuged for stratification; 300 μL of the lower aqueous phase solution in the extraction tube was taken, 900 μL of pre-cooled methanol solution was added, and after protein precipitation, 1000 μL of the supernatant was dried by rotation, and 200 μL of water was added for reconstitution. After reconstitution, it was used for LC-MS detection of water-soluble polar substances.
[0059] 3) Detection of Water-Soluble Small Molecule Metabolites:
[0060] Use Waters ACQUTTY HSS T3 1.8μm 2.1*100mm column for small molecule separation; both the liquid chromatography and mass spectrometer used the ACQUITY UPLC I-Class liquid chromatography system (Waters) and the Q-Exactive mass spectrometry system (Thermo Fisher Scientific).
[0061] The mobile phase parameters are as follows:
[0062] Mobile phase A is an aqueous solution containing 0.1% formic acid; mobile phase B is an acetonitrile solution containing 0.1% formic acid. The separation and elution gradient is as follows: 1% - 70% mobile phase B from 0 to 13 minutes, and 99% mobile phase B from 13 to 18 minutes;
[0063] The mass spectrometry parameters are as follows:
[0064] The mass spectrometry data was collected in the Full MS and Full MS / dd-MS2 modes (both positive and negative modes), and the parameters used for QExactive are as follows: In the Full MS mode, the resolution is 70,000, the scanning range is 100 - 1500 m / z, the Automatic Gain Control (AGC) is 3E+6, and the Maximum IT is 200 milliseconds; in the Full MS / dd-MS2 mode, the resolution of the second-stage mass spectrometry is 17,500, the quadrupole window is 1.5 m / z, the AGC is 1E+5, the maximum ion injection time is 50 ms, and the relative collision energy (Higher Energy Collisional Dissociation, HCD) is 30 eV.
[0065] 3. Metabolomics data preprocessing and metabolite identification
[0066] 1) Metabolomics data processing:
[0067] (1) The RAW format files obtained from the mass spectrometry were subjected to peak extraction to obtain FeatureXML format files, reducing the dimension of the original mass spectrometry data and improving the signal-to-noise ratio;
[0068] (2) Using the peak alignment algorithm of the OpenMS software, the retention times of the extracted peak format data were corrected and aligned among samples, thereby converting the mass spectrometry data into a data matrix;
[0069] (3) The isotope peaks in the data matrix obtained in step 2 were matched and filtered, and abnormal data (0, negative values, background noise, etc.) were replaced with missing values;
[0070] (4) Among all the characteristic peaks obtained in step 3, those with a detection rate < 80% were excluded, and the characteristic peaks with a detection rate > 80% were filled with the median value of the characteristic peak and added with 5% random noise (following a standard normal distribution);
[0071] (5) In order to reduce the differences in metabolite concentrations among samples and make the data distribution more symmetrical, NormalizationAutoencoder (NormAE) was used for normalization to remove systematic errors such as batch effects.
[0072] 2) Identification of metabolites:
[0073] Based on the spectral information of the compound's primary parent ions (MS1) and secondary fragment ions (MS2) obtained after software parsing of the original data, such as the mass-to-charge ratio (m / z) of the primary mass spectrum and the fragments of the secondary ions, the spectral information of the primary and secondary metabolites in the public database is matched to qualitatively analyze the metabolites. Commonly used metabolite databases include the Human Metabolome Database (HMDB, www.hmdb.ca), the Metabolomics Database (Metlin, metlin.scripps.edu), and the Mass Spectrometry Database (www.massbank.jp); for the metabolites identified based on the relevant databases, the retention time, MS1, and MS2 mass spectrometry information during separation under the same chromatographic column and mass spectrometry conditions as the standard samples are used to finally verify the metabolites. The criteria for metabolite identification are that the retention time difference is within 0.1 min, and the difference between the theoretical value and the measured value of the metabolite molecular weight is less than 10 ppm.
[0074] 4. Data analysis
[0075] 1) Biomarker screening
[0076] After metabolite detection and annotation of the above samples, a total of 360 metabolites were obtained. For the data of the modeling group, LASSO (Least Absolute Shrinkage and Selection Operator) regression analysis was performed. The average error corresponding to each regularization parameter alpha was calculated using 5-fold cross-validation, and the optimal alpha with the minimum error was obtained as 0.0296. The metabolites with non-zero regression coefficients in the corresponding model were retained, and finally a total of 21 differential metabolites (Table 2) were screened out as important metabolic biomarkers for distinguishing the LNCD group from the LC group.
[0077] Table 2 21 important metabolic biomarkers for distinguishing LC from LNCD
[0078]
[0079]
[0080] 5. Construction of a diagnostic model for distinguishing non-lung cancer diseases from lung cancer
[0081] To verify the diagnostic efficacy of these 21 selected markers in differentiating LNCD vs LC, multivariate ROC curve analysis was performed on these 21 markers in the modeling group. Randomly, 3 / 4 of the sample data of LNCD and LC groups in the modeling group were used as the training set, and 1 / 4 as the test set for learning. Using the machine learning support vector machine (SVM), it was randomly iterated 1000 times. By statistically averaging the accuracy of the final model, a diagnostic model for differentiating non-lung cancer diseases from lung cancer was constructed.
[0082] The ROC curve is a method for studying the relationship between the sensitivity and specificity of a model. With sensitivity as the ordinate and 1 - specificity as the abscissa, the evaluation is based on comparing the area under the curve AUC. When AUC is greater than 0.5, the closer AUC is to 1, the better the model performance and the better the diagnostic efficacy. If it is less than 0.5, it indicates poor model accuracy. The ROC classification prediction model includes not only common parameters such as the receiver operating characteristic curve (ROC) and the area under the curve (AUC), but also sensitivity and specificity.
[0083] The calculation formula for sensitivity is shown in Formula I:
[0084]
[0085] The calculation formula for specificity is shown in Formula II:
[0086]
[0087] Among them, TP (True Positive): True positive, the number of samples that are actually positive cases and are correctly predicted as positive cases;
[0088] TN (Ture Negative): True negative, the number of samples that are actually negative cases and are correctly predicted as negative cases;
[0089] FP (False Positive): False positive, the number of samples that are actually negative cases and are wrongly predicted as positive cases;
[0090] FN (False Negative): False negative, the number of samples that are actually positive cases and are wrongly predicted as negative cases;
[0091] The results are as Figure 1As shown, AUC = 0.895 (sensitivity = 0.841, specificity = 0.821), and the results indicate that the constructed model has high diagnostic efficacy.
[0092] Example 2
[0093] Verification of the diagnostic model for differentiating non-lung cancer diseases from the lung cancer group
[0094] To further verify the diagnostic model for differentiating non-lung cancer diseases from the lung cancer group constructed based on the data of the modeling group, the above model was verified using the data of the verification group, and multivariate ROC curve analysis was performed to evaluate the independent verification effect of the model on the unknown data set outside the modeling group data set.
[0095] After putting the samples of the verification group into the model constructed by the modeling group, according to the detection data of each sample for 21 important metabolic markers differentiating LC from LNCD, the corresponding probability value (Probability) will be output. Taking the probability value of each sample as the diagnostic threshold, a set of confusion matrices (including true positives, true negatives, false positives, and false negatives) can be obtained. According to Formula I and Formula II, the sensitivity and specificity can be calculated. In the ROC analysis graph with sensitivity (Sensitivity) as the vertical coordinate and 1 - specificity (1 - Specificity) as the horizontal coordinate, a point can be marked. Similarly, when the probability value of each sample is used as the diagnostic threshold, we will obtain multiple different points in the ROC analysis graph. Connecting these points will draw the ROC curve ( Figure 2 ). Among them, the point composed of the best-performing sensitivity and specificity is selected, and the diagnostic threshold at this time is 0.6957.
[0096] As shown in the confusion matrix results in Table 3, in the prediction model constructed based on the above 21 metabolic markers, with 0.6957 as the diagnostic threshold, among 108 lung cancer patients, 92 were judged to have lung cancer, and 16 were misjudged as the non-lung cancer disease group; among 43 non-lung cancer disease patients, 35 were correctly judged, and 8 were misjudged as the lung cancer group; the ROC analysis results of the diagnostic model in the verification group are as Figure 2 shown. Based on the results of the confusion matrix, the sensitivity and specificity were calculated, and the results showed that AUC = 0.908 (sensitivity = 0.852, specificity = 0.814). The above results indicate that the established diagnostic model for differentiating non-lung cancer diseases from the lung cancer group also has good diagnostic effects in the verification group.
[0097] Table 3 Confusion matrix of the diagnostic model for differentiating non-lung cancer diseases from lung cancer
[0098] Disease type Lung cancer Non-lung cancer disease Lung cancer, N = 108 92 (TP) 16 (FN) Non-lung cancer disease, N = 43 8 (FP) 35 (TN)
[0099] Example 3
[0100] For Example 1, different combinations of metabolic markers were screened to construct a diagnostic model according to the method of Example 1. Fifteen metabolic markers were used, specifically including 4-pyridoxic acid, 5-aminolevulinic acid, choline, hippuric acid, L-carnosine, O-acetyl-L-homoserine, L-alanyl-L-phenylalanine, pyrrole-2-carboxylic acid, tartaric acid, carbamoyl-aspartic acid, 1-methyladenosine, anthranilic acid, phenylalanine, quinolinic acid, and gluconic acid. ROC curve analysis was performed in the modeling group.
[0101] The results showed that when using 15 metabolic markers, AUC = 0.887 (sensitivity = 0.864, specificity = 0.718), which has clinical diagnostic significance.
[0102] Verification was carried out in the validation group, and multivariate ROC curve analysis was performed. The results showed that for the above 15 metabolic markers (4-pyridoxic acid, 5-aminolevulinic acid, choline, hippuric acid, L-carnosine, O-acetyl-L-homoserine, L-alanyl-L-phenylalanine, pyrrole-2-carboxylic acid, tartaric acid, carbamoyl-aspartic acid, 1-methyladenosine, anthranilic acid, phenylalanine, quinolinic acid, and gluconic acid), AUC = 0.868 (sensitivity = 0.824, specificity = 0.791) in the validation group.
[0103] Example 4
[0104] For Example 1, different combinations of metabolic markers were screened to construct a diagnostic model according to the method of Example 1. Ten metabolic markers were used, specifically including choline, hippuric acid, L-carnosine, O-acetyl-L-homoserine, tartaric acid, carbamoyl-aspartic acid, 1-methyladenosine, anthranilic acid, phenylalanine, and quinolinic acid. ROC curve analysis was performed in the modeling group.
[0105] When using 10 metabolic markers, AUC = 0.818 (sensitivity = 0.750, specificity = 0.825), which has clinical diagnostic significance.
[0106] Verification was carried out in the validation group, and multivariate ROC curve analysis was performed. The results showed that for the 10 metabolic markers (choline, hippuric acid, L-carnosine, O-acetyl-L-homoserine, tartaric acid, carbamoyl-aspartic acid, 1-methyladenosine, anthranilic acid, phenylalanine, and quinolinic acid), AUC = 0.888 (sensitivity = 0.769, specificity = 0.837) in the validation group.
[0107] Example 5
[0108] The diagnostic model was constructed according to the method of Example 1 using different combinations of metabolic markers screened in Example 1. Five metabolic markers were used, specifically including hippuric acid, tartaric acid, carbamoyl-aspartic acid, anthranilic acid, and phenylalanine. ROC curve analysis was performed in the modeling group.
[0109] When using five metabolic markers, AUC = 0.814 (sensitivity = 0.773, specificity = 0.718), which has clinical diagnostic significance.
[0110] Verification was carried out in the validation group, and multivariate ROC curve analysis was performed. The results showed that when using five metabolic markers (hippuric acid, tartaric acid, carbamoyl-aspartic acid, anthranilic acid, and phenylalanine), AUC = 0.842 (sensitivity = 0.741, specificity = 0.721) in the validation group.
[0111] The results of the above examples show that the prediction model established by the present invention also has good diagnostic effects in the validation group.
[0112] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A metabolic marker for differentiating lung cancer from non-lung cancer diseases, characterized in that, including at least the following 5: 4-pyridoxic acid, 5-aminolevulinic acid, choline, hippuric acid, L-carnosine, O-acetyl-L-homoserine, L-alanyl-L-phenylalanine, pyrrole-2-carboxylic acid, tartaric acid, carbamoyl-aspartic acid, 1-methyladenosine, anthranilic acid, phenylalanine, quinolinic acid, gluconic acid, normetanephrine, ribitol, D-glyceric acid, guanidinosuccinic acid, pantothenic acid, and 3-methoxybenzene-1,2-diol.
2. The metabolic marker for differentiating lung cancer and non-lung cancer diseases according to claim 1, wherein including hippuric acid, tartaric acid, carbamoyl-aspartic acid, anthranilic acid, and phenylalanine.
3. The metabolic marker for differentiating lung cancer and non-lung cancer diseases according to claim 1, wherein including choline, hippuric acid, L-carnosine, O-acetyl-L-homoserine, tartaric acid, carbamoyl-aspartic acid, 1-methyladenosine, anthranilic acid, phenylalanine, and quinolinic acid.
4. The metabolic marker for differentiating lung cancer from non-lung cancer diseases according to claim 1, wherein including 4-pyridoxic acid, 5-aminolevulinic acid, choline, hippuric acid, L-carnosine, O-acetyl-L-homoserine, L-alanyl-L-phenylalanine, pyrrole-2-carboxylic acid, tartaric acid, carbamoyl-aspartic acid, 1-methyladenosine, anthranilic acid, phenylalanine, quinolinic acid, and gluconic acid.
5. Use of the metabolic marker for differentiating lung cancer and non-lung cancer diseases according to claim 1 in preparing a diagnostic prediction model for differentiating lung cancer and non-lung cancer diseases.
6. The application according to claim 5, characterized in that The non-lung cancer diseases include at least one of the following: pneumonia, pulmonary tuberculosis, and pulmonary infection.
7. The application according to claim 5, wherein The diagnostic prediction model is constructed based on the machine learning support vector machine algorithm.
8. A diagnostic system for differentiating lung cancer from non-lung cancer diseases, characterized in that, Connect the following modules in a communicable manner: A data acquisition module, configured to acquire sample information of a sample to be tested and measurement data of the metabolic marker according to claim 1 of the sample to be tested. A data diagnosis module, configured to input the measurement data of the data acquisition module into the diagnostic prediction model constructed by the application according to any one of claims 5 to 7 for analysis to obtain a diagnosis result. A data output module, configured to output the diagnosis result obtained by the data diagnosis module to a terminal for display.
9. A diagnostic device for differentiating lung cancer and non-lung cancer diseases, characterized in that, Equipped with the diagnostic system according to claim 8.
10. The diagnostic device for differentiating lung cancer and non-lung cancer diseases according to claim 9, characterized in that, It further includes a metabolic marker measuring device of the sample to be tested electrically connected to the data acquisition module in the diagnostic system and / or a terminal display device of the data output module in the diagnostic system.