Blood marker for distinguishing lung cancer and non-lung cancer diseases and application thereof
By screening blood metabolites through metabolomics and building a machine learning model, the problem of accurately distinguishing lung cancer and non-lung cancer diseases has been solved, the sensitivity and specificity of early diagnosis have been improved, and the diagnostic process has been simplified.
Patent Information
- Application Number
- CN202510900785.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies have difficulty accurately distinguishing between lung cancer and non-lung cancer diseases, resulting in low early diagnosis rates, reliance on invasive examinations with the risk of complications, and difficulty in sampling tiny nodules.
Metabolomics technology was used to screen out a variety of blood metabolites, such as phosphatidylethanolamine, lysophosphatidylcholine, sphingomyelin, dodecanedioic acid, etc., and combined with machine learning support vector machine to construct a diagnostic model for distinguishing lung cancer and non-lung cancer diseases.
It achieves high sensitivity and high specificity in lung cancer diagnosis, simplifies experimental operations, and improves the accuracy and reliability of early diagnosis.
Smart Images

Figure CN120741856A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of biomarkers, and in particular relates to a blood marker for distinguishing lung cancer from non-lung cancer diseases and its application. Background Art
[0002] Lung cancer (LC) is one of the malignant tumors with the highest morbidity and mortality rates worldwide. Currently, clinical lung cancer diagnosis mainly relies on low-dose computed tomography (LDCT) screening, but it has problems such as high false-positive rates and overdiagnosis. The development and progression of lung cancer is a complex and long-term process, usually including stages such as chronic lung inflammation and hyperplasia to lung cancer. Non-lung cancer diseases (Lung non-cancer disease, LNCD) mainly include tuberculosis, pneumonia, lung nodules and lung cysts, which belong to a certain stage in the development of lung cancer. In the early stages of lung cancer, there are no obvious symptoms, or non-specific symptoms such as cough and chest pain may occur. They are often confused with the symptoms of lung diseases such as chronic bronchitis and pneumonia, and are easily overlooked, resulting in a low early diagnosis rate of lung cancer and frequent misdiagnosis or missed diagnosis. Existing differential diagnosis mainly relies on invasive examinations (such as puncture biopsy), but there is a risk of complications and it is difficult to sample small nodules (<1 cm).
[0003] Metabolomics, an emerging omics technology, focuses on the comprehensive analysis of small molecule metabolites in biological samples, providing real-time insights into the physiological and pathological states of organisms. By analyzing changes in metabolite profiles, metabolomics can reveal unique metabolic signatures associated with diseases. Combined with artificial intelligence and deep learning algorithms, metabolomics can screen for disease-related metabolic markers for diagnosis and classification. However, the current lack of metabolic markers that can accurately distinguish between lung cancer and non-lung cancer diseases undoubtedly hinders timely treatment. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a blood marker for distinguishing lung cancer from non-lung cancer diseases, which has high detection sensitivity and specificity and can be used for accurate screening of lung cancer.
[0005] The present invention provides a blood marker for distinguishing lung cancer from non-lung cancer diseases, comprising at least five of the following metabolites: phosphatidylethanolamine (36:4), lysophosphatidylcholine 18:1e, phosphatidylethanolamine 38:7e, sphingomyelin d38:5, dodecanedioic acid, N-acetylputrescine, L-glutamyl-L-glutamine, hippuric acid, trans-cyclohexane-1,2-dicarboxylic acid, glycochenodeoxycholic acid, allantoic acid, L-valine, methylleucine and 4-guanidinobutyric acid.
[0006] Preferably, when the blood markers include 5 types, the following metabolites are included: a first blood marker combination consisting of lysophosphatidylcholine 18:1e, phosphatidylethanolamine 38:7e, N-acetylputrescine, L-glutamyl-L-glutamine and hippuric acid; a second blood marker combination consisting of phosphatidylethanolamine (36:4), phosphatidylethanolamine 38:7e, dodecanedioic acid, trans-cyclohexane-1,2-dicarboxylic acid and glycochenodeoxycholic acid; a second blood marker combination consisting of phosphatidylethanolamine (3 6:4), a third blood marker combination consisting of lysophosphatidylcholine 18:1e, sphingomyelin d38:5, dodecanedioic acid and N-acetylputrescine, a fourth blood marker combination consisting of phosphatidylethanolamine 38:7e, L-glutamyl-L-glutamine, hippuric acid, glycochenodeoxycholic acid and allantoic acid, and a fifth blood marker combination consisting of lysophosphatidylcholine 18:1e, hippuric acid, trans-cyclohexane-1,2-dicarboxylic acid, glycochenodeoxycholic acid and methylleucine.
[0007] Preferably, when the blood markers include 10 types, the following metabolites are included: a sixth blood marker combination consisting of lysophosphatidylcholine 18:1e, phosphatidylethanolamine 38:7e, sphingomyelin d38:5, N-acetylputrescine, L-glutamyl-L-glutamine, hippuric acid, trans-cyclohexane-1,2-dicarboxylic acid, glycochenodeoxycholic acid, L-valine and 4-guanidinobutyric acid; a sixth blood marker combination consisting of phosphatidylethanolamine (36:4), lysophosphatidylcholine 18:1e, sphingomyelin d38:5, N-acetylputrescine, L-glutamyl-L-glutamine, hippuric acid, trans-cyclohexane-1,2-dicarboxylic acid, glycochenodeoxycholic acid, L-valine and 4-guanidinobutyric acid; The seventh blood marker combination consists of phosphatidylethanolamine 38:7e, sphingomyelin d38:5, dodecanedioic acid, N-acetylputrescine, L-glutamyl-L-glutamine, hippuric acid, trans-cyclohexane-1,2-dicarboxylic acid, glycochenodeoxycholic acid and L-valine, and the eighth blood marker combination consists of phosphatidylethanolamine 38:7e, sphingomyelin d38:5, dodecanedioic acid, N-acetylputrescine, trans-cyclohexane-1,2-dicarboxylic acid, glycochenodeoxycholic acid, allantoic acid, L-valine, methylleucine and 4-guanidinobutyric acid.
[0008] The present invention provides the use of the blood marker in constructing a diagnostic model for distinguishing non-lung cancer diseases from lung cancer.
[0009] Preferably, the diagnostic model is constructed using a machine learning support vector machine.
[0010] Preferably, the non-lung cancer disease includes at least one of the following lung diseases: pulmonary tuberculosis, pneumonia, pulmonary nodules and pulmonary cysts.
[0011] The present invention provides a method for screening the blood marker, comprising the following steps:
[0012] Metabolite detection was performed on samples from the non-lung cancer disease group and the lung cancer group, respectively, to obtain two sets of metabolite data;
[0013] The two groups of metabolite data are subjected to compound analysis and identification to obtain metabolite data of a non-lung cancer disease group and metabolite data of a lung cancer group;
[0014] Differential metabolites were screened from the metabolite data of the non-lung cancer disease group and the lung cancer group to obtain blood markers.
[0015] Preferably, the non-lung cancer disease group sample or lung cancer group sample further comprises extraction and impurity removal to obtain an organic phase and an aqueous phase;
[0016] The metabolite detection method is completed by liquid chromatography-mass spectrometry;
[0017] The organic phase was detected by liquid chromatography using Waters ACQUTTY BEH C8 1.7μm 2.1×100mm column; mobile phase A: an aqueous solution containing 0.1% acetic acid by volume and 0.1% ammonium acetate by weight; mobile phase B: an acetonitrile-isopropanol solution containing 0.1% acetic acid by volume and 0.1% ammonium acetate by weight, with a volume ratio of acetonitrile to isopropanol of 7:3; separation and elution gradient: 55%-89% mobile phase B (volume percentage) from 0 to 12 minutes, 100% mobile phase B (volume percentage) from 12 to 19.5 minutes;
[0018] The aqueous phase was detected by liquid chromatography using Waters ACQUTTY HSS T3 1.8μm 2.1×100mm column, mobile phase A is an aqueous solution containing 0.1% by volume formic acid; mobile phase B is an acetonitrile solution containing 0.1% by volume formic acid, and the separation elution gradient is as follows: 0-13min, 1% to 70% by volume mobile phase B, 13-18min, 99% by volume mobile phase B.
[0019] The mass spectrometry detection parameters of the organic or aqueous phase are as follows: mass spectrometry data are collected in Full MS and Full MS / dd-MS2 modes, and the parameters used by Q Exactive are as follows: Full MS mode resolution is 70,000, scan range is 100-1500 m / z, AGC is 3E+6, and Maximum IT is 200 milliseconds; in Full MS / dd-MS2 mode, the secondary mass spectrometry resolution is 17,500, the quadrupole window is 1.5 m / z, the AGC is 1E+5, the maximum ion injection time is 50 ms, and the HCD relative collision energy is 30 eV.
[0020] Preferably, the method for screening differential metabolites is to perform regression analysis on the metabolite data of the non-lung cancer disease group and the metabolite data of the lung cancer group using the least absolute shrinkage and selection algorithm, perform 5-fold cross-validation on the analysis results, retain the metabolites with non-zero regression coefficients in the corresponding model, and obtain the blood markers.
[0021] The present invention provides a blood marker for distinguishing lung cancer from non-lung cancer diseases, comprising at least five of the following metabolites: phosphatidylethanolamine (36:4), lysophosphatidylcholine 18:1e, phosphatidylethanolamine 38:7e, sphingomyelin d38:5, dodecanedioic acid, N-acetylputrescine, L-glutamyl-L-glutamine, hippuric acid, trans-cyclohexane-1,2-dicarboxylic acid, glycochenodeoxycholic acid, allantoic acid, L-valine, methylleucine, and 4-guanidinobutyric acid. The present invention obtains the blood markers through precise screening of lung cancer and non-lung cancer disease samples. A machine learning prediction model based on the blood markers is constructed to distinguish lung cancer from non-lung cancer diseases, with high diagnostic accuracy. Experiments have shown that the diagnostic models constructed based on five, ten, and fourteen of the blood markers have an area under the curve (AUC) greater than 0.77, with a sensitivity greater than 0.72 and a specificity greater than 0.82, indicating that the diagnostic results are accurate and reliable. This indicates that the blood markers provided by the present invention provide effective markers for the diagnosis of lung cancer and the clinical differentiation of lung cancer and non-lung cancer diseases, greatly simplifying the experimental operation and having important clinical diagnostic significance. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 Multivariate ROC curve analysis of the diagnostic model constructed for 14 important markers in the modeling group;
[0023] Figure 2 Multivariate ROC curve analysis of the diagnostic model constructed for 14 important markers in the validation group. DETAILED DESCRIPTION
[0024] The present invention provides a method for screening blood markers for distinguishing lung cancer from non-lung cancer diseases, comprising the following steps:
[0025] Metabolite detection was performed on samples from the non-lung cancer disease group and the lung cancer group, respectively, to obtain two sets of metabolite data;
[0026] The two groups of metabolite data are subjected to compound analysis and identification to obtain metabolite data of a non-lung cancer disease group and metabolite data of a lung cancer group;
[0027] Differential metabolites were screened from the metabolite data of the non-lung cancer disease group and the lung cancer group to obtain blood markers.
[0028] In the present invention, the non-lung cancer disease preferably includes at least one of the following lung diseases: pulmonary tuberculosis, pneumonia, pulmonary nodules and lung infection. The type of the non-lung cancer disease group sample or lung cancer group sample includes blood or plasma.
[0029] In the present invention, before metabolite detection, the non-lung cancer disease group sample or lung cancer group sample preferably further includes extraction and impurity removal to obtain an organic phase and an aqueous phase. The method for extracting and removing impurities of the non-lung cancer disease group sample or lung cancer group sample preferably comprises mixing the sample and the extracting solution to obtain the extracting solution and ultrasonically extracting with methanol aqueous solution to obtain an organic phase and an aqueous phase; mixing the organic phase with an isopropanol acetonitrile solution, incubating the mixture, centrifuging, and collecting the supernatant for later use; mixing the aqueous phase with ice methanol to remove protein, collecting the aqueous phase, removing water, and resuspending it in water for later use. The extracting solution is preferably a mixture of methyl tert-butyl ether and methanol in a volume ratio of 3:1. The volume ratio of methanol to water in the methanol aqueous solution is 3:1. The volume ratio of isopropanol to acetonitrile in the isopropanol acetonitrile solution is 1:3.
[0030] In the present invention, the metabolite detection method is to use liquid chromatography-mass spectrometry technology to complete; the organic phase is detected by liquid chromatography using Waters ACQUTTY BEH C8 1.7μm 2.1×100mm column, mobile phase A is an aqueous solution containing 0.1% acetic acid by volume and 0.1% ammonium acetate by mass; mobile phase B is an acetonitrile-isopropanol (7:3 v / v) solution containing 0.1% acetic acid by volume and 0.1% ammonium acetate by mass, and the separation elution gradient is as follows: 55% to 89% mobile phase B by volume from 0 to 12 minutes, and 100% mobile phase B from 12 to 19.5 minutes. The aqueous phase was detected by liquid chromatography using Waters ACQUTTY HSS T3 1.8μm 2.1×100mm column; mobile phase A was an aqueous solution containing 0.1% by volume formic acid; mobile phase B was an acetonitrile solution containing 0.1% by volume formic acid; the separation elution gradient was as follows: 0-13 minutes, 1% to 70% by volume mobile phase B, 13-18 minutes, 99% by volume mobile phase B. The mass spectrometry detection parameters of the organic or aqueous phase are as follows: mass spectrometry data are collected in Full MS and Full MS / dd-MS2 modes, and the parameters used by Q Exactive are as follows: Full MS mode resolution is 70,000, scan range is 100-1500 m / z, AGC is 3E+6, and Maximum IT is 200 milliseconds; in Full MS / dd-MS2 mode, the secondary mass spectrometry resolution is 17,500, the quadrupole window is 1.5 m / z, the AGC is 1E+5, the maximum ion injection time is 50 ms, and the HCD relative collision energy is 30 eV.
[0031] In the present invention, the two sets of metabolite data preferably further include data processing before compound analysis. The method of data processing is preferably to extract the peak of the RAW format file obtained by the determination into a FeatureXML format file, reduce the dimension of the original mass spectrometry data, and improve the signal-to-noise ratio; using the peak alignment algorithm of OpenMS software, the retention time of the peak format data after extraction is corrected and aligned between samples, thereby converting the mass spectrometry data into a data matrix; matching and filtering the isotope peaks in the data matrix, and then replacing the abnormal data (0, negative value, background noise, etc.) with a vacant value; among all the above-mentioned characteristic peaks, those with a detection rate of <80% are eliminated, and those with a detection rate of >80% are filled with the median value of the characteristic peak, and 5% random noise (obeying a standard normal distribution) is added; using Normalization Autoencoder (NormAE) for homogenization, to remove systematic errors such as batch effects, so as to reduce the difference in metabolite concentrations between samples and make the data distribution more symmetrical.
[0032] In the present invention, the two sets of metabolite data are used for compound analysis, preferably matching the spectral information of the compound primary parent ion (MS1) and the secondary fragment ion (MS2) with the metabolite spectral information in the public database for qualitative judgment. The public database preferably includes the Human Metabolite Database (HMDB, www.hmdb.ca), the Metabolomics Database (Metlin, metlin.scripps.edu), the Mass Spectrum Database (www.massbank.jp) and the Lipid Metabolite Database (Lipidmap, www.lipidmaps.org). The identification method of the two sets of metabolite data is preferably based on the retention time, MS1, and MS2 mass spectrometry information of the compound standard when separated under the same chromatographic column and mass spectrometry conditions. The metabolites are finally verified, specifically based on the standard for metabolite identification, the retention time is within 0.1min difference, and the theoretical value and the measured value of the metabolite molecular weight are less than 10ppm to identify them as the same compound.
[0033] In the present invention, the method for screening differential metabolites is preferably to perform regression analysis on the metabolite data of the non-lung cancer disease group and the metabolite data of the lung cancer group using the least absolute shrinkage and selection algorithm, and to use 5-fold cross-validation on the analysis results, and to retain the metabolites with non-zero regression coefficients in the corresponding model to obtain the blood markers.
[0034] The present invention provides a blood marker for distinguishing lung cancer from non-lung cancer diseases, comprising at least five of the following metabolites: phosphatidylethanolamine (36:4), lysophosphatidylcholine 18:1e, phosphatidylethanolamine 38:7e, sphingomyelin d38:5, dodecanedioic acid, N-acetylputrescine, L-glutamyl-L-glutamine, hippuric acid, trans-cyclohexane-1,2-dicarboxylic acid, glycochenodeoxycholic acid, allantoic acid, L-valine, methylleucine and 4-guanidinobutyric acid.
[0035] In the present invention, when the blood markers include 5 types, the following metabolites are preferably included: a first blood marker combination consisting of lysophosphatidylcholine 18:1e, phosphatidylethanolamine 38:7e, N-acetylputrescine, L-glutamyl-L-glutamine and hippuric acid, a second blood marker combination consisting of phosphatidylethanolamine (36:4), phosphatidylethanolamine 38:7e, dodecanedioic acid, trans-cyclohexane-1,2-dicarboxylic acid and glycochenodeoxycholic acid, and a second blood marker combination consisting of phosphatidylethanolamine (36:4), phosphatidylethanolamine 38:7e, dodecanedioic acid, trans-cyclohexane-1,2-dicarboxylic acid and glycochenodeoxycholic acid. (36:4), lysophosphatidylcholine 18:1e, sphingomyelin d38:5, dodecanedioic acid and N-acetylputrescine; a fourth blood marker combination consisting of phosphatidylethanolamine 38:7e, L-glutamyl-L-glutamine, hippuric acid, glycochenodeoxycholic acid and allantoic acid; and a fifth blood marker combination consisting of lysophosphatidylcholine 18:1e, hippuric acid, trans-cyclohexane-1,2-dicarboxylic acid, glycochenodeoxycholic acid and methylleucine.
[0036] In the present invention, when the blood markers include 10 kinds, the following kinds of metabolites are preferably included: a sixth blood marker combination consisting of lysophosphatidylcholine 18:1e, phosphatidylethanolamine 38:7e, sphingomyelin d38:5, N-acetylputrescine, L-glutamyl-L-glutamine, hippuric acid, trans-cyclohexane-1,2-dicarboxylic acid, glycochenodeoxycholic acid, L-valine and 4-guanidinobutyric acid; a sixth blood marker combination consisting of phosphatidylethanolamine (36:4), lysophosphatidylcholine 18:1e, sphingomyelin d38:5, N-acetylputrescine, L-glutamyl-L-glutamine, hippuric acid, trans-cyclohexane-1,2-dicarboxylic acid, glycochenodeoxycholic acid, L-valine and 4-guanidinobutyric acid; a seventh blood marker combination consisting of d38:5, dodecanedioic acid, N-acetylputrescine, L-glutamyl-L-glutamine, hippuric acid, trans-cyclohexane-1,2-dicarboxylic acid, glycochenodeoxycholic acid, and L-valine, and an eighth blood marker combination consisting of phosphatidylethanolamine 38:7e, sphingomyelin d38:5, dodecanedioic acid, N-acetylputrescine, trans-cyclohexane-1,2-dicarboxylic acid, glycochenodeoxycholic acid, allantoic acid, L-valine, methylleucine, and 4-guanidinobutyric acid.
[0037] The present invention provides the use of the blood marker in constructing a diagnostic model for distinguishing non-lung cancer diseases from lung cancer.
[0038] The types of non-lung cancer diseases addressed by the present invention are the same as those of the above technical solutions and will not be described in detail here.
[0039] In the present invention, the diagnostic model is preferably constructed using a machine learning support vector machine. When constructing the diagnostic model using the machine learning support vector machine, a random loop iteration is performed 1000 times, and the average accuracy of the final model is calculated to construct a diagnostic model for distinguishing non-lung cancer diseases from lung cancer.
[0040] In the present invention, after constructing the diagnostic model, it is preferred to further verify the validity of the diagnostic model. The verification method is to use the validation group sample data for independent verification. In the embodiment of the present invention, the diagnostic threshold for constructing the diagnostic model is 0.6868.
[0041] In an embodiment of the present invention, taking the first blood marker combination as an example, a machine learning model is constructed for diagnosing and distinguishing lung cancer and non-lung cancer diseases. The results show that the diagnostic model constructed by the combination of 5 metabolic markers has a model performance result of AUC = 0.776 (sensitivity = 0.702, specificity = 0.769) in the modeling group and a model performance result of AUC = 0.881 (sensitivity = 0.75, specificity = 0.903) in the validation group.
[0042] In an embodiment of the present invention, the sixth blood marker combination is taken as an example to construct a machine learning model for diagnosing and distinguishing lung cancer and non-lung cancer diseases. The results show that the diagnostic model constructed by the combination of 10 metabolic markers has a model performance result of AUC = 0.83 (sensitivity = 0.726, specificity = 0.821) in the modeling group and a model performance result of AUC = 0.898 (sensitivity = 0.83, specificity = 0.865) in the validation group.
[0043] In an embodiment of the present invention, an experiment was also conducted to construct a diagnostic model using a combination of 14 blood markers. The results showed that the model performance in the modeling group was AUC = 0.855 (sensitivity = 0.75, specificity = 0.846), and the model performance in the validation group was AUC = 0.874 (sensitivity = 0.813, specificity = 0.827). It can be seen that the diagnostic model constructed based on the blood markers provided by the present invention has a good diagnostic effect in distinguishing non-lung cancer diseases from lung cancer.
[0044] The following describes in detail a blood marker for distinguishing lung cancer from non-lung cancer diseases and its application provided by the present invention in conjunction with examples. However, they should not be construed as limiting the scope of protection of the present invention.
[0045] Example 1
[0046] A screening method for blood markers for distinguishing lung cancer from non-lung cancer diseases
[0047] 1. Subjects
[0048] 1) Sample inclusion criteria:
[0049] Subjects must meet all of the following inclusion criteria to be eligible for this study:
[0050] (1) Male or female aged ≥18 years;
[0051] (2) read and fully understand, sign the informed consent form, and be able to provide blood samples for metabolomics testing;
[0052] (3) All lung cancer patients were diagnosed using the gold standard method of histopathological examination;
[0053] (4) Patients diagnosed with benign lung diseases by biopsy / postoperative pathology or by comprehensive evaluation by clinicians, including patients with pneumonia, tuberculosis, pulmonary nodules, lung infections, etc.
[0054] 2) Sample exclusion criteria:
[0055] Subjects who meet any of the following exclusion criteria are not eligible to participate in this study:
[0056] (1) Pregnancy or lactation;
[0057] (2) Emergency or emergency treatment required;
[0058] (3) history of malignant tumor or any anti-tumor treatment before sampling;
[0059] (4) Patients with multiple primary malignant tumors at the same time.
[0060] 3) Subjects' conditions
[0061] In this example, plasma samples were collected from 656 subjects at two medical centers, including 208 non-lung cancer (LNCD) samples and 448 lung cancer (LC) samples. The plasma samples used for the modeling group were: 156 subjects in the non-lung cancer (LNCD) group and 336 subjects in the lung cancer (LC) group; the plasma samples used for the validation group were: 52 subjects in the non-lung cancer (LNCD) group and 112 subjects in the lung cancer (LC) group (see Table 1).
[0062] Table 1 Subjects' information
[0063] Non-lung cancer disease group (LNCD) Lung cancer group (LC) Modeling team members 156 336 Number of people in the verification group 52 112 total 208 448
[0064] 2. Plasma metabolite detection
[0065] 1) Reagents:
[0066] Methanol, acetonitrile, water, acetic acid, and isopropanol of mass spectrometry grade, and formic acid, ammonium acetate, and methyl tert-butyl ether of HPLC grade were purchased from Sigma-Aldrich, USA.
[0067] 2) Sample preparation:
[0068] Take 100 μL of plasma, place it in 1000 μL of pre-cooled solution (the volume ratio of methyl tert-butyl ether and methanol is 3:1), vortex mix the extracted blood sample to obtain a sample extract; add 500 μL of solution (the volume ratio of methanol and water is 3:1) to the sample extract, sonicate, let stand, vortex and centrifuge to separate the layers, the upper layer is the organic phase, and the lower layer is the aqueous phase.
[0069] Organic phase: After sample separation, transfer 500 μL of the upper organic phase to a centrifuge tube. After drying, add 200 μL of acetonitrile and isopropanol (volume ratio 3:1) and incubate at room temperature for 15 minutes. After incubation, vortex the centrifuge tube to mix thoroughly, ultrasonically treat for 5 minutes, and then centrifuge the centrifuge tube at room temperature for 5 minutes (12,000 rpm). Transfer 180 μL of the supernatant from the centrifuge tube to a 2 mL glass injection vial as the organic phase test solution for LC-MS detection.
[0070] Aqueous phase: After the sample is separated, 400 μL of the lower aqueous phase is transferred to a centrifuge tube, and 1100 μL of ice methanol is added thereto to precipitate the protein. After the protein in the centrifuge tube is precipitated, the centrifuge tube is centrifuged, 1000 μL of the supernatant is transferred to a new centrifuge tube, and dried overnight. 200 μL of water is added to the dried centrifuge tube and incubated at room temperature for 15 minutes. After incubation, the centrifuge tube is vortexed and ultrasonically treated for 5 minutes, and then centrifuged at room temperature for 5 minutes (12000 rpm). 180 μL of the supernatant is transferred from the centrifuge tube to a 2 mL glass injection vial as the aqueous phase test solution for LC-MS detection.
[0071] 3) Small molecule metabolite detection:
[0072] The organic phase was purified by Waters ACQUTTY BEH C8 1.7 μm 2.1 × 100 mm column, the aqueous phase uses Waters ACQUTTY An HSS T3 1.8 μm 2.1 × 100 mm column was used for small molecule separation. Liquid chromatography and mass spectrometry were performed using an ACQUITY UPLC I-Class liquid chromatography system (Waters) and a Q-Exactive mass spectrometry system (Thermo Fisher Scientific).
[0073] The mobile phase parameters are as follows:
[0074] Mobile phase parameters for the organic phase test solution: mobile phase A is an aqueous solution containing 0.1% acetic acid and 0.1% ammonium acetate; mobile phase B is an acetonitrile-isopropanol (7:3 v / v) solution containing 0.1% acetic acid and 0.1% ammonium acetate, and the separation elution gradient is as follows: 55%-89% mobile phase B from 0 to 12 minutes, and 100% mobile phase B from 12 to 19.5 minutes.
[0075] Mobile phase parameters for aqueous test solution: mobile phase A is an aqueous solution containing 0.1% formic acid; mobile phase B is an acetonitrile solution containing 0.1% formic acid, and the separation elution gradient is as follows: 0-13 minutes is 1%-70% mobile phase B, 13-18 minutes is 99% mobile phase B.
[0076] The mass spectrometry parameters are as follows:
[0077] Mass spectrometric data were acquired using Full MS and Full MS / dd-MS2 modes (each in positive and negative modes). The QExactive parameters used were as follows: Full MS mode with a resolution of 70,000 s, a scan range of 100–1500 m / z, an AGC of 3E+6, and a Maximum Interval (IT) of 200 ms. In Full MS / dd-MS2 mode, the secondary mass spectrometer had a resolution of 17,500 s, a quadrupole window of 1.5 m / z, an AGC of 1E+5, a maximum ion injection time of 50 ms, and an HCD relative collision energy of 30 eV.
[0078] 3. Metabolomics Data Preprocessing and Metabolite Identification
[0079] 1) Metabolomics data processing:
[0080] (1) Extract peaks from the RAW format files of mass spectrometry data into FeatureXML format files to reduce the dimension of the original mass spectrometry data and improve the signal-to-noise ratio;
[0081] (2) Using the peak alignment algorithm of OpenMS software, the retention times of the extracted peak format data were corrected and aligned between samples, thereby converting the mass spectrometry data into a data matrix;
[0082] (3) Match and filter the isotope peaks in the data matrix obtained in step 2, and then replace abnormal data (0, negative values, background noise, etc.) with blank values;
[0083] (4) Among all the characteristic peaks obtained in step 3, those with a detection rate of <80% are eliminated, and those with a detection rate of >80% are filled with the median value of the characteristic peak and 5% random noise (obeying the standard normal distribution) is added;
[0084] (5) In order to reduce the differences in metabolite concentrations between samples and make the data distribution more symmetrical, Normalization Autoencoder (NormAE) was used for normalization processing to remove systematic errors such as batch effects.
[0085] 2) Identification of metabolites:
[0086] The raw data is parsed by software to obtain the spectral information of the compound's primary parent ion (MS1) and secondary fragment ion (MS2), such as the mass-to-charge ratio (m / z) of the primary mass spectrum and the fragments of the secondary ion, which are matched with the spectral information of primary and secondary metabolites in the public database to qualitatively identify the metabolites. Commonly used metabolite databases include the Human Metabolite Database (HMDB, www.hmdb.ca), the Metabolomics Database (Metlin, metlin.scripps.edu), the Mass Spectrum Database (www.massbank.jp), and the Lipid Metabolite Database (Lipidmap, www.lipidmaps.org). Based on the metabolites identified in the relevant databases, the metabolites are finally verified based on the retention time, MS1, and MS2 mass spectrometric information when the standard is separated under the same chromatographic column and mass spectrometry conditions. The criteria for metabolite identification are that the retention time is within 0.1 min and the theoretical and measured values of the metabolite molecular weight are less than 10 ppm.
[0087] 4. Data Analysis
[0088] 1) Screening of metabolic markers for distinguishing non-lung cancer diseases from lung cancer
[0089] Metabolite detection was performed on the above samples, and a total of 515 metabolites were obtained after annotation. LASSO (Least Absolute Shrinkage and Selection Operator) regression analysis was performed on the data of the modeling group. The average error corresponding to each regularization parameter alpha was calculated using 5-fold cross-validation, and the optimal alpha = 0.03 with the smallest error was obtained. Metabolites with non-zero regression coefficients in their corresponding models were retained. Finally, a total of 14 differential metabolites were screened and obtained (Table 2) as important metabolic markers to distinguish the LNCD group from the LC group.
[0090] Table 2 14 metabolic markers for distinguishing non-lung cancer diseases from lung cancer
[0091]
[0092]
[0093] 2) Construction of a diagnostic model for distinguishing non-lung cancer diseases from lung cancer
[0094] To verify the diagnostic efficacy of the 14 screened markers in distinguishing LNCD from LC, a multivariate receiver operating characteristic (ROC) curve analysis was performed on these 14 markers in the modeling group. Three-quarters of the sample data from the LNCD and LC groups in the modeling group were randomly used as the training set, and one-quarter was used as the test set. A machine learning support vector machine (SVM) was used with 1,000 random iterations. A diagnostic model for distinguishing non-lung cancer diseases from lung cancer was constructed by calculating the average accuracy of the final model.
[0095] The ROC curve is a method for studying the relationship between model sensitivity and specificity. It uses sensitivity as the vertical axis and 1-specificity as the horizontal axis. Evaluation is based on the area under the curve (AUC). When the AUC is greater than 0.5, the closer the AUC is to 1, the better the model performance and the better the diagnostic effect. A value less than 0.5 indicates poor model accuracy. In addition to common parameters such as the receiver operating characteristic (ROC) curve and the area under the curve (AUC), ROC classification prediction models also include sensitivity and specificity.
[0096] The sensitivity calculation method is shown in Formula I. The specificity calculation method is shown in Formula II.
[0097]
[0098] Among them, TP (True Positive): true positive, the number of samples that are actually positive examples that are correctly predicted as positive examples;
[0099] TN (Ture Negative): True negative, the number of samples that are actually negative examples but are correctly predicted as negative examples;
[0100] FP (False Positive): False positive, the number of samples that are actually negative examples but are mistakenly predicted as positive examples;
[0101] FN (False Negative): False negatives, the number of samples that are actually positive examples but are mistakenly predicted as negative examples.
[0102] The results are as follows Figure 1 As shown in the figure, AUC = 0.855 (sensitivity = 0.75, specificity = 0.846), which shows that the constructed diagnostic model has high diagnostic efficacy.
[0103] 3) Validation of diagnostic models for distinguishing non-lung cancer diseases from lung cancer
[0104] In order to further verify the effectiveness of the diagnostic model for distinguishing non-lung cancer diseases from lung cancer based on the modeling group data, the validation group data was used to verify the above diagnostic model. A multivariate ROC curve analysis was specifically performed to evaluate the independent verification effect of the diagnostic model on unknown data sets other than the modeling group data set. After the validation group samples were placed in the diagnostic model constructed by the modeling group, the corresponding probability value (Probability) was output based on the detection data of 14 important metabolic markers for distinguishing non-lung cancer diseases from lung cancer for each sample. Based on the probability value of each sample as the diagnostic threshold, a set of confusion matrices (including true positive, true negative, false positive and false negative) were obtained. Sensitivity and specificity can be calculated according to the formula, and a point can be marked in the ROC analysis graph with sensitivity (sensitivity) as the vertical coordinate and 1-specificity (1-specificity) as the horizontal coordinate. Similarly, when the probability value of each sample is used as the diagnostic threshold, multiple different points are obtained in the ROC analysis graph. These points are linked to draw the ROC curve graph ( Figure 2 ). Among them, the point with the best sensitivity and specificity was selected, and the diagnostic threshold was 0.6868.
[0105] As shown in the confusion matrix results in Table 3, in the diagnostic model constructed based on the above 14 metabolic markers, with a diagnostic threshold of 0.6868, among the 52 lung cancer patients, 43 were diagnosed as lung cancer, and 9 were misdiagnosed as non-lung cancer patients; among the 112 non-lung cancer disease subjects, 91 were correctly diagnosed, and 21 were misdiagnosed as lung cancer. The ROC analysis results of the diagnostic model in the validation group are shown in Figure 3. Figure 2 As shown, based on the confusion matrix results, sensitivity and specificity were calculated, and the results showed AUC = 0.874 (sensitivity = 0.813, specificity = 0.827). The above results show that the diagnostic model constructed for distinguishing non-lung cancer diseases from lung cancer also has a good diagnostic effect in the validation group.
[0106] Table 3 Confusion matrix of the diagnostic model used to distinguish non-lung cancer diseases from lung cancer
[0107] lung cancer Subjects with non-lung cancer diseases 52 cases of lung cancer 43(TP) 9(FN) 112 cases of non-lung cancer diseases 21(FP) 91(TN)
[0108] Example 2
[0109] In order to further simplify the number of metabolic markers during diagnosis, this example selects the following 10 metabolic marker combinations according to the method of Example 1 to construct a diagnostic model to verify the diagnostic effect: hemolytic phosphatidylcholine 18:1e, phosphatidylethanolamine 38:7e, sphingomyelin d38:5, N-acetylputrescine, L-glutamyl-L-glutamine, hippuric acid, trans-cyclohexane-1,2-dicarboxylic acid, glycochenodeoxycholic acid, L-valine, and 4-guanidinobutyric acid combination.
[0110] The results showed that the diagnostic model constructed based on the combination of 10 metabolite markers had an AUC of 0.83 (sensitivity = 0.726, specificity = 0.821). The results showed that the diagnostic models constructed based on the combination of the above 10 metabolite markers all had high diagnostic efficacy.
[0111] The diagnostic model constructed by combining the 10 metabolite markers was validated. The results showed that the diagnostic model constructed by combining the 10 metabolite markers had an AUC of 0.898 (sensitivity = 0.83, specificity = 0.865) in the validation group.
[0112] Example 3
[0113] In order to further simplify the number of metabolic markers during diagnosis, this example selects the following five metabolic marker combinations according to the method of Example 1 to construct a diagnostic model to verify the diagnostic effect: hemolytic phosphatidylcholine 18:1e, phosphatidylethanolamine 38:7e, N-acetylputrescine, L-glutamyl-L-glutamine, and hippuric acid combination.
[0114] The results showed that the diagnostic model constructed based on the combination of the five metabolite markers had an AUC of 0.776 (sensitivity = 0.702, specificity = 0.769). The results showed that the diagnostic models constructed based on the combination of the above five metabolite markers all had high diagnostic efficacy and clinical diagnostic significance.
[0115] The diagnostic model constructed by combining the five metabolite markers was validated in the validation group. The diagnostic model constructed by combining the five metabolite markers had an AUC of 0.881 (sensitivity = 0.75, specificity = 0.903).
[0116] The above results show that the diagnostic model constructed to distinguish non-lung cancer diseases from lung cancer also has a good diagnostic effect in the validation group.
[0117] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A blood marker for distinguishing lung cancer from non-lung cancer diseases, characterized in that: Includes at least five of the following metabolites: phosphatidylethanolamine (36:4), lysophosphatidylcholine 18:1e, phosphatidylethanolamine 38:7e, sphingomyelin d38:5, dodecanedioic acid, N-acetylputrescine, L-glutamyl-L-glutamine, hippuric acid, trans-cyclohexane-1,2-dicarboxylic acid, glycochenodeoxycholic acid, allantoic acid, L-valine, methylleucine, and 4-guanidinobutyric acid.
2. The blood marker according to claim 1, characterized in that When the blood markers include five types, the following metabolites are included: a first blood marker combination consisting of lysophosphatidylcholine 18:1e, phosphatidylethanolamine 38:7e, N-acetylputrescine, L-glutamyl-L-glutamine and hippuric acid; a second blood marker combination consisting of phosphatidylethanolamine (36:4), phosphatidylethanolamine 38:7e, dodecanedioic acid, trans-cyclohexane-1,2-dicarboxylic acid and glycochenodeoxycholic acid; a second blood marker combination consisting of phosphatidylethanolamine (36: 4), a third blood marker combination consisting of lysophosphatidylcholine 18:1e, sphingomyelin d38:5, dodecanedioic acid, and N-acetylputrescine, a fourth blood marker combination consisting of phosphatidylethanolamine 38:7e, L-glutamyl-L-glutamine, hippuric acid, glycochenodeoxycholic acid, and allantoic acid, and a fifth blood marker combination consisting of lysophosphatidylcholine 18:1e, hippuric acid, trans-cyclohexane-1,2-dicarboxylic acid, glycochenodeoxycholic acid, and methylleucine.
3. The blood marker according to claim 1, characterized in that When the blood markers include 10 types, the following metabolites are included: a sixth blood marker combination consisting of lysophosphatidylcholine 18:1e, phosphatidylethanolamine 38:7e, sphingomyelin d38:5, N-acetylputrescine, L-glutamyl-L-glutamine, hippuric acid, trans-cyclohexane-1,2-dicarboxylic acid, glycochenodeoxycholic acid, L-valine and 4-guanidinobutyric acid; a sixth blood marker combination consisting of phosphatidylethanolamine (36:4), lysophosphatidylcholine 18:1e, sphingomyelin d38:5, N-acetylputrescine, L-glutamyl-L-glutamine, hippuric acid, trans-cyclohexane-1,2-dicarboxylic acid, glycochenodeoxycholic acid, L-valine and 4-guanidinobutyric acid.
5. The seventh blood marker combination consisting of dodecanedioic acid, N-acetylputrescine, L-glutamyl-L-glutamine, hippuric acid, trans-cyclohexane-1,2-dicarboxylic acid, glycochenodeoxycholic acid and L-valine; and the eighth blood marker combination consisting of phosphatidylethanolamine 38:7e, sphingomyelin d38:5, dodecanedioic acid, N-acetylputrescine, trans-cyclohexane-1,2-dicarboxylic acid, glycochenodeoxycholic acid, allantoic acid, L-valine, methylleucine and 4-guanidinobutyric acid.
4. Use of the blood marker according to any one of claims 1 to 3 in constructing a diagnostic model for distinguishing non-lung cancer diseases from lung cancer.
5. The application according to claim 4, characterized in that: The diagnostic model is constructed using a machine learning support vector machine.
6. The application according to claim 4, characterized in that: The non-lung cancer disease includes at least one of the following lung diseases: pulmonary tuberculosis, pneumonia, pulmonary nodules and pulmonary cysts.
7. The method for screening blood markers according to any one of claims 1 to 3, characterized in that: The following steps are involved: Metabolite detection was performed on samples from the non-lung cancer disease group and the lung cancer group, respectively, to obtain two sets of metabolite data; The two groups of metabolite data are subjected to compound analysis and identification to obtain metabolite data of a non-lung cancer disease group and metabolite data of a lung cancer group; Differential metabolites were screened from the metabolite data of the non-lung cancer disease group and the lung cancer group to obtain blood markers.
8. The screening method according to claim 7, characterized in that The non-lung cancer disease group sample or the lung cancer group sample further comprises extraction and impurity removal to obtain an organic phase and an aqueous phase; The metabolite detection method is completed by liquid chromatography-mass spectrometry; The organic phase was detected by liquid chromatography using Waters BEH C8 1.7μm 2.1×100mm column; mobile phase A: an aqueous solution containing 0.1% acetic acid by volume and 0.1% ammonium acetate by weight; mobile phase B: an acetonitrile-isopropanol solution containing 0.1% acetic acid by volume and 0.1% ammonium acetate by weight, with a volume ratio of acetonitrile to isopropanol of 7:3; separation and elution gradient: 55%-89% mobile phase B by volume from 0 to 12 minutes, 100% mobile phase B by volume from 12 to 19.5 minutes; The aqueous phase was detected by liquid chromatography using Waters HSS T3 1.8μm 2.1×100mm column, mobile phase A is an aqueous solution containing 0.1% by volume formic acid; mobile phase B is an acetonitrile solution containing 0.1% by volume formic acid, and the separation elution gradient is as follows: 0-13min, 1% to 70% by volume mobile phase B, 13-18min, 99% by volume mobile phase B.
9. The screening method according to claim 8, characterized in that The mass spectrometry detection parameters of the organic or aqueous phase are as follows: mass spectrometry data are collected in Full MS and Full MS / dd-MS2 modes, and the parameters used by Q Exactive are as follows: Full MS mode resolution is 70,000, scan range is 100-1500 m / z, AGC is 3E+6, and Maximum IT is 200 milliseconds; in Full MS / dd-MS2 mode, the secondary mass spectrometry resolution is 17,500, the quadrupole window is 1.5 m / z, the AGC is 1E+5, the maximum ion injection time is 50 ms, and the HCD relative collision energy is 30 eV.
10. The screening method according to any one of claims 7 to 9, characterized in that The method for screening differential metabolites is to perform regression analysis on the metabolite data of the non-lung cancer disease group and the metabolite data of the lung cancer group using the least absolute shrinkage and selection algorithm, perform 5-fold cross-validation on the analysis results, retain the metabolites with non-zero regression coefficients in the corresponding model, and obtain the blood markers.