Non-small cell lung cancer tissue discrimination model based on differential metabolite combination as well as construction method and application of non-small cell lung cancer tissue discrimination model
By constructing a logistic regression model based on the differential metabolite combination of lactic acid, ethanolamine phosphate, decenyl carnitine and symmetric dimethylarginine, the problem of time-consuming and insufficient sensitivity for diagnosis of non-small cell lung cancer is solved, and efficient lung cancer tissue detection and surgical decision-making assistance is achieved.
Patent Information
- Application Number
- CN202510649789.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the diagnosis method of non-small cell lung cancer is time-consuming and subjective. Traditional imaging and biomarkers such as CEA are insufficient in sensitivity, making it difficult to effectively assist surgeons in determining the scope of tumor resection during surgery.
A logistic regression discriminant model based on the differential metabolites combination of lactic acid, ethanolamine phosphate, decenyl carnitine and symmetric dimethylarginine was constructed, and a key metabolites were screened through UPLC-HRMS detection and machine learning algorithm was used to establish an efficient non-small cell lung cancer tissue discriminant model.
A high sensitivity and specificity of non-small cell lung cancer tissue detection is achieved, assisting surgeons to accurately decide the range of resection during surgery, reduce model complexity and improve detection performance.
Smart Images

Figure CN120490369A_ABST
Abstract
Description
Technical field:
[0001] The present invention belongs to the field of biomedical technology, and specifically relates to a non-small cell lung cancer tissue discrimination model based on differential metabolite combinations, as well as its construction method and application. The logistic regression discrimination model based on four differential metabolite combinations of the present invention can effectively distinguish non-small cell lung cancer tumor tissue from normal tissue. Background technology:
[0002] Lung cancer has a high morbidity and mortality rate worldwide, with non-small cell lung cancer (NSCLC) accounting for approximately 85%. Common diagnostic methods for lung cancer primarily include CT imaging and histopathological analysis. These traditional examination methods are time-consuming and subjective, often relying on the physician's experience and expertise. During surgery, physicians must excise the tumor as completely as possible while sparing normal tissue to the greatest extent possible to minimize trauma to the patient. Pathological diagnosis can help physicians determine the extent of tumor resection. However, diagnostic concordance among different pathologists for the same NSCLC specimen ranges from 67.1% to 89.6%. Tumor markers are substances produced during tumor cell development and proliferation and can aid in diagnosis. Existing lung cancer blood biomarkers, such as carcinoembryonic antigen (CEA), have clinical value, but they still suffer from significant inter-individual variability and limited sensitivity. The identification of lung cancer tissue biomarkers is crucial for identifying lung cancer biomarkers and assisting in clinical decision-making regarding the extent of tumor resection.
[0003] Lung cancer is a disease closely related to metabolism, and its tumor cells show significant differences in multiple metabolic pathways, including sugar metabolism, nucleic acid metabolism, and lipid metabolism. Metabolic reprogramming is one of the characteristics of tumors. Tumor cells gain advantages in proliferation and metastasis through metabolic reprogramming. Uncovering the specific metabolites of tumor cells is of great significance for discovering new diagnostic markers and improving treatment efficacy. The combination of metabolomics and machine learning provides a new technical means for the discovery of metabolic markers and disease diagnosis. Metabolomics data are characterized by a large variety of metabolites and large differences in abundance values. Machine learning shows significant advantages in processing such complex data. It can quickly identify potential patterns in the data, extract key information and build predictive models. For example, Chinese patent CN202311708074.6 discloses an application and kit of metabolic markers for diagnosing lung cancer staging. By extracting blood samples from healthy and lung cancer subjects for metabolite detection, processing and identification, 19 metabolic markers were screened out through modeling analysis, including dehydroepiandrosterone sulfate, 4-hydroxy-1-(3-pyridine)-1-butanone, L-leucine dipeptide, octanoyl-L-carnitine, phosphatidylcholine 34:3e, acylcarnitine 10:0, N-formylmethionine, 5-methylthioadenosine, phosphatidylcholine 36:3e, fatty acid 22:6, hydrocinnamic acid, acylcarnitine 10:1, hypoxanthine, L-alanyl-L-aspartic acid, iminodiacetic acid, betaine, choline, palmitoyl-L-carnitine and L-glutamate. Chinese patent CN202311651833.X discloses a lung cancer metabolite marker combination, its screening method and application. The lung cancer metabolite marker combination includes the following compounds: glucose, urea, creatinine, uric acid, L-phenylalanine, L-valine, L-histidine, L-arginine and L-leucine. The screening method is to collect serum or plasma samples from non-lung cancer and lung cancer patients, extract metabolites by MALDI mass spectrometry analysis and perform data preprocessing to obtain alternative metabolite marker characteristics. After training and verification using multiple machine learning models, the metabolite markers in the two machine learning models with the best performance are ranked by importance, and the intersection is taken to determine the potential metabolite marker combination.
[0004] The above-mentioned existing technologies screen out a large number of metabolites, which results in a complex screening process. Therefore, it is necessary to research and establish a lung cancer discrimination model based on a small number of metabolite combinations with good discrimination performance. Summary of the invention:
[0005] The purpose of the present invention is to provide a non-small cell lung cancer tissue discrimination model based on differential metabolite combinations, as well as its construction method and application. The differential metabolite combination has the potential to assist in the clinical detection and classification of non-small cell lung cancer tissue, and the potential to assist surgeons in deciding the extent of resection during tumor resection.
[0006] To achieve the above objectives, the present invention provides a non-small cell lung cancer tissue discrimination model based on a differential metabolite combination consisting of lactic acid, phosphoethanolamine, 9-decenoylcarnitine, and symmetric dimethylarginine (SDMA). The logistic regression discriminant model based on the differential metabolite combination of the present invention has an AUC value greater than 0.85 for distinguishing lung cancer from normal tissue samples.
[0007] The present invention also provides a method for constructing a non-small cell lung cancer tissue discrimination model based on differential metabolite combinations, using a non-targeted metabolomics method based on ultra-high performance liquid chromatography-high resolution mass spectrometry (UPLC-HRMS) to detect and identify metabolites in tumor and normal samples of non-small cell lung cancer patients, screen key differential metabolites, and construct a discrimination model; the specific steps are as follows:
[0008] (1) Collection and processing of tissue samples: Non-small cell lung cancer tumor tissue and paired normal tissue samples were collected and metabolites were extracted;
[0009] (2) UPLC-HRMS detection: UPLC-HRMS method was used to perform non-targeted metabolomics analysis of tissue metabolites to obtain raw mass spectrometry data; Progenesis QI software was used to preprocess the mass spectrometry data, perform peak extraction and peak alignment, and search the metabolite charge ratio (m / z) matching candidate metabolites based on the Human Metabolomics Database (HMDB); the mass spectrometry peaks were further processed for decontamination, blank removal, adduct and source lysate discrimination, QC correction and segmented normalization (Zhang J, Zang X, Jiao P, et al. Alterations of Ceramides, Acylcarnitines, GlyceroLPLs, and Amines in NSCLC Tissues. J. Proteome Res. 2024; 23: 4343-4358.) to generate the sample metabolite abundance matrix;
[0010] (3) Metabolite identification: UPLC combined with secondary mass spectrometry was used to identify the detected metabolites;
[0011] (4) Differential metabolite screening and model testing. The specific steps are as follows:
[0012] (a) The paired signed-rank test method was used to calculate the difference in metabolite abundance between tumor tissues and normal tissues. The Benjamini-Hochberg correction method was used to control the false discovery rate (FDR) to less than 0.05. In combination with the criterion of fold difference (FC) greater than 1.25 or less than 0.8, differential metabolites between lung cancer tissues and normal tissues were screened.
[0013] (b) The tumor group and normal group sample data were randomly split into a training set and a validation set at a ratio of 70% and 30%; based on the training set data, the LASSO method was used to screen the key differential metabolites;
[0014] (c) Based on the identified key differential metabolites, the performance of six discriminant models, including decision tree, random forest, K-nearest neighbor, naive Bayes, support vector machine, and logistic regression, was tested. The models were trained using the training set and validated using the test set. The robustness and discriminant performance of the models were comprehensively evaluated based on accuracy, sensitivity, specificity, precision, F1 score, and AUC (area under the receiver operating characteristic curve). The optimal model was the logistic regression model, which was selected as the discriminant model for non-small cell lung cancer tissue.
[0015] The non-small cell lung cancer tissue discrimination model equation based on the key differential metabolite combination of the present invention is:
[0016] ln[p / (1-p)]=0.513+1.189×lactic acid+1.771×phosphoethanolamine-0.490×decenoylcarnitine+1.108×symmetric dimethylarginine;
[0017] The compound name in the above formula represents the relative abundance of the compound. The optimal cut-off value (cut-off) was determined to be 0.468 based on the Youden index. When p ≤ 0.468, the sample was judged as normal tissue, and when p > 0.468, the sample was judged as non-small cell lung cancer tumor tissue.
[0018] In step (4) (c) of the present invention, the further step includes: randomly selecting 50% of the samples from the tumor group and the normal group to form a test set, repeating 10 times to obtain 10 different test sets for verifying the robustness of the selected model.
[0019] The present invention also provides the use of the differential metabolite combination in preparing a non-small cell lung cancer tissue detection kit. The detection method of the kit includes a UPLC-HRMS method, and the reagents of the kit include a sample metabolite extraction solvent and a UPLC-HRMS detection solvent.
[0020] Compared with the prior art, the present invention has the following beneficial effects:
[0021] (1) Highly sensitive and specific detection performance for non-small cell lung cancer tissue
[0022] In response to the limitations of existing biomarkers (such as CEA, etc.) in terms of insufficient diagnostic sensitivity and the time-consuming and highly subjective nature of traditional histopathological analysis and imaging methods, the present invention screened out four key differential metabolites: lactate, phosphoethanolamine, decenoylcarnitine, and symmetric dimethylarginine. Based on these key differential metabolites, a logistic regression discriminant model was constructed for the detection of non-small cell lung cancer. The model demonstrated stable performance in the training set, validation set, and 10 randomized test sets, with accuracy, sensitivity, specificity, precision, F1 score, and AUC all greater than 0.8. This metabolite combination supports high-throughput detection based on UPLC-HRMS, enabling the application of kits. It has the potential to assist in the detection and classification of non-small cell lung cancer tissues, and to assist surgeons in deciding the extent of resection during tumor resection.
[0023] (2) Efficient screening method
[0024] In response to the challenges of high dimensionality and high noise in metabolomics data, the present invention combines LASSO regularization constraints and a ten-fold cross-validation strategy to screen out four key differential metabolites, significantly reducing the complexity of the model (the number of features dropped from 72 to 4, a decrease of 94.4%). The performance of the decision tree, random forest, K nearest neighbor, naive Bayes, support vector machine and logistic regression discriminant models based on the key differential metabolites was then tested. The logistic regression model based on the four key differential metabolites had the best discriminant performance: the AUC reached 0.890 in the training set, an increase of 8.3% over the optimal single metabolite (decenoylcarnitine, AUC = 0.807). The test set (AUC>0.85) showed the excellent robustness of the model, and the Hosmer-Lemeshow test (p = 0.109) confirmed the goodness of fit of the model.
[0025] In summary, the present invention screened out four key differential metabolites, which are of a small variety; the logistic regression model based on the combination of the four key differential metabolites can effectively distinguish non-small cell lung cancer tumor tissue and normal tissue, with good sensitivity, specificity and accuracy, showing potential in assisting the detection and classification of non-small cell lung cancer tissue, and has the potential to assist surgeons in deciding the extent of resection during tumor resection. Description of the drawings:
[0026] Figure 1 The figure is a schematic diagram of the screening process of key differential metabolites involved in the present invention.
[0027] Figure 2 Figure 2 is the relative abundance of four key differential metabolites in tumor and normal tissues of patients with non-small cell lung cancer, where A is lactate, B is phosphoethanolamine, C is decenoylcarnitine, and D is symmetric dimethylarginine. *** indicates a statistically significant difference between the two groups (p < 0.001).
[0028] Figure 3 Shown are the structures and MS / MS spectra of four key differential metabolites.
[0029] Figure 4 Figure 3 ROC curves of the models based on four key differential metabolites established using different machine learning algorithms. A is the training set and B is the validation set. Specific implementation method:
[0030] The present invention will be further described in detail below through specific embodiments in conjunction with the accompanying drawings.
[0031] Example 1:
[0032] This example relates to a method for constructing a non-small cell lung cancer tissue discrimination model based on differential metabolite combinations, and the specific steps are as follows:
[0033] 1. Collect and process tissue samples
[0034] (a) Collection of tissue samples
[0035] This study was approved by the Scientific Ethics Committee of the Academic Committee of Ocean University of China (OUC-HM-2022-005) and the Ethics Committee of Qingdao Municipal Hospital (2022-Y014). Tumor lesion tissue (as tumor samples) and distal normal lung tissue (as normal samples) were collected from 104 patients with non-small cell lung cancer (NSCLC), including 88 patients with lung adenocarcinoma and 16 with lung squamous cell carcinoma. Blood was removed from the collected tissue samples using gauze and then quickly transferred to liquid nitrogen for short-term storage before being stored in a -80°C freezer for long-term storage.
[0036] Table 1 Clinical information of patients with non-small cell lung cancer
[0037]
[0038] (b) Tissue metabolite extraction
[0039] Tissue samples were thawed on ice at room temperature and cut into small pieces with scissors. 100 mg of minced tissue sample was placed in a homogenizer tube and homogenized with 1000 μL of pre-chilled 80% (v / v) methanol in water (60 Hz, 90 seconds). The sample was vortexed for 1 minute and then incubated in a −80°C freezer for 4 hours. After incubation, the sample was centrifuged at 16,000 g for 10 minutes at 4°C. The supernatant was transferred to a new EP tube and concentrated using a vacuum centrifuge. After the sample dried, 100 μL of water was added for reconstitution, vortexed again for 15 seconds, and centrifuged at 16,000 g for 10 minutes. The supernatant was finally transferred to a sample injection vial. In addition, 100 μL of water was used instead of tissue sample as a blank control to eliminate background ion interference. Quality control (QC) samples were prepared by mixing the samples to monitor and correct for variations in instrument response.
[0040] 2. UPLC-HRMS detection
[0041] (a) Untargeted metabolomics analysis of the extracted metabolites was performed using UPLC-LTQ Orbitrap XL to obtain raw mass spectrometry data.
[0042] The HPLC conditions are as follows: the chromatographic column is The elution was performed on a BEH C18 column (2.1×50 mm, 1.7 μm) with an injection volume of 10 μL. The column temperature was 35°C, the mobile phase A was an aqueous solution containing 0.1% (v / v) acetic acid, the mobile phase B was an acetonitrile solution, and the gradient elution conditions were: 0 min (100% A, 0.25 mL / min), 1 min (90% A, 0.25 mL / min), 2.5 min (85% A, 0.25 mL / min), 4 min (78% A, 0.25 mL / min), 6 min (62% A, 0.25 mL / min), 9 min (35% A, 0.25 mL / min), 12 min (20% A, 0.25 mL / min), 16 min (0% A, 0.30 mL / min), and 18 min (0% A, 0.45 mL / min).
[0043] The mass spectrometer LTQ Orbitrap XL used ESI ionization, with detection in both positive and negative ion modes. The spray voltage was 4 kV for the positive electrode and 2.5 kV for the negative electrode. The capillary temperature was 300°C, the sheath gas flow rate was 40 arb, the auxiliary gas flow rate was 10 arb, the capillary voltage was 20 V for the positive electrode and -40 V for the negative electrode, and the tube lens voltage was 35 V for the positive electrode and -80 V for the negative electrode. The scan mode was Full Scan, the scan range was 50–1000 (m / z), and the resolution was 60,000 (FWHM).
[0044] (b) Mass spectrometry data processing
[0045] The mass spectrometry data were preprocessed using Progenesis QI software, and the adduct ion was set to [M+H] in the positive ion mode. + 、[M+Na] + 、[M+K] + 、[M+2H] 2+ 、[M+Na+H] 2+ 、[M+2Na] 2+ 、[M+3H] 3+ 、[M+2H+Na] 3+ and [M+2Na+H] 3+ , negative ions are set to [MH] - , [M-2H] 2- and [M-3H] 3- Peak extraction and peak alignment were performed, and metabolite mass-to-charge ratio (m / z) matching candidate metabolites was retrieved based on the Human Metabolomics Database (HMDB). Mass spectrometry peaks were further processed for contamination removal, blank removal, adduct and source lysate discrimination, QC correction, and segment normalization (Zhang J, Zang X, Jiao P, et al. Alterations of Ceramides, Acylcarnitines, GlyceroLPLs, and Amines in NSCLC Tissues. J. Proteome Res. 2024; 23: 4343-4358.) to generate a sample metabolite abundance matrix.
[0046] 3. Metabolite identification
[0047] Compound identification was performed using UPLC coupled with secondary mass spectrometry (MS / MS). Specifically, after UPLC separation, fragment ions of metabolites were obtained by MS / MS at a resolution of 30,000 (FWHM) in MS / MS mode. These fragments were then compared with spectra in the HMDB and MassBank databases, and the compounds were identified with high confidence by manually analyzing the fragmentation patterns and fragment ion structures.
[0048] 4. Screening of differential metabolites and model testing. The specific steps are as follows:
[0049] (a) Screening for differential metabolites between tumor tissue and normal tissue
[0050] For metabolites identified in step 3, the Wilcoxon matched-pairs signed-rank test was first used to calculate the P value for tumor tissue versus paired normal tissue. The false discovery rate (FDR) was then calculated using the Benjamini-Hochberg multiple testing correction method. The fold change (FC) between the tumor and normal groups was also calculated based on the median metabolite abundance. A FDR < 0.05 and (FC > 1.25 or FC < 0.8) threshold for statistical significance was used. A total of 72 differentially expressed metabolites were identified.
[0051] To evaluate the classification performance of the 72 differentially expressed metabolites for non-small cell lung cancer tumor and normal tissues, receiver operating characteristic (ROC) curves were constructed, and sensitivity, specificity, and area under the ROC curve (AUC) were calculated. Five metabolites, including octanoylcarnitine, guanosine monophosphate (GMP), N2,N2-dimethylguanosine, 9-decenoylcarnitine, and adenosine monophosphate (AMP), were found to have sensitivity, specificity, and AUC greater than 70% (Table 2). Because individual metabolites often lack sufficient sensitivity and specificity, their value in clinical applications is limited.
[0052] Table 2 Discriminant performance of individual metabolites in non-small cell lung cancer tissue
[0053]
[0054] Note: * indicates that the median of normal samples in the FC (tumor / normal) calculation process is 0.
[0055] (b) LASSO method was used to screen out key differential metabolites
[0056] After z-score normalization of the abundance matrix of the 72 differential metabolites obtained in step 4(a), the data set was randomly split into a training set and a validation set at a ratio of 70% and 30%. Based on the training set data, LASSO regression was used to screen variables, and the optimal regularization parameter was selected by the ten-fold cross-validation method. Four key differential metabolites were obtained: lactic acid, phosphoethanolamine, 9-decenoylcarnitine, and symmetric dimethylarginine (SDMA). As shown in Table 3 and Figure 2 As shown in Figure 2, compared with normal tissues, lactate, phosphoethanolamine, and symmetric dimethylarginine were significantly increased in non-small cell lung cancer tumor tissues, while decenoylcarnitine was significantly decreased. The structures and MS / MS spectra of the four metabolites are shown in Figure 2. Figure 3 shown.
[0057] Table 3 Four key differential metabolites
[0058]
[0059] (c) Performance of the discriminant model based on the combination of key differential metabolites
[0060] Based on the four key differential metabolites in Table 3, with tumor and normal as response variables, six machine learning algorithms, including decision tree (DT), random forest (RF), K-nearest neighbors (KNN), naive Bayes (NB), support vector machine (SVM) and logistic regression (LR), were used to construct discriminant models and perform performance evaluation. Modeling was performed using the training set and validation set. The results showed (Table 4 and Figure 4 ), these algorithms demonstrated similar classification performance, with validation set AUC values ranging from 0.841 to 0.882 for each model, indicating that this differential metabolite combination possessed strong generalization capabilities and stable discriminant performance. Comprehensively considering accuracy, sensitivity, specificity, precision, F1 score, and AUC, the logistic regression model was ultimately selected as the optimal discriminant model.
[0061] Table 4. Discrimination performance of models based on four key differential metabolites established using different machine learning algorithms.
[0062]
[0063]
[0064] A non-small cell lung cancer tissue discrimination model was constructed based on the logistic regression algorithm, and the obtained logistic regression equation was:
[0065] ln[P / (1-P)]=0.513+1.189×lactic acid+1.771×phosphoethanolamine-0.490×decenoylcarnitine+1.108×symmetric dimethylarginine;
[0066] In the above formula, the compound name represents the relative abundance of the compound. The optimal cut-off value (cut-off) was determined to be 0.468 according to the Youden index. When p≤0.468, the sample was judged to be normal tissue, and when p>0.468, the sample was judged to be non-small cell lung cancer tumor tissue. The accuracy of the training set data in this model was 0.849, the sensitivity was 0.836, the specificity was 0.863, the precision was 0.859, the F1 score was 0.847, and the AUC was 0.890. The goodness of fit of the logistic regression model was tested using Hosmer-Lemeshow, p=0.109, indicating that the model fit was good. The validation set was input into the above model, with an accuracy of 0.871, a sensitivity of 0.839, a specificity of 0.903, a precision of 0.897, an F1 score of 0.867, and an AUC of 0.874. To further validate the robustness of the model, we randomly selected 50% of the samples to form a test set. This was repeated 10 times, resulting in 10 different test sets. The model was fed these 10 test set data. The corresponding accuracy, sensitivity, specificity, precision, and F1 score were all greater than 0.8, and the AUC was greater than 0.85, as shown in Table 5.
[0067] Table 5 Discriminant performance of the logistic regression model based on the combination of four key differential metabolites
[0068]
[0069]
[0070] (d) Comparison of the detection performance of three commonly used clinical lung cancer serum protein markers with the detection performance of the discrimination model of the present invention
[0071] The performance of three commonly used serum protein markers for lung cancer was studied, and the results are provided as a reference. The included protein markers and their normal reference ranges are: carcinoembryonic antigen (CEA: 0-5 ng / mL), primarily used for the diagnosis of lung adenocarcinoma; squamous cell carcinoma antigen (SCC: 0-1.5 ng / mL), specifically for the diagnosis of lung squamous cell carcinoma; and soluble cytokeratin 19 fragment (CYFRA21-1: 0-3.3 ng / mL), which has diagnostic value for both lung adenocarcinoma and lung squamous cell carcinoma. A total of 74 patients with complete protein marker test data were included in this study, including 63 patients with lung adenocarcinoma and 11 patients with lung squamous cell carcinoma. The results showed that the sensitivity of the three lung cancer serum protein markers in the 74 patients in this study was limited: the sensitivity of CEA in detecting lung adenocarcinoma was 0.111; the sensitivity of SCC in detecting lung squamous cell carcinoma was 0.364; the sensitivity of CYFRA21-1 in detecting lung adenocarcinoma was 0.159, the sensitivity in detecting lung squamous cell carcinoma was 0.455, and the sensitivity in detecting non-small cell lung cancer (lung adenocarcinoma and lung squamous cell carcinoma) was 0.203. Using a combination of markers, where any marker in the combination exceeded its normal threshold, the sensitivity of CEA combined with CYFRA21-1 in detecting lung adenocarcinoma was 0.222, the sensitivity of SCC combined with CYFRA21-1 in detecting lung squamous cell carcinoma was 0.455, and the sensitivity of the three-marker combination (CEA+SCC+CYFRA21-1) in detecting non-small cell lung cancer (lung adenocarcinoma and lung squamous cell carcinoma) was 0.270. In comparison, the non-small cell lung cancer tissue discrimination model constructed in the present invention based on a combination of four differential metabolites had a sensitivity of 0.889 in lung adenocarcinoma detection, a sensitivity of 0.727 in lung squamous cell carcinoma detection, and an overall (lung adenocarcinoma and lung squamous cell carcinoma) sensitivity of 0.865. In this study, it was superior to the non-small cell lung cancer detection performance of single or combined serum protein markers.
Claims
1. A non-small cell lung cancer tissue discrimination model based on differential metabolite combinations, characterized in that: The differential metabolite combination consists of lactate, phosphoethanolamine, decenoylcarnitine and symmetric dimethylarginine.
2. The non-small cell lung cancer tissue discrimination model based on differential metabolite combinations according to claim 1, characterized in that: The discriminant model is a logistic regression model.
3. A method for constructing a non-small cell lung cancer tissue discrimination model based on differential metabolite combinations, the specific steps are as follows: (1) Collection and processing of tissue samples: Non-small cell lung cancer tumor tissue and paired normal tissue samples were collected and metabolites were extracted; (2) Ultra-high performance liquid chromatography-high-resolution mass spectrometry: Ultra-high performance liquid chromatography-high-resolution mass spectrometry was used to perform non-targeted metabolomics analysis of tissue metabolites to obtain raw mass spectrometry data. Progenesis QI software was used to preprocess the mass spectrometry data, perform peak extraction and peak alignment, and search for metabolite-to-charge ratio matching candidate metabolites based on the human metabolomics database. The mass spectrometry peaks were further processed for contamination removal, blank removal, adduct and source lysate discrimination, QC correction, and segmented normalization to generate a sample metabolite abundance matrix. (3) Metabolite identification: Ultra-high performance liquid chromatography combined with secondary mass spectrometry was used to identify the detected metabolites; (4) Differential metabolite screening and model testing. The specific steps are as follows: (a) The Wilcoxon paired signed-rank test was used to calculate the differences in metabolite abundance between tumor tissues and normal tissues. The Benjamini-Hochberg correction method was used to control the false discovery rate (FDR) to less than 0.
05. In combination with the criterion of fold difference (FC) greater than 1.25 or less than 0.8, differential metabolites between lung cancer tissues and normal tissues were screened. (b) The tumor group and normal group sample data were randomly split into a training set and a validation set at a ratio of 70% and 30%; based on the training set data, the LASSO method was used to screen the key differential metabolites; (c) Based on the identified key differential metabolites, the performance of six discriminant models, including decision tree, random forest, K-nearest neighbor, naive Bayes, support vector machine, and logistic regression, was tested. The models were trained using the training set and validated using the validation set. The robustness and discriminant performance of the models were comprehensively evaluated based on accuracy, sensitivity, specificity, precision, F1 score, and AUC, and the optimal model was selected.
4. The method for constructing a non-small cell lung cancer tissue discrimination model based on differential metabolite combinations according to claim 3, characterized in that: The optimal model is a logistic regression model, and the equation is: ln[p / (1-p)]=0.513+1.189×lactic acid+1.771×phosphoethanolamine-0.490×decenoylcarnitine+1.108×symmetric dimethylarginine; In the above formula, the compound name represents the relative abundance of the compound; when p≤0.468, the sample is judged to be normal tissue, and when p>0.468, the sample is judged to be tumor tissue.
5. The method for constructing a non-small cell lung cancer tissue discrimination model based on differential metabolite combinations according to claim 3, characterized in that: The step (4) (c) further includes: randomly selecting 50% of the samples from the tumor group and the normal group to form a test set, repeating 10 times to obtain 10 different test sets for verifying the robustness of the selected model.
6. Use of the non-small cell lung cancer tissue discrimination model based on differential metabolite combinations according to any one of claims 1-2 in preparing a non-small cell lung cancer tissue detection kit.
7. The use according to claim 6, characterized in that The detection method of the kit includes an ultra-high performance liquid chromatography-high-resolution mass spectrometry method, and the reagents of the kit include a sample metabolite extraction solvent and an ultra-high performance liquid chromatography-high-resolution mass spectrometry detection solvent.
Citation Information
Patent Citations
A combination of lung cancer metabolic markers and its screening method and application
CN117352064B
Application of metabolic marker for diagnosing stages of lung cancer and kit
CN117388495A