A non-small cell lung cancer tissue discrimination model based on differential metabolite combination, a construction method thereof and application thereof

By constructing a logistic regression model through screening key differential metabolite combinations, the problems of long time consumption and insufficient sensitivity in lung cancer diagnosis have been solved, achieving efficient and accurate detection and classification of non-small cell lung cancer tissues, and supporting surgical decisions.

CN121499719BActive Publication Date: 2026-06-02OCEAN UNIV OF CHINA

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
OCEAN UNIV OF CHINA
Filing Date
2025-11-24
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies for lung cancer diagnosis suffer from problems such as long processing time, high subjectivity, and insufficient sensitivity of biomarkers. Furthermore, existing metabolite screening processes are complex and difficult to effectively distinguish between non-small cell lung cancer tissue and normal tissue.

Method used

A non-targeted metabolomics approach based on ultra-high performance liquid chromatography-high resolution mass spectrometry (UPLC-HRMS) was used to screen key differential metabolite combinations. A logistic regression discriminant model was constructed, and combinations of lactate, ethanolamine phosphate, decenoylcarnitine and symmetric dimethylarginine or xanthine, ethanolamine phosphate, N-acetylaspartate, decenoylcarnitine and symmetric dimethylarginine were used to distinguish lung cancer from normal tissues through the logistic regression model.

Benefits of technology

It achieves high sensitivity and specificity in the detection of non-small cell lung cancer tissue, assists surgeons in deciding the extent of resection during tumor resection, reduces model complexity, and improves the accuracy and stability of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121499719B_ABST
    Figure CN121499719B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of biomedical technology, specifically relating to a non-small cell lung cancer (NSCLC) tissue discrimination model based on a combination of differentially expressed metabolites, its construction method, and its application. The differentially expressed metabolite combination consists of lactic acid, ethanolamine phosphate, decenoylcarnitine, and symmetric dimethylarginine, or xanthine, ethanolamine phosphate, N-acetylaspartic acid, decenoylcarnitine, and symmetric dimethylarginine. The discrimination model establishment method involves: collecting and processing lung cancer tissue samples, UPLC-HRMS detection, compound identification, differential metabolite screening, selecting key differentially expressed metabolites using the LASSO method, testing the performance of each model based on the key differentially expressed metabolites, and selecting the optimal model as the logistic regression model. The logistic regression discrimination model of this invention based on 4 or 5 differentially expressed metabolites can effectively distinguish between NSCLC tumor tissue and normal tissue, with accuracy, sensitivity, specificity, F1 score, and AUC all greater than 0.8. This model has a small number of metabolites, demonstrating its potential to assist in the clinical detection of NSCLC tissue and to assist surgeons in determining the extent of tumor resection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biomedical technology, specifically relating to a non-small cell lung cancer tissue discrimination model based on differential metabolite combinations, its construction method, and its application. The logistic regression discrimination model constructed based on differential metabolite combinations can effectively distinguish between non-small cell lung cancer tumor tissue and normal tissue. Background Technology

[0002] Lung cancer has a high incidence and mortality rate worldwide, with non-small cell lung cancer (NSCLC) accounting for approximately 85% (Nicholson AG, Chansky K, Crowley J, et al. The International Association for the Study of Lung Cancer Lung Cancer Staging Project: Proposals for the Revision of the Clinical and Pathologic Staging of Small Cell Lung Cancer in the Forthcoming Eighth Edition of the TNM Classification for Lung Cancer. J Thorac Oncol. 2016;11(3):300-311.). Common diagnostic methods for lung cancer mainly include CT imaging and histopathological analysis. These traditional examination methods are time-consuming and somewhat subjective, often relying on the doctor's experience and professional knowledge. During surgery, doctors need to remove the tumor tissue as completely as possible while preserving as much normal tissue as possible to reduce surgical trauma to the patient's body. Pathological diagnosis can help doctors decide the extent of tumor resection. However, the diagnostic consistency among different pathologists for the same non-small cell lung cancer sample ranged from 67.1% to 89.6% (Paech DC, Weston AR, Pavlakis N, et al. A systematic review of the interobserver variability for histology in the differentiation between squamous and nonsquamous non-small cell lung cancer. J. Thorac. Oncol. 2011;6, 55−63.). Tumor markers are substances produced during the occurrence and proliferation of tumor cells and are helpful in diagnosis.Existing blood biomarkers for lung cancer, such as carcinoembryonic antigen (CEA), have clinical application value, but still suffer from problems such as large individual variability and insufficient sensitivity (Molina R, Auge JM, Escudero JM, et al. Mucins CA 125, CA 19.9, CA 15.3 and TAG-72.3 as tumor markers in patients with lung cancer: comparison with CYFRA 21-1, CEA, SCC and NSE. Tumour Biol. 2008;29(6):371-380.). Finding tissue biomarkers for lung cancer is of great significance for assisting in the clinical detection of tumor tissue and in determining the extent of clinical tumor resection.

[0003] Lung cancer, a disease closely related to metabolism, exhibits significant differences in multiple metabolic pathways in its tumor cells, including glucose metabolism, nucleic acid metabolism, and lipid metabolism (Ke X, Wu H, Chen YX, et al. Individualized pathway activity algorithm identifies oncogenic pathways in pan-cancer analysis. EBioMedicine. 2022;79:104014.). Metabolic reprogramming is a characteristic of tumors, allowing tumor cells to gain advantages in proliferation and metastasis. Revealing the specific metabolites of tumor cells is crucial for discovering new diagnostic biomarkers and improving treatment efficacy. The combination of metabolomics and machine learning provides new technical means for the discovery of metabolic biomarkers and disease diagnosis. Metabolomics data is characterized by a wide variety of metabolites and large differences in abundance values. Machine learning demonstrates significant advantages in processing such complex data, enabling it to quickly identify potential patterns, extract key information, and build predictive models. For example, Chinese patent CN202311708074.6 discloses the application and kit of metabolic biomarkers for diagnosing lung cancer staging. By extracting blood samples from healthy and lung cancer subjects, metabolites are detected, processed, and identified. Through modeling analysis, 19 metabolic biomarkers are screened out, including dehydroepiandrosterone sulfate, 4-hydroxy-1-(3-pyridine)-1-butanone, L-leucine dipeptide, octanoyl-L-carnitine, phosphocholine 34:3e, acylcarnitine 10:0, N-methionine, 5-methionine, phosphocholine 36:3e, fatty acid 22:6, hydrogenated cinnamic acid, acylcarnitine 10:1, hypoxanthine, L-alanyl-L-aspartic acid, iminodiacetic acid, betaine, choline, palmitoyl-L-L-carnitine, and L-glutamic acid. Chinese patent CN202311651833.X discloses a combination of metabolic biomarkers for lung cancer, its screening method, and its application. The combination of metabolic biomarkers for lung cancer includes the following compounds: glucose, urea, creatinine, uric acid, L-phenylalanine, L-valine, L-histidine, L-arginine, and L-leucine. The screening method involves collecting serum or plasma samples from non-lung cancer and lung cancer patients, analyzing the extracted metabolites by MALDI mass spectrometry, and performing data preprocessing to obtain the characteristics of candidate metabolic biomarkers. After training and validating with multiple machine learning models, the metabolic biomarkers in the two machine learning models with the best performance are ranked by importance, and the intersection is taken to determine the potential combination of metabolic biomarkers.

[0004] The existing technologies described above screen for a large number of metabolites, resulting in a complex screening process. Therefore, it is necessary to research and establish a lung cancer discrimination model based on a small number of metabolite combinations with good discriminative performance. Summary of the Invention

[0005] The purpose of this invention is to provide a non-small cell lung cancer tissue discrimination model based on differential metabolite combinations, its construction method and application. These differential metabolite combinations have the potential to assist in the clinical detection and classification of non-small cell lung cancer tissues, and to assist surgeons in deciding the extent of resection during tumor resection.

[0006] To achieve the above objectives, this invention provides a non-small cell lung cancer tissue discrimination model based on differential metabolite combinations. These differential metabolite combinations are either Key Differential Metabolite Combination I or Key Differential Metabolite Combination II. Key Differential Metabolite Combination I consists of lactic acid, phosphate ethanolamine, decenoylcarnitine, and symmetric dimethylarginine; Key Differential Metabolite Combination II consists of xanthine, phosphate ethanolamine, N-acetyl-aspartic acid, decenoylcarnitine, and symmetric dimethylarginine. Based on the logistic regression discrimination model using these two different key differential metabolite combinations, the AUC values ​​for distinguishing lung cancer from normal tissue samples are both greater than 0.85.

[0007] This invention also provides a method for constructing the non-small cell lung cancer tissue discrimination model based on differential metabolite combinations. This method utilizes a non-targeted metabolomics approach based on ultra-high performance liquid chromatography-high resolution mass spectrometry (UPLC-HRMS) to detect and identify metabolites in tumor and normal samples from non-small cell lung cancer patients, screen key differential metabolites, and construct a discrimination model. The specific steps are as follows:

[0008] (1) Collection and processing of tissue samples: Non-small cell lung cancer tumor tissue and paired normal tissue samples were collected separately, and metabolites were extracted;

[0009] (2) UPLC-HRMS detection: Untargeted metabolomics analysis of tissue metabolites was performed using the UPLC-HRMS method to obtain raw mass spectrometry data; the mass spectrometry data were preprocessed using Progenesis QI software, peak extraction and peak alignment were performed, and candidate metabolites were matched based on the human metabolomics database (HMDB) by metabolite charge ratio (m / z); further, the mass spectrometry peaks were decontaminated, blank removed, adducts and source lysates were identified, QC correction and segmented normalization were performed (Zhang J, Zang X, Jiao P, et al. Alterations of Ceramides, Acylcarnitines, GlyceroLPLs, and Amines in NSCLC Tissues. J. Proteome Res. 2024; 23:4343-4358.), and a sample metabolite abundance matrix was generated;

[0010] (3) Metabolite identification: The detected metabolites were identified using UPLC-MS / MS.

[0011] (4) Screening of differential metabolites and model testing, the specific steps are as follows:

[0012] (a) The difference in metabolite abundance between tumor tissue and normal tissue was calculated using the paired signed-rank test. The Benjamini-Hochberg correction method was used to control the false discovery rate (FDR) to be less than 0.05. The differential metabolites between lung cancer tumor tissue and normal tissue were screened based on the criteria of fold difference (FC) greater than 1.25 or less than 0.8.

[0013] (b) The sample data of the tumor group and the normal group were randomly split into training set and validation set at a ratio of 70% and 30%, respectively; based on the training set data, the least absolute shrinkage and selection operator (LASSO) method was used to screen out key differential metabolites.

[0014] (c) Based on the selected key differential metabolites, the performance of six discriminant models—decision tree, random forest, K-nearest neighbors, Naive Bayes, support vector machine, and logistic regression—was tested. The models were trained using the training set and validated using the test set. The robustness and discriminant performance of the models were comprehensively evaluated based on accuracy, sensitivity, specificity, precision, F1 score, and AUC (area under the receiver operating characteristic curve). The optimal model was selected as the logistic regression model, which was used as the tissue discrimination model for non-small cell lung cancer.

[0015] The non-small cell lung cancer tissue discrimination model equation based on key differential metabolite combinations described in this invention is as follows: When the differential metabolite combination is key differential metabolite combination I, the model equation is:

[0016] ln[p / (1-p)] = 0.513 + 1.189 × lactate + 1.771 × ethanolamine phosphate - 0.490 × decenoylcarnitine + 1.108 × symmetric dimethylarginine; the optimal cut-off value was determined to be 0.468 based on the Youden index. When p ≤ 0.468, the sample was judged as normal tissue, and when p > 0.468, the sample was judged as non-small cell lung cancer tumor tissue.

[0017] When the differential metabolite combination is the key differential metabolite combination II, the model equation is:

[0018] ln[p / (1-p)] = 0.976 + 1.240 × xanthine + 1.086 × ethanolamine phosphate + 2.931 × N-acetylaspartic acid - 0.499 × decenoylcarnitine + 0.467 × symmetric dimethylarginine; The optimal cutoff value determined by the Youden index is 0.393. When p ≤ 0.393, the sample is judged as normal tissue, and when p > 0.393, the sample is judged as non-small cell lung cancer tumor tissue.

[0019] In the above equations, the compound names represent the relative abundance of the compounds.

[0020] Step (4) (c) of the present invention further includes: randomly selecting 50% of the samples from the tumor group and the normal group to form a test set, repeating 10 times to obtain 10 different test sets, which are used to verify the robustness of the selected model.

[0021] This invention also provides the application of the aforementioned differential metabolite combination in the preparation of a non-small cell lung cancer tissue detection kit. The detection method of the kit includes UPLC-HRMS, and the reagents of the kit include a sample metabolite extraction solvent and a UPLC-HRMS detection solvent.

[0022] Compared with the prior art, the beneficial effects of this invention are as follows:

[0023] (1) High sensitivity and specificity in detecting non-small cell lung cancer tissue

[0024] To address the limitations of existing biomarkers (such as CEA) in terms of diagnostic sensitivity and the time-consuming and subjective nature of traditional histopathological and imaging methods, this invention provides two key differential metabolite combinations: Combination I and Combination II. Combination I includes four key differential metabolites: lactate, ethanolamine phosphate, decenoylcarnitine, and symmetric dimethylarginine. Combination II includes five key differential metabolites: xanthine, ethanolamine phosphate, N-acetylaspartate, decenoylcarnitine, and symmetric dimethylarginine. Logistic regression discriminant models were then constructed based on each of the two combinations for the detection of non-small cell lung cancer (NSCLC). Stable performance was demonstrated on the training set, validation set, and 10 randomized test sets, with accuracy, sensitivity, specificity, precision, F1 score, and AUC all greater than 0.8. This metabolite combination supports high-throughput detection based on UPLC-HRMS, enabling kit-based applications. It has the potential to assist in the detection and classification of NSCLC tissues and to assist surgeons in deciding the extent of tumor resection during tumor resection.

[0025] (2) Efficient screening method

[0026] To address the challenges of high dimensionality and high noise in metabolomics data, this invention combines LASSO regularization constraints and a 10-fold cross-validation strategy to select two key differential metabolite combinations, significantly reducing model complexity (feature count reduced from 71 and 70 to 4 and 5, respectively). The performance of decision trees, random forests, K-nearest neighbors, Naive Bayes, support vector machines, and logistic regression discriminative models based on these two differential metabolite combinations was then tested. The logistic regression model based on the two key differential metabolite combinations showed the best discriminative performance, achieving AUCs of 0.890 and 0.895 in the training set, respectively, both higher than the best single metabolite (decenoylcarnitine, AUC = 0.807). The AUC > 0.85 in the test set indicates excellent model robustness, and the Hosmer-Lemeshow test (p > 0.05) confirms the model's goodness of fit.

[0027] In summary, this invention screened out two key differential metabolite combinations (containing 4 and 5 metabolites respectively), which are relatively few in number; the logistic regression models based on the two combinations can effectively distinguish between non-small cell lung cancer tumor tissue and normal tissue, with good sensitivity, specificity and accuracy, showing potential in assisting the detection and classification of non-small cell lung cancer tissue, and also has the potential to assist surgeons in deciding the extent of resection during tumor resection. Attached Figure Description

[0028] Figure 1 This is a schematic diagram of the key differential metabolite screening process involved in this invention.

[0029] Figure 2The relative abundance of key differentially expressed metabolites in key differentially expressed metabolite combinations I and II in tumors and normal tissues of non-small cell lung cancer patients is shown. A is lactate, B is phosphoethanolamine, C is decenoylcarnitine, D is symmetric dimethylarginine, E is xanthine, and F is N-acetylaspartate. *** indicates a statistically significant difference between the two groups (p < 0.001). The y-axis represents normalized abundance.

[0030] Figure 3 The structures and MS / MS spectra of key differential metabolites in key differential metabolite combinations I and II are shown.

[0031] Figure 4 The ROC curves are for models based on key differential metabolite combinations I and II, built using different machine learning algorithms. A is the training set of key differential metabolite combination I, B is the validation set of key differential metabolite combination I, C is the training set of key differential metabolite combination II, and D is the validation set of key differential metabolite combination II. Detailed Implementation

[0032] The present invention will be further described in detail below with reference to specific embodiments and accompanying drawings.

[0033] Example 1:

[0034] This embodiment relates to a method for constructing a non-small cell lung cancer tissue discrimination model based on differential metabolite combinations. The specific steps are as follows:

[0035] 1. Collection and processing of tissue samples

[0036] (a) Collection of tissue samples

[0037] This study has been approved by the Scientific Ethics Committee of the Academic Committee of Ocean University of China (OUC-HM-2022-005) and the Ethics Committee of Qingdao Municipal Hospital (2022-Y014). Tumor lesion tissue (as tumor group samples) and distal normal lung tissue (as normal group samples) were collected from 104 patients with non-small cell lung cancer, including 88 cases of lung adenocarcinoma and 16 cases of lung squamous cell carcinoma. The collected tissue samples were first moistened with gauze to remove surface blood, then rapidly transferred to liquid nitrogen for short-term preservation, and finally transferred to a -80°C freezer for long-term storage.

[0038] Table 1 Clinical information of patients with non-small cell lung cancer

[0039]

[0040] (b) Extraction of tissue metabolites

[0041] Tissue samples were thawed on ice at room temperature and cut into small pieces with scissors. 100 mg of the chopped tissue sample was placed in a homogenization tube, and 1000 μL of pre-cooled 80% (v / v) methanol aqueous solution was added for homogenization (60 Hz, 90 seconds). The sample was then vortexed for 1 minute and incubated at -80°C for 4 hours. After incubation, the sample was centrifuged at 16000 g for 10 minutes at 4°C, and the supernatant was transferred to a new EP tube for concentration using a vacuum centrifuge. After drying, 100 μL of water was added to reconstitute the sample, and the mixture was vortexed again for 15 seconds and centrifuged at 16000 g for 10 minutes. Finally, the supernatant was transferred to a sample vial. Additionally, 100 μL of water was used as a blank control instead of the tissue sample to eliminate background ion interference. Quality control (QC) samples were prepared by mixing the samples to monitor and correct for changes in instrument response.

[0042] 2. UPLC-HRMS testing

[0043] (a) Untargeted metabolomics analysis of the extracted metabolites was performed using UPLC-LTQ Orbitrap XL to obtain raw mass spectrometry data.

[0044] The high-performance liquid chromatography (HPLC) conditions were as follows: the column was an ACQUITY UPLC® BEH C18 column (2.1 × 50 mm, 1.7 μm), the injection volume was 10 μL, the column temperature was 35℃, mobile phase A was an aqueous solution containing 0.1% (v / v) acetic acid, mobile phase B was acetonitrile solution, and the gradient elution conditions were: 0 min (100% A, 0.25 mL / min), 1 min (90% A, 0.25 mL / min), 2.5 min (85% A, 0.25 mL / min), 4 min (78% A, 0.25 mL / min), 6 min (62% A, 0.25 mL / min), 9 min (35% A, 0.25 mL / min), 12 min (20% A, 0.25 mL / min), 16 min (0% A, 0.30 mL / min), and 18 min (0% A, 0.45 mL / min).

[0045] The LTQ Orbitrap XL mass spectrometer employs ESI ionization and performs detection in both positive and negative ion modes. Spray voltage: 4 kV for positive and 2.5-2.8 kV for negative; capillary temperature: 300°C; sheath gas flow rate: 40 arb; auxiliary gas flow rate: 10 arb; capillary voltage: 20 V for positive and -40 V for negative; tube lens voltage: 35 V for positive and -80 V for negative; full scan mode: 50-1000 m / z; resolution: 60000 FWHM.

[0046] (b) Mass spectrometry data processing

[0047] The obtained mass spectrometry data were preprocessed using Progenesis QI software, with the adduct ion set to [M+H] in positive ion mode. + [M+Na] + [M+K] + [M+2H] 2+ [M+Na+H] 2+ [M+2Na] 2+ [M+3H] 3+ [M+2H+Na] 3+ and [M+2Na+H] 3+ The negative ion concentration is set to [MH]. - [M-2H] 2- and [M-3H] 3- Peak extraction and alignment were performed, and candidate metabolites were matched based on their m / z ratios using the Human Metabolomics Database (HMDB). Further processing of the mass spectrometry peaks included decontamination, blank removal, adduct and source lysate discrimination, QC correction, and segmented normalization (Zhang J, Zang X, Jiao P, et al. Alterations of Ceramides, Acylcarnitines, GlyceroLPLs, and Amines in NSCLC Tissues. J. Proteome Res. 2024; 23:4343-4358.), generating a metabolite abundance matrix for the samples.

[0048] 3. Metabolite identification

[0049] Compound identification was performed using UPLC combined with secondary mass spectrometry (MS / MS). Specifically, after UPLC separation, fragment ions of metabolites were obtained by MS / MS at a resolution of 30,000 (FWHM) in MS / MS mode. These fragment ions were then compared with spectra in the databases HMDB and MassBank. The compounds were identified with high confidence by manually analyzing the fragmentation patterns and fragment ion structures.

[0050] 4. Differential metabolite screening and model testing, the specific steps are as follows:

[0051] (a) Screening for differentially expressed metabolites in tumor tissue compared to normal tissue

[0052] For the metabolites identified in step 3, the p-values ​​of tumor tissue versus paired normal tissue were first calculated using the Wilcoxon paired-pairs signed-rank test. Then, the false discovery rate (FDR) was calculated using the Benjamini-Hochberg multiple test correction method. Simultaneously, the fold change (FC) between the tumor and normal groups was calculated based on the median metabolite abundance. A statistical significance threshold of FDR < 0.05 (FC > 1.25 or FC < 0.8) was used. Ultimately, 71 differentially expressed metabolites were identified.

[0053] To evaluate the classification performance of the 71 differentially expressed metabolites for non-small cell lung cancer tumor tissues and normal tissues, ROC (Receiver Operating Characteristic) curves were constructed, and sensitivity, specificity, and area under the ROC curve (AUC) were calculated. Four metabolites were found to have sensitivity, specificity, and AUC greater than 70%, including octanoylcarnitine, guanosine monophosphate, decenoylcarnitine, and adenosine monophosphate (Table 2). Since individual metabolites typically lack sufficient sensitivity and specificity, their value in clinical applications is limited.

[0054] Table 2. Discriminative properties of individual metabolites in non-small cell lung cancer tissues

[0055]

[0056] Note: * indicates that the median of normal samples is 0 during the FC (tumor / normal) calculation.

[0057] (b) Option 1: Performance testing of the discriminant model based on key differential metabolite combinations screened by LASSO.

[0058] After z-score standardization of the abundance matrix of the 71 differentially expressed metabolites obtained in step 4(a), the dataset was randomly split into training and validation sets at a ratio of 70% to 30%. Based on the training set data, LASSO regression was used to screen variables, and the optimal regularization parameter was selected using ten-fold cross-validation. Four key differentially expressed metabolites were obtained: lactic acid, phosphate ethanolamine, decenoylcarnitine, and symmetric dimethylarginine, which were designated as key differentially expressed metabolite combination I. (See Table 3 and...) Figure 2 As shown, compared with normal tissue, lactate, phosphoethanolamine, and symmetric dimethylarginine were significantly elevated in non-small cell lung cancer tumor tissue, while decenoylcarnitine was significantly decreased. The structures and MS / MS spectra of the four metabolites are shown below. Figure 3 As shown.

[0059] Table 3. Four key differential metabolites in Scheme 1

[0060]

[0061] Based on the four key differentially expressed metabolites in Table 3, and using tumor and normal tissue as response variables, six machine learning algorithms—Decision Tree (DT), Random Forest (RF), K-Nearest Neighbors (KNN), Naive Bayes (NB), Support Vector Machine (SVM), and Logistic Regression (LR)—were used to construct discriminative models and evaluate their performance. Modeling was performed using the training set, and validation was conducted using the validation set. The results are shown in Table 4 and... Figure 4 These algorithms exhibited similar classification performance, with validation set AUC values ​​ranging from 0.841 to 0.882, indicating that the differential metabolite combination has strong generalization ability and stable discrimination effect. Considering accuracy, sensitivity, specificity, precision, F1 score, and AUC, the logistic regression model was ultimately selected as the optimal discrimination model for the key differential metabolite combination I.

[0062] Table 4 shows the discriminative performance of models based on key differential metabolite combination I built using different machine learning algorithms.

[0063]

[0064] A tissue discrimination model for non-small cell lung cancer was constructed based on key differential metabolite combination I and logistic regression algorithm. The resulting logistic regression equation is as follows:

[0065] ln[p / (1-p)] = 0.513 + 1.189 × lactic acid + 1.771 × ethanolamine phosphate - 0.490 × decenoylcarnitine + 1.108 × symmetric dimethylarginine;

[0066] In the above formula, the compound name represents the relative abundance of the compound. The optimal cut-off value was determined to be 0.468 based on the Youden index. When p ≤ 0.468, the sample was classified as normal tissue; when p > 0.468, the sample was classified as non-small cell lung cancer tumor tissue. The training set data showed an accuracy of 0.849, sensitivity of 0.836, specificity of 0.863, precision of 0.859, F1 score of 0.847, and AUC of 0.890 in this model. The Hosmer-Lemeshow test was used to assess the goodness of fit of the logistic regression model, with p = 0.109, indicating a good fit. When the validation set was input into the model, the accuracy was 0.871, sensitivity was 0.839, specificity was 0.903, precision was 0.897, F1 score was 0.867, and AUC was 0.874. To further verify the robustness of the model, 50% of the samples were randomly selected to form a test set, and this was repeated 10 times to obtain 10 different test sets. The data from these 10 test sets were input into the model, and the corresponding accuracy, sensitivity, specificity, precision, and F1 score were all greater than 0.8, and the AUC was greater than 0.85, as shown in Table 5.

[0067] Table 5 Discriminant performance of the logistic regression model based on key differential metabolite combination I.

[0068]

[0069] (c) Option 2: After removing lactate, test the performance of the discriminant model based on the key differential metabolite combinations screened by LASSO.

[0070] Of the four key differential metabolites identified in Scheme 1, lactate exhibited background residue in the LC-MS system, potentially leading to inaccurate quantitative data. To ensure the reliability of the modeling data, this invention proposes Scheme 2, which, after removing lactate, analyzes the remaining 70 metabolites using the same construction process as Scheme 1. LASSO regression screening yielded five compounds: xanthine, ethanolamine phosphate, N-acetyl-aspartic acid, decenoylcarnitine, and symmetric dimethylarginine, which were designated as key differential metabolite combination II. (See Table 6 and...) Figure 2 As shown, compared with normal tissues, xanthine, ethanolamine phosphate, N-acetylaspartate, and symmetric dimethylarginine were significantly elevated in non-small cell lung cancer tumor tissues, while decenoylcarnitine was significantly decreased. Their structures and MS / MS spectra are shown below. Figure 3 As shown in Table 7, based on the above five metabolites, six algorithms—decision tree, random forest, K-nearest neighbors, Naive Bayes, support vector machine, and logistic regression—were used to construct a discriminative model. The model results are shown in Table 7. Figure 4 As shown. Considering all performance indicators, logistic regression was ultimately determined to be the optimal model for the key differential metabolite combination II.

[0071] A discriminant model was constructed based on the key differential metabolite combination II and the logistic regression algorithm, and the resulting logistic regression equation is as follows:

[0072] ln[p / (1-p)] = 0.976 + 1.240 × xanthine + 1.086 × ethanolamine phosphate + 2.931 × N-acetylaspartic acid - 0.499 × decenoylcarnitine + 0.467 × symmetric dimethylarginine.

[0073] The optimal cutoff value was 0.393, used to distinguish normal tissue (p ≤ 0.393) from tumor tissue (p > 0.393). The logistic regression model achieved an accuracy of 0.849 and an AUC of 0.895 on the training set, and an accuracy of 0.871 and an AUC of 0.879 on the validation set. The Hosmer-Lemeshow test showed a good model fit (p = 0.306). Further stability testing was conducted using 10 random sampling runs; all performance metrics on the test sets were greater than 0.8, with AUCs all greater than 0.85 (Table 7).

[0074] Table 6. Five key differential metabolites in Key Differential Metabolite Combination II

[0075]

[0076] Table 7. Discriminant performance of models based on five key differentially expressed metabolites built using different machine learning algorithms in Scheme 2.

[0077]

[0078] (d) Comparison of the detection performance of three commonly used clinical lung cancer serum protein markers with the detection performance of the discriminant model of this invention.

[0079] The detection performance of three commonly used clinical serum protein biomarkers for lung cancer was studied, and the results are provided for reference. The included protein biomarkers and their normal reference ranges are as follows: Carcinoembryonic antigen (CEA: 0-5 ng / mL), mainly used for the diagnosis of lung adenocarcinoma; Squamous cell carcinoma antigen (SCC: 0-1.5 ng / mL), specifically for the diagnosis of lung squamous cell carcinoma; and Soluble cytokeratin 19 fragment (CYFRA21-1: 0-3.3 ng / mL), which has diagnostic value for both lung adenocarcinoma and lung squamous cell carcinoma. A total of 74 patients with complete protein biomarker detection data were included in this study, including 63 cases of lung adenocarcinoma and 11 cases of lung squamous cell carcinoma. The results showed that the detection sensitivity of the three serum protein biomarkers for lung cancer in the 74 patients in this study was limited: CEA had a sensitivity of 0.111 in lung adenocarcinoma; SCC had a sensitivity of 0.364 in lung squamous cell carcinoma; and CYFRA21-1 had a sensitivity of 0.159 in lung adenocarcinoma, 0.455 in lung squamous cell carcinoma, and 0.203 in non-small cell lung cancer (both adenocarcinoma and squamous cell carcinoma). Using a biomarker combination method, where any biomarker exceeding its normal threshold was considered positive, the sensitivity of CEA combined with CYFRA21-1 in lung adenocarcinoma was 0.222, SCC combined with CYFRA21-1 in lung squamous cell carcinoma was 0.455, and the sensitivity of the three-indicator combination (CEA + SCC + CYFRA21-1) in non-small cell lung cancer (both adenocarcinoma and squamous cell carcinoma) was 0.270. In comparison, the non-small cell lung cancer tissue discrimination models constructed by Scheme 1 and Scheme 2 of this invention both have a sensitivity of 0.889 in lung adenocarcinoma detection and 0.727 in lung squamous cell carcinoma detection, with an overall sensitivity of 0.865 (for both lung adenocarcinoma and lung squamous cell carcinoma). In this study, these models outperform the non-small cell lung cancer detection performance of using serum protein markers alone or in combination.

Claims

1. A non-small cell lung cancer tissue discrimination model based on differential metabolite combinations, characterized in that, The differential metabolite combination is either Key Differential Metabolite Combination I or Key Differential Metabolite Combination II. Key Differential Metabolite Combination I consists of lactate, ethanolamine phosphate, decenoylcarnitine, and symmetric dimethylarginine. Key Differential Metabolite Combination II consists of xanthine, ethanolamine phosphate, N-acetylaspartic acid, decenoylcarnitine, and symmetric dimethylarginine. The discriminant model is a logistic regression model.

2. The method for constructing a non-small cell lung cancer tissue discrimination model based on differential metabolite combinations as described in claim 1, comprising the following specific steps: (1) Collection and processing of tissue samples: Non-small cell lung cancer tumor tissue and paired normal tissue samples were collected separately, and metabolites were extracted; (2) Ultra-high performance liquid chromatography-high resolution mass spectrometry detection: Ultra-high performance liquid chromatography-high resolution mass spectrometry was used to perform non-targeted metabolomics analysis of tissue metabolites to obtain raw mass spectrometry data; Progenesis QI software was used to preprocess the mass spectrometry data, perform peak extraction and peak alignment, and search for candidate metabolites with matching metabolite charge ratios based on the human metabolomics database; further, the mass spectrometry peaks were decontaminated, blank removed, adducts and source lysates were identified, QC correction and segmented normalization were performed to generate a sample metabolite abundance matrix; (3) Metabolite identification: Ultra-high performance liquid chromatography combined with secondary mass spectrometry was used to identify the detected metabolites; (4) Screening of differential metabolites and model testing, the specific steps are as follows: (a) The Wilcoxon paired signed-rank test was used to calculate the difference in metabolite abundance between tumor tissue and normal tissue. The Benjamini-Hochberg correction method was used to control the false discovery rate (FDR) to be less than 0.

05. The differential metabolites between lung cancer tumor tissue and normal tissue were screened by combining the criteria of fold difference (FC) greater than 1.25 or less than 0.

8. (b) The sample data of the tumor group and the normal group were randomly split into training set and validation set at a ratio of 70% and 30%, respectively; based on the training set data, the LASSO method was used to screen out key differential metabolites; (c) Based on the selected key differential metabolites, the performance of six discriminative models—decision tree, random forest, K-nearest neighbors, Naive Bayes, support vector machine, and logistic regression—was tested. The models were trained using the training set and validated using the validation set. The robustness and discriminative performance of the models were comprehensively evaluated based on accuracy, sensitivity, specificity, precision, F1 score, and AUC, and the optimal model was selected.

3. The construction method according to claim 2, characterized in that, The optimal model is a logistic regression model. When the differential metabolite combination is the key differential metabolite combination I, the model equation is: ln[p / (1-p)] = 0.513 + 1.189 × lactate + 1.771 × ethanolamine phosphate - 0.490 × decenoylcarnitine + 1.108 × symmetric dimethylarginine; when p ≤ 0.468, the sample is judged as normal tissue, and when p > 0.468, the sample is judged as tumor tissue; When the differential metabolite combination is the key differential metabolite combination II, the model equation is: ln[p / (1-p)] = 0.976 + 1.240 × xanthine + 1.086 × ethanolamine phosphate + 2.931 × N-acetylaspartic acid - 0.499 × decenoylcarnitine + 0.467 × symmetric dimethylarginine; when p ≤ 0.393, the sample is judged as normal tissue, and when p > 0.393, the sample is judged as tumor tissue; In the above equations, the compound names represent the relative abundance of the compounds.

4. The construction method according to claim 2, characterized in that, Step (4) (c) further includes: randomly selecting 50% of the samples from the tumor group and the normal group to form a test set, repeating this process 10 times to obtain 10 different test sets, which are used to verify the robustness of the selected model.

5. The application of the non-small cell lung cancer tissue discrimination model based on differential metabolite combinations as described in claim 1 in the preparation of a non-small cell lung cancer tissue detection kit.

6. The application according to claim 5, characterized in that, The detection method of the kit includes ultra-high performance liquid chromatography-high resolution mass spectrometry, and the reagents of the kit include sample metabolite extraction solvent and ultra-high performance liquid chromatography-high resolution mass spectrometry detection solvent.