Combined marker for prognosis prediction of lung adenocarcinoma, risk scoring model and construction method and application thereof
By constructing a prognostic model based on a combined biomarker of 12 genes and a random survival forest, the problem of insufficient accuracy in predicting the prognosis of lung adenocarcinoma was solved, enabling high-precision stratification of patient survival risk and support for individualized treatment. The upregulation effect of ANGPTL4 in lactic acid environment was also verified.
Patent Information
- Application Number
- CN202511591165.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-03
- Publication Date
- 2026-02-06
AI Technical Summary
Existing technologies lack accuracy in predicting the prognosis of lung adenocarcinoma. Traditional biomarkers such as PD-L1 expression levels and tumor mutation burden have limited predictive power. There is an urgent need to introduce multi-omics approaches to reveal molecular subtypes and operable pathways, especially since research on the immune regulatory mechanism of lactation in the tumor microenvironment is insufficient.
A set of joint biomarkers based on 12 genes, including 9 risk genes and 3 protective genes, was constructed. Key genes were screened out through single-cell analysis, lactation activity assessment and machine learning model, and a random survival forest prognostic model was constructed for risk scoring and prognosis prediction of lung adenocarcinoma patients.
This study achieved high-precision stratification of survival risk in patients with lung adenocarcinoma, providing a reliable reference for clinical treatment strategies, improving prediction accuracy and model stability, identifying new lactation-related genes, and validating the significant upregulation of ANGPTL4 in a lactic environment, supporting the potential link to the lactation process.
Smart Images

Figure CN121472406A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedical and molecular diagnostic technology, specifically relating to a set of combined biomarkers, risk scoring models, and their construction methods and applications for predicting the prognosis of lung adenocarcinoma based on lactation-related genes. Background Technology
[0002] Lung adenocarcinoma (LUAD) is the most common subtype of non-small cell lung cancer, accounting for approximately 45.6% of lung cancer cases in men and 59.7% in women worldwide. In recent years, the incidence of LUAD has been steadily increasing, particularly among women and non-smokers. Because early-stage lesions often lack typical clinical symptoms, patients are frequently diagnosed at middle or late stages, resulting in a generally poor prognosis. Although the five-year survival rate for early-stage LUAD patients is significantly higher than that for late-stage patients, overall survival improvement remains limited.
[0003] Currently, treatment options for LUAD primarily include surgical resection, chemotherapy, radiotherapy, targeted therapy, and immunotherapy based on immune checkpoint inhibitors (ICIs). Immune checkpoint inhibitors targeting the PD-1 / PD-L1 pathway (such as nivolumab and pembrolizumab) have significantly improved survival in some patients without operable driver gene mutations. However, LUAD generally presents with a "cold tumor" immunophenotype, characterized by insufficient CD8⁺ T cell infiltration and low PD-L1 expression, limiting the overall efficacy of ICIs. Furthermore, long-term benefits are only observed in a minority of patients, and resistance mechanisms involve the immunosuppressive tumor microenvironment (such as regulatory T cell infiltration), B cell dysfunction, and abnormal epigenetic regulation. To improve efficacy, the academic community has explored combination therapy regimens of ICIs with chemotherapy, anti-angiogenic drugs, or targeted drugs, but these have limited efficacy improvements and are often accompanied by increased toxicity. Furthermore, the accuracy of traditional biomarkers such as PD-L1 expression levels and tumor mutation burden in predicting treatment efficacy is limited, suggesting the need to introduce multi-omics approaches to reveal molecular subtypes and actionable pathways, thereby promoting the development of personalized treatment strategies.
[0004] In recent years, lactation has been discovered as a novel post-translational modification of lysine. Its mechanism involves the lactyl group binding to histones or non-histone proteins via acylation, thereby affecting gene transcription and protein function. In cancer cells, metabolic reprogramming shifts towards aerobic glycolysis (the Warburg effect), leading to a large accumulation of lactate and inducing lactation modification, thus forming a metabolic-epigenetic feedback loop that promotes adaptive growth and malignant progression of tumor cells. Previous studies have shown a positive correlation between lactation levels and glycolytic flux, with lactate concentration directly determining the degree of modification. In the tumor microenvironment, lactation not only serves as an integrative link in metabolic signaling and cell fate regulation but also participates in immune regulation, DNA repair, and cellular metabolism, driving tumor progression, metastasis, and drug resistance.
[0005] However, the role of lactation in LUAD (lumbar inflammatory disease) remains poorly understood, particularly regarding its heterogeneity at the cellular level and its immunomodulatory mechanisms within the tumor microenvironment. Therefore, a deeper understanding of the biological significance and clinical value of lactation in LUAD remains a crucial scientific challenge. Simultaneously, there is an urgent need to screen and identify novel lactation-related genes and, based on this, construct reliable prognostic models to achieve more accurate predictions of LUAD patient outcomes and provide a basis for personalized treatment. Summary of the Invention
[0006] The purpose of this invention is to provide a combined biomarker, a risk scoring model, and a method for constructing the model for predicting the prognosis of lung adenocarcinoma, in order to address the shortcomings of current prognosis prediction methods for lung adenocarcinoma.
[0007] The present invention adopts the following technical solution:
[0008] Firstly, it provides combined biomarkers for predicting the prognosis of lung adenocarcinoma, including the following 12 genes:
[0009] ANGPTL4, ITGA6, SOD1, CCL20, FKBP3, DECR1, SEMA3C, TRIM28, VEGFC, TNNC2, CRTAC1, HGF.
[0010] These include nine risk genes (ANGPTL4, ITGA6, SOD1, CCL20, FKBP3, DECR1, SEMA3C, TRIM28, VEGFC) and three protective genes (TNNC2, CRTAC1, HGF). These genes originate from: the lactation gene set (SOD1, FKBP3, DECR1, TRIM28), receptor-ligand pairs (ANGPTL4, ITGA6, SEMA3C, VEGFC, HGF), and characteristic genes of highly lactated cells (CCL20, CRTAC1, TNNC2).
[0011] The aforementioned joint markers were obtained through the following method.
[0012] 1. Data Collection and Single-Cell Analysis
[0013] The single-cell transcriptome dataset GSE131907 from LUAD patients and normal lung tissue was obtained from a public database. Quality control, normalization, dimensionality reduction, clustering, and cell type annotation were performed on the single-cell data to obtain the main cell subpopulations, including epithelial, stromal, and immune cells, and their sub-subpopulations.
[0014] Based on 332 published lactation-related genes, the lactation activity score (AUC) of each cell was calculated using AUCell to quantify the lactation level of each cell. Cell subpopulations with high lactation levels were labeled according to their AUC scores within each subpopulation and threshold classification. Results showed significantly elevated lactation levels in epithelial cells (epithelial cells 10, 3, 4, and 5) and stromal cells (lymphatic endothelial cells, fibroblast 3, and myofibroblast) in tumor tissues. Subsequently, the Top 20 characteristic genes were extracted from highly lactated cells as one of the candidate input features for modeling.
[0015] 2. Receptor-ligand pair analysis of highly lactated cells and immune cells
[0016] CellChat was used to perform ligand-receptor interaction analysis on annotated cell subpopulations, and significant interaction pairs between hyperlactated epithelial cells and immune cells were screened at a statistical significance threshold of p < 0.01. Corresponding genes were extracted and added to a candidate feature set.
[0017] 3. Candidate gene integration and machine learning modeling
[0018] In the initial modeling process, we integrated three types of genes as input features: (1) the reported set of lactation-related genes; (2) the top 20 characteristic genes of highly lactated cells; and (3) genes with significant receptor-ligand pairs between highly lactated cells and immune cells.
[0019] By screening 101 machine learning models, 12 key genes closely related to the lactation process were finally identified.
[0020] Subsequently, we used these 12 genes as new input features to conduct secondary modeling. The model performed well on both the training set (TCGA-LUAD) and the independent validation set (GSE87340), with an AUC of 0.92 on the training set and 0.71 on the test set. The C-index also showed high predictive performance. Based on the risk score, patient survival stratification significantly distinguished patients with different prognoses on both the training set (p<0.001, HR=0.03) and the test set (p=0.036, HR=0.16).
[0021] PCR experiments verified that the expression trends of the screened genes in tumor samples were consistent with the risk or protective attributes in the model; that is, high expression corresponded to risk genes, and low expression corresponded to protective genes. Furthermore, in a cellular hypoxia model simulating a lactic acid environment, the results showed that ANGPTL4 expression gradually increased with increasing hypoxia time, further validating the potential link between this gene and the lactation process.
[0022] In a second aspect, based on the aforementioned combined marker genes, a risk scoring model for predicting the prognosis of lung adenocarcinoma is provided, which is constructed based on the aforementioned combined markers.
[0023] A third aspect of the present invention provides a method for constructing the above-mentioned prognostic risk scoring model for patients with lung adenocarcinoma, comprising the following steps:
[0024] (a) Obtain tumor tissue samples from patients with lung adenocarcinoma and detect the expression levels of the above 12 genes;
[0025] (b) Input the patient’s gene expression level data into the prognostic model trained based on the Random Survival Forest (RSF) algorithm, and calculate the mean of the patient’s Cumulative Hazard Function (CHF) as the risk score;
[0026] (c) Patients are stratified by risk based on a preset cutoff value of 68.67: if the risk score is below 68.67, they are predicted to be in the low-risk group with a better prognosis; if the risk score is above or equal to 68.67, they are predicted to be in the high-risk group with a poor prognosis.
[0027] In a fourth aspect, the present invention also provides a diagnostic reagent for predicting the prognosis of lung adenocarcinoma, comprising a primer set or probe set for detecting the expression levels of the aforementioned combined marker genes.
[0028] A fifth aspect of the invention also provides a diagnostic system for predicting the prognosis of lung adenocarcinoma, comprising:
[0029] (1) Sequencing module: used to collect tumor tissue samples from patients and detect the expression levels of the above-mentioned genes;
[0030] (2) Output prediction module: used to input the test results into the random survival forest prognostic model, calculate the patient's risk score, perform risk stratification based on the score, and finally output the prognostic result of overall survival (OS).
[0031] By implementing the above technical solution, the present invention has the following beneficial effects:
[0032] This invention, based on sequencing data from a large number of lung adenocarcinoma samples, systematically screened a group of lactation-related genes closely associated with overall survival and constructed a risk scoring model. This model can accurately stratify the survival risk of lung adenocarcinoma patients, providing a reliable reference for clinical treatment strategy development and individualized management. Furthermore, the newly discovered lactation-related genes are beneficial for understanding the lactation process in lung adenocarcinoma.
[0033] Specifically, this is reflected in:
[0034] Constructing a high-performance prognostic model related to lactation:
[0035] This invention integrates known lactation genes, genes characteristic of highly lactated cells, and genes showing significant receptor-ligand pairings between highly lactated cells and immune cells. Using 101 machine learning algorithms, 12 key genes were selected, and secondary modeling was performed. The constructed random survival forest prognostic model showed the best performance. The model demonstrated high predictive performance on both the training set (TCGA-LUAD) and the independent test set (GSE87340) (training set AUC=0.92, test set AUC=0.71), and the C-index also showed good consistency. It accurately predicted the overall survival of patients and achieved significant stratification between high-risk and low-risk patients.
[0036] Includes novel lactation-related genes:
[0037] By incorporating characteristic genes of highly lactated cells and receptor-ligand pair genes between highly lactated cells and immune cells, this invention can identify and integrate more potential lactation-related genes when the study of the lactation process is insufficient, expand the dimensions of model information, improve prediction accuracy, and provide a reference for the study of lactation-related mechanisms.
[0038] Improving model stability and clinical applicability:
[0039] The 12 key genes selected can be detected in clinical specimens using conventional PCR methods. The model can directly input patient gene expression data and generate risk scores, enabling rapid and quantitative assessment of survival risk. Through risk score stratification (cutoff value 68.67), patients can be reliably classified into high-risk or low-risk groups, providing an operational reference for clinical treatment decisions and individualized management.
[0040] Identify new lactation-related genes:
[0041] Hypoxic cell model experiments showed that ANGPTL4 was significantly upregulated in a lactic acid environment, validating its potential link with the lactation process. This not only improves the predictive accuracy of lung adenocarcinoma prognosis but also provides a scientific basis for further research on the role of lactation in tumor immune escape.
[0042] Experimental verification shows consistency with predictions:
[0043] PCR experiments validated that high expression of risk genes in the model corresponds to high risk, while low expression of protective genes corresponds to low risk, consistent with the model's predictions, further demonstrating the model's reliability and reproducibility.
[0044] The source is clear and the reproducibility is high:
[0045] All 12 characteristic genes were derived from public databases and literature reports. The screening process was standardized and can be reproduced in different research or clinical settings, making it easy to promote and apply. Attached Figure Description
[0046] Figure 1 The results of secondary modeling based on 101 machine learning algorithms show the AUC values of the lactation-related prognostic model under different algorithms.
[0047] Figure 2 The results of secondary modeling based on 101 machine learning algorithms are shown, along with the corresponding C-index values.
[0048] Figure 3 To train Kaplan-Meier overall survival curves for patients in different risk groups.
[0049] Figure 4 To test the Kaplan-Meier overall survival curves of patients in different risk groups.
[0050] Figure 5 The graph shows the results of a univariate Cox regression analysis used to determine the impact of 12 key genes on the prognosis of lung adenocarcinoma.
[0051] Figure 6 This image shows the results of detecting the mRNA expression levels of 12 genes in clinical tissue samples.
[0052] Figure 7The mRNA expression levels of glycolysis-related genes in lung adenocarcinoma cell lines under hypoxic conditions (4 h, 8 h, and 12 h) were measured: (A) HK1; (B) PKM2; (C) G6PD; (D) LDHA; (E) The expression of the potential lactation-related gene ANGPTL4 was detected under the same conditions. Five independent biological replicates were performed for each gene. Detailed Implementation
[0053] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0054] This specific embodiment is merely an explanation of the present invention and is not intended to limit the invention. After reading this specification, those skilled in the art can make modifications to this embodiment without contributing any inventive step, but such modifications are protected by patent law as long as they are within the scope of the claims of the present invention.
[0055] To verify the effectiveness of the lung adenocarcinoma prognostic risk scoring model based on lactation genes proposed in this invention, the applicant conducted systematic bioinformatics analysis and experimental verification based on publicly available datasets. The specific implementation process is as follows:
[0056] 1. Data Collection
[0057] Bulk RNA expression matrices were obtained from the Cancer Genome Atlas (TCGA, https: / / portal.gdc.cancer.gov / ), including transcriptomic data from 568 LUAD patients and their corresponding clinical follow-up information. The external validation cohort was selected from the GSE87340 dataset in the GEO database (https: / / www.ncbi.nlm.nih.gov / geo / ), containing tumor tissue and paired adjacent normal lung tissue samples from 27 LUAD patients, with complete clinical annotations. Single-cell transcriptomic data were obtained from the GEO dataset GSE131907, including 11 untreated LUAD tumor samples and 11 normal lung tissue samples. The lactation gene set was derived from published literature (PMID: 37242427; PMID: 35761067), comprising 332 genes.
[0058] 2. Preprocessing of single-cell transcriptome data
[0059] Single-cell transcriptome data were processed using the R packages Seurat (v5.1.0) and Harmony (v1.2.1). Quality control criteria were: >300 genes detected per cell and <20% mitochondrial genes. The expression matrix was then normalized, and 2,000 hypervariable genes were screened. Mitochondrial genes, hemoglobin genes, and ribosomal genes were removed to reduce technical bias.
[0060] Data scaling was applied, and regression analysis was used to remove confounding effects such as nCount_RNA, nFeature_RNA, and cell cycle stage. Principal component analysis (PCA) was performed using the first 30 principal components, and batch effect correction was performed using Harmony. Clustering parameters were set to 40 k-nearest neighbors and 0.1 resolution, and two-dimensional visualization was achieved using Uniform Manifold Approximation and Projection (UMAP). Cell type annotation was based on reported classic marker genes (such as KRT7 and SFTPC for epithelial cells, ACTA2 and TAGLN for myofibroblasts, and CD3E for T cells), and validated using the CellMarker 2.0 database.
[0061] To further investigate the functional status of key cell subpopulations, epithelial cells, stromal cells, myeloid cells, and lymphocytes were extracted and subjected to fine-grained subpopulation cluster analysis.
[0062] 3. Evaluation of lactation activity
[0063] The R package AUCell (v1.24.0) was used to assess the lactation activity of different cell subpopulations. This method first sorts the expression profiles of each cell in descending order of gene expression level, and then calculates an AUC score based on a set of 332 lactation-related genes to quantify the enrichment of this gene set in a single cell. A higher AUC value indicates stronger lactation activity. Results showed significantly increased lactation levels in epithelial cells (epithelial cells 10, 3, 4, 5) and stromal cells (lymphaticendothelial cells, fibroblast 3, myofibroblast) in tumor tissue. Subsequently, the Top 20 characteristic genes were extracted from highly lactated cells as one of the candidate input features for modeling.
[0064] 4. Cell communication analysis
[0065] To explore the potential interactions between lactate-overexpressing epithelial cell subsets and immune cells, a cell communication network was constructed using CellChat (v1.6.1). The input consisted of annotated cell subsets and information on the high lactation status. Ligand-receptor interaction pairs significantly enriched between high lactate-overexpressing epithelial cells and T cells, as well as myeloid cells, were screened (P < 0.01, permutation test). Analysis revealed that signaling pathways such as ANGPTL4–ITGA6, VEGFC–VEGFR, and HGF–cMET were significantly active in the high lactate tumor microenvironment. Genes corresponding to these significant interaction pairs were included in a candidate feature set for subsequent machine learning modeling.
[0066] 5. Prognostic model construction and survival analysis
[0067] Using the TCGA-LUAD cohort as the training set, a prognostic model related to lactation was constructed, and its predictive performance was evaluated on the independent validation set GSE87340. Modeling was performed using the R package Mime1 (v0.99.9), which integrates 10 machine learning algorithms, including Random Survival Forest (RSF), Least Absolute Shrinkage and Selection Operator (LASSO), and Support Vector Machine (SVM), generating a total of 101 combined models.
[0068] In the first modeling, the input features included: (i) a set of 332 lactation-related genes; (ii) the top 20 characteristic genes in epithelial cells with high lactation activity; and (iii) ligand-receptor gene pairs significantly enriched between lactated epithelial cells and immune cells. Based on the model results, genes repeatedly selected in ≥15 models were identified and defined as key genes, forming a candidate set of 12 genes: ANGPTL4, ITGA6, SOD1, CCL20, FKBP3, DECR1, SEMA3C, TRIM28, VEGFC, TNNC2, CRTAC1, and HGF.
[0069] Subsequently, a second modeling was conducted based on this key gene set, constructing 101 new combinatorial models. The results are shown below. Figure 1 and 2 As shown, comparing the AUC and C-index under different algorithms, the results show that the model based on Random Survival Forest (RSF) achieves an AUC of 0.92 and a C-index of 0.85 in the training set, demonstrating the best performance.
[0070] Based on this RSF model, the mean of the Cumulative Hazard Function (CHF) for each patient is calculated as a risk score. For example... Figure 3 As shown, in the TCGA-LUAD training set, to determine the optimal cutoff value for the risk score, we used the `surv_cutpoint` function from the R package `survminer`. Aiming to maximize the log-rank test statistic, we iterated through all possible split points, dividing patients multiple times and calculating the corresponding survival differences. This method can automatically identify the threshold that best distinguishes survival outcomes in continuous variables. The final optimal cutoff value was determined to be 68.67, based on which patients were divided into high-risk and low-risk groups. Kaplan-Meier survival analysis showed a significant difference in overall survival between the two groups (log-rank p < 0.001, HR = 0.03), indicating that this cutoff value has good discriminative power.
[0071] In the independent validation set GSE87340, the same model and cutoff values were used for prediction. For example... Figure 4 As shown, the overall survival of patients in the high-risk group was significantly shorter than that in the low-risk group (log-rank p = 0.036, HR = 0.16), indicating that the model has good generalization ability.
[0072] In addition, such as Figure 5 As shown, univariate Cox regression analysis further validated the independent impact of these 12 genes on the prognosis of lung adenocarcinoma. High expression of 9 genes (such as ANGPTL4 and ITGA6) was significantly associated with poor prognosis (risk genes), while low expression of 3 genes (such as TNNC2 and HGF) was associated with poor prognosis (protective genes).
[0073] 6. Patient Sample Collection
[0074] LUAD tissue samples are obtained from patients undergoing surgical resection. Tumors and adjacent non-tumor tissues are collected intraoperatively, immediately and rapidly frozen in liquid nitrogen, and stored at −80 °C for analysis.
[0075] 7. Quantitative Real-Time PCR (qRT-PCR)
[0076] RNA was extracted from A494 cells using the TRIzol method (TaKaRa RNAiso Reagent, 9180). RNA quality was assessed using a NanoDrop 2000 spectrophotometer (Thermo Fisher) and 1.5% agarose gel electrophoresis. cDNA was synthesized using the PrimeScript™ RT kit (RR047A, Takara) according to the manufacturer's instructions. qRT-PCR was performed using TBGreen Premix Ex Taq™ II (RR420A, Takara) on a Bio-Rad CFX96 system. The reaction volume was 20 μL, and the conditions were: 95 °C for 30 s; followed by 45 cycles (95 °C for 5 s, 60 °C for 30 s); finally, melting curve analysis was performed. Relative expression levels were determined by 2... ⁻ΔΔCt The method was used for calculation, and all reactions were repeated three times. Primer sequences are shown in Table 1.
[0077] like Figure 6 As shown, qRT-PCR results indicated that in LUAD tumor tissue, the mRNA expression levels of risk genes (such as ANGPTL4 and ITGA6) were significantly higher than those in paired normal tissue (p < 0.01), while the expression of protective genes (such as TNNC2 and CRTAC1) was significantly reduced (p < 0.05), consistent with the risk attributes predicted by the model, thus validating the reliability of the model.
[0078] Table 1. Primers used for qRT-PCR Primer name Sequence ANGPTL4 forward (tissue sample) 5'-GCCTCCTGGACCACAAGCAC-3' ANGPTL4 reverse (tissue sample) 5'-GAACAGCTCCTGGCAATCCCT-3' ITGA6 forward 5'-TGGTTATTGAACTGCTTTTATCG-3' ITGA6 reverse 5'-GGCTCACAAGTTACCTTTTCC-3' SOD1 forward 5'-TGGGGAAGCATTAAAGGACT-3' SOD1 reverse 5'-TACACCACAAGCCAAACGAC-3' CCL20 forward 5'-AGTTTGCTCCTGGCTGCTTT-3' CCL20 reverse 5'-CCAAGTCTGTTTTGGATTTGC-3' FKBP3 forward 5'-AAGACAGCTAACAAGGACCACT-3' FKBP3 reverse 5'-ATAACTTTGCCTACTCCGACC-3' DECR1 forward 5'-ACTGACATAGTTCTAAATGGCACA-3' DECR1 reverse 5'-CTACAAAGGAAAGCAGCAAGAT-3' SEMA3C forward 5'-AATACTTCAGCCTTTCCCACC-3' SEMA3C reverse 5'-AAACTTGGTCCTCTGATCTCCT-3' TRIM28 forward 5'-CACAGCCCTTTTGCTTTCTA-3' TRIM28 reverse 5'-CCTGAGCGGGACCGTTTCAC-3' VEGFC forward 5'-GCCCCAAACCAGTAACAATC-3' VEGFC reverse 5'-GCATCCGAGGAAAACATAAA-3' TNNC2 forward 5'-AGGCTGAGGCCAGGTCCTAC-3' TNNC2 reverse 5'-GTGCCCAACTCCTTGACGCT-3' CRTAC1 forward 5'-AAGTTGTTCAAGTTCCGCAATA-3' CRTAC1 reverse 5'-TTGTGGAAAAGGAAGTTAGGC-3' HGF forward 5'-TATCTTGTGCCAAAACGAAAC-3' HGF reverse 5'-AACAAAATCATCCAGGACAGC-3' actinl forward 5'-GCGGGAAATCGTGCGTGACATT-3' actinl reverse 5'-GATGGAGTTGAA GGTAGTTTCGTG-3' GAPDH forward 5'-ACGGATTTGGTCGTATTGGG-3' GAPDH reverse 5'-GGGATCTCGCTCCTGGAAG-3' ANGPTL4 forward (cell sample) 5'-GGCTCAGTGGACTTCAACCG-3' ANGPTL4 reverse (cell sample) 5'-CCGTGATGCTATGCACCTTCT-3' G6PD forward 5'-CTACCGCATCGACCACTACC-3' G6PD reverse 5'-TGTTGTCCCGGTTCCAGATG-3' PKM2 forward 5'-CTGGTGACGGAGGTGGAAAA-3' PKM2 reverse 5'-GCTCGACCCCAAACTTCAGA-3' HK1 forward 5'-GAGGGTCCTGGATCAGAGGT-3' HK1 reverse 5'-GACCTTCAGGAATGGCTGCT-3' LDHA forward 5'-ACACCCAAACGTCGATATTCCT-3' LDHA reverse 5'-GGGGGTCTGTTCTTCCTTTAGA-3'
[0079] 8. Cell Culture and Processing
[0080] Human LUAD cell line A549 was purchased from Xiamen Yimo Biotechnology Co., Ltd. (IM-H113). Cells were cultured in DMEM (Gibco, C11995) supplemented with 10% FBS (HyClone, SH30406.05) and 1% penicillin-streptomycin (Haicheng Yuanhong, HCCC102) and incubated at 37°C and 5% CO2. Cells were passaged with 0.25% trypsin (Invitrogen, 25200056) at approximately 80% confluence, and only cells with >90% viability were used for experiments. In functional experiments, cells were cultured at 2 × 10⁶ cells / year. 6 Cells were seeded at 50–60% density and subjected to hypoxia treatment (0, 4, 8, 12 h). Cells were then collected, lysed with TRIzol, and RNA was extracted and analyzed by qRT-PCR.
[0081] like Figure 7 As shown in (AD), with prolonged hypoxia, the expression levels of key glycolysis genes HK1, PKM2, G6PD, and LDHA were significantly upregulated (p < 0.01), indicating enhanced cellular glycolytic activity and increased lactate production. Against this backdrop, as... Figure 7 As shown in (E), the expression of the potential lactation-related gene ANGPTL4 was also significantly upregulated in a time-dependent manner with hypoxia time (p < 0.01), consistent with the results of five independent biological replicates, further supporting the potential biological link between ANGPTL4 and the lactation process.
[0082] 9. Statistical Analysis
[0083] All statistical analyses were performed in R software (v4.2.0). Wilcoxon's rank-sum test was used for comparisons of continuous variables, and chi-square test or Fisher's exact test was used for comparisons of categorical variables. Survival curves were plotted using the Kaplan-Meier method, and log-rank test was used to analyze differences between groups. Univariate prognostic analysis was performed using a Cox proportional hazards regression model. A two-sided p-value <0.05 was considered statistically significant.
Claims
1. A combined biomarker for predicting the prognosis of lung adenocarcinoma, characterized in that, Including the following 12 genes: ANGPTL4, ITGA6, SOD1, CCL20, FKBP3, DECR1, SEMA3C, TRIM28, VEGFC, TNNC2, CRTAC1, HGF.
2. A risk scoring model for predicting the prognosis of lung adenocarcinoma, characterized in that, It is obtained by constructing based on the joint marker as described in claim 1.
3. The method for constructing a prognostic risk scoring model for patients with lung adenocarcinoma as described in claim 2, characterized in that, Includes the following steps: (a) Obtain tumor tissue samples from patients with lung adenocarcinoma and detect the expression levels of 12 genes; (b) Input the patient’s gene expression level data into the prognostic model trained based on the random survival forest algorithm, and calculate the mean of the patient’s cumulative risk function as the risk score; (c) Stratify patients by risk based on preset cutoff values.
4. The method for constructing a prognostic risk scoring model for patients with lung adenocarcinoma according to claim 3, characterized in that, The risk stratification is as follows: if the risk score is lower than the preset cutoff value, it is predicted to be in the low-risk group with a better prognosis; if the risk score is higher than or equal to the preset cutoff value, it is predicted to be in the high-risk group with a poor prognosis.
5. The method for constructing a prognostic risk scoring model for patients with lung adenocarcinoma according to claim 4, characterized in that, The preset cutoff value is 68.
67.
6. A diagnostic reagent for predicting the prognosis of lung adenocarcinoma, characterized in that, It includes a primer set or probe set for detecting the expression level of the combined biomarker as described in claim 1.
7. A diagnostic system for predicting the prognosis of lung adenocarcinoma, characterized in that, include: (1) Sequencing module: used to collect tumor tissue samples from patients and detect the expression level of the combined biomarker as described in claim 1; (2) Output prediction module: It is used to input the test results into the random survival forest prognostic model, calculate the patient's risk score, and perform risk stratification accordingly, and finally output the prognostic judgment result of the overall survival.
8. The diagnostic system for predicting the prognosis of patients with lung adenocarcinoma according to claim 7, characterized in that, The risk stratification is as follows: if the risk score is lower than the preset cutoff value, it is predicted to be in the low-risk group with a better prognosis; if the risk score is higher than or equal to the preset cutoff value, it is predicted to be in the high-risk group with a poor prognosis.
9. The diagnostic system for predicting the prognosis of patients with lung adenocarcinoma according to claim 7, characterized in that, The preset cutoff value is 68.67.