A proteasome-related gene-based gene prognosis risk assessment model and application thereof

CN122762289APending Publication Date: 2026-09-15ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610979502.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-15

Smart Images

  • Figure CN122762289A_ABST
    Figure CN122762289A_ABST
Patent Text Reader

Abstract

The application discloses a gene prognosis risk assessment model based on proteasome-related genes and application thereof. The proteasome-related genes are applied to constructing a lung adenocarcinoma gene prognosis risk assessment model; a gene data set is formed by integrating lung adenocarcinoma related gene data, the gene data set is processed to obtain the correlation characteristics of the proteasome-related genes and lung adenocarcinoma gene data, a risk assessment model is established by using the correlation characteristics, the lung adenocarcinoma related gene data is used for training and verification, and a model with the optimal consistency index is screened as the lung adenocarcinoma gene prognosis risk assessment model. The application can effectively identify a high-risk LUAD population, has a robust prognosis discrimination and immune phenotype stratification capability, and reveals the key regulation between the malignant epithelium and immune cells at a single cell level; under the background of immunotherapy, the model is expected to improve the accuracy of response prediction and provide a new path for future precise intervention of lung adenocarcinoma.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedical technology, specifically to a gene prognostic risk assessment model based on proteasome-related genes and its application. Background Technology

[0002] Lung cancer is one of the most common and deadliest malignant tumors worldwide, posing a significant threat to public health and safety. Due to the lack of obvious symptoms in the early stages, approximately 70% of NSCLC patients are diagnosed at an advanced stage. The high heterogeneity of tumors, complex mutational landscapes, and resistance to chemotherapy, targeted therapy, and immunotherapy contribute to a poor overall prognosis, with a five-year survival rate consistently around 20%. In recent years, the widespread use of immunotherapy strategies, such as immune checkpoint inhibitors (ICIs), has improved the survival rate of NSCLC patients to some extent. Lung adenocarcinoma (LUAD), the most common subtype of non-small cell lung cancer (NSCLC), remains associated with low survival rates. However, the high heterogeneity of the tumor microenvironment and the presence of immune escape mechanisms result in unsatisfactory clinical efficacy of ICIs. More importantly, the current lack of reliable biomarker systems for accurate patient stratification and efficacy prediction further limits the clinical application of ICIs.

[0003] The proteasome-ubiquitin system (UPS) is a core mechanism regulating protein homeostasis and degradation in eukaryotic cells, participating in multiple processes including cell cycle, apoptosis, DNA repair, and stress response. The 26S proteasome, as the core functional structure of this system, is mainly composed of the 20S core granule and the 19S regulatory granule. The PSMD family of non-ATPase subunits (PSMD1-14) of the 19S granule plays a crucial role in substrate recognition, depolymerization, and deubiquitination. In recent years, many studies have shown that UPS-related genes are abnormally active in various solid tumors, promoting tumor cell growth and immunosuppression. While the proteasome plays a central role in protein degradation and immune regulation, its prognostic and immunological relevance in LUAD (lumbar inflammatory disease) remains unclear. Proteasome-related genes show an upregulated trend in various oncology genes and are closely associated with poor patient prognosis. Previous studies have indicated that proteasome-related genes not only affect cell proliferation by regulating protein degradation rates but may also participate in regulating key cell cycle molecules, influencing tumor cell tolerance to stress and chemotherapy drugs. Although the role of proteasome-related genes in various tumors has attracted attention, systematic studies in lung cancer remain scarce, and their potential prognostic value and immunomodulatory mechanisms have not yet been systematically elucidated. Summary of the Invention

[0004] To ensure a stable and reliable biomarker system for precise stratification and efficacy prediction, and to prevent limitations on the further clinical application of ICIs, this invention aims to construct and validate a gene prognostic risk assessment model based on proteasome-related genes (PRGs). This model can effectively identify high-risk individuals among lung adenocarcinoma (LUAD) patients and possesses stable prognostic assessment and immunophenotypic analysis capabilities. The score is closely related to immunosuppression status, driver gene mutations, and cell communication networks, and reveals its crucial regulatory role between malignant epithelial cells and immune cells at the single-cell level. In the context of immunotherapy, this model is expected to supplement traditional indicators such as programmed death-ligand 1 (PD-L1), improving the accuracy of treatment response prediction. Furthermore, its good clinical applicability provides a new direction for future precision intervention and treatment of lung adenocarcinoma.

[0005] The technical solution adopted in this invention is: I. Application of a proteasome-associated gene (PRG) Applications in gene prognostic risk assessment and the construction of a computer-based gene prognostic risk assessment model for lung adenocarcinoma. The gene prognostic risk assessment model includes a lung adenocarcinoma gene risk model and a lung adenocarcinoma prognostic risk assessment model.

[0006] II. A method for constructing a gene prognostic risk assessment model based on proteasome-related genes. The construction method steps are as follows: S1. Extract and integrate data from the Lung Adenocarcinoma Genome Atlas Cohort (TCGA) and five independent Gene Expression Comprehensive Database Cohorts (GEO) from the Gene Expression Comprehensive Database to form a gene dataset, which is then stored in the computer's internal memory.

[0007] S2. Using computing tools such as processors inside the computer, Cox regression analysis is performed on the gene dataset to identify proteasome-related genes. Functional enrichment analysis is performed on proteasome-related genes, and then the association features between the risk assessment of proteasome-related genes and the lung adenocarcinoma genome map cohort data are obtained. The association features are used as key molecular features. This study investigates the association between risk assessment of proteasome-related genes and clinical characteristics, tumor mutational burden (TMB), and immune characteristics. Functional enrichment analysis is used to reveal the KEGG pathway and GO function involved in these genes, thereby constructing a prognostic risk model for lung adenocarcinoma based on proteasome-related genes.

[0008] S3. Run the program in the computer, use key molecular features, and use multiple algorithms to establish corresponding risk assessment models through machine learning. Use the data from the Cancer Genome Atlas Cohort (TCGA) as the training set and the data from the Gene Expression Comprehensive Database Cohort (GEO) as the validation set. After training the established risk assessment models using the training set, validate them using the validation set to evaluate the predictive performance of these models. Select the model with the best consistency index as the optimal model, which is the final lung adenocarcinoma gene prognostic risk assessment model, and display it on the user terminal screen.

[0009] Finally, single-cell RNA (scRNA-seq) was used to map the PRG expression profile in the tumor microenvironment.

[0010] Step S1 specifically involves: acquiring RNA expression profile data and clinical follow-up data from the existing platform's genomic atlas cohort for lung adenocarcinoma and five independent gene expression comprehensive database cohorts; normalizing the RNA expression profile data and clinical follow-up data; mapping probe IDs to official gene symbols to unify gene identification; and then calculating the average expression value of multiple probes for the same gene for standardization to obtain the final gene dataset.

[0011] The genomic atlas cohort data for lung adenocarcinoma used data from the UCSC Xena platform; the five independent gene expression comprehensive database cohorts used five different gene datasets from the GEO-GSEx cohort, and the gene dataset numbers GSE and sample numbers n in the GEO-GSEx cohorts used were as follows: GSE13213, n=117; GSE29016, n=68; GSE30219, n=293; GSE31210, n=226; GSE42127, n=133.

[0012] In step S2, the proteasome-related genes are those that are significantly associated with overall survival. The associated features are those closely related to the risk assessment of proteasome-related genes that change significantly due to changes in clinical features, tumor mutational burden (TMB), and immune features in the lung adenocarcinoma cohort. The proteasome-related genes are mainly one or more combinations of USP5, UBC, PSMD2, PSMD12, PSMD11, and CDC20.

[0013] Key molecular signatures serve as reliable prognostic indicators in LUAD, reflecting the immunosuppressive state of the tumor microenvironment. These findings provide a foundation for improving patient stratification and targeting the proteasome pathway in therapy.

[0014] Step S3 specifically involves: S3.1 Utilizing key molecular features, the data from the Cancer Genome Atlas was used as the training set. Multiple ensemble machine learning methods were employed in the Cancer Genome Atlas cohort, including glmnet for Lasso feature selection and plsRcox model for variable screening, to establish multiple risk assessment models. The final model was determined based on the optimal parameter combination (lambda.min).

[0015] S3.2. Data from the Gene Expression Omnibus (GEO) database is used as a validation set for external validation, and the risk assessment model is evaluated for its risk prediction performance to obtain risk samples. S3.3 Calculate the risk assessment of the risk sample, and divide the sample into high-risk group and low-risk group according to the median of the risk assessment; use Kaplan-Meier survival analysis to assess the survival difference between the high-risk group and the low-risk group, and calculate the area under the receiver operating characteristic curve (AUC) to judge the predictive accuracy of the risk assessment model, and obtain the risk assessment model with the best consistency index. S3.4 Select the risk assessment model with the best consistency index as the optimal model, which is the final prognostic risk assessment model.

[0016] In step S3.3, the more significant the survival differences, the more accurate the prediction of the risk assessment model, that is, the optimal consistency index of the risk assessment model.

[0017] III. Construction Method of Gene Prognostic Risk Assessment Model Based on Proteasome-Related Genes and Application of the Constructed Gene Risk Assessment Model Applications in assisting risk stratification, prognosis assessment, and immune prediction in patients with lung adenocarcinoma.

[0018] IV. A machine learning device Includes the following components: Processor, used for regression analysis and machine learning model training; Storage device used to store gene datasets; A display screen is used to present the modeling results of the lung adenocarcinoma gene prognostic risk assessment model; A computer program, stored in memory and executable by a processor, which, when executed by the processor, enables the construction of a gene prognostic risk assessment model based on proteasome-related genes.

[0019] PRG-based models have shown consistent prognostic accuracy. High PRG scores correlate with CD8. +Decreased T cell and B cell infiltration, enrichment of regulatory T cells and myeloid-derived suppressor cells (MDSCs), and increased expression of immune checkpoints were associated with tumors. High-risk tumors exhibited greater genomic instability and higher tumor tumor mass (TMB). scRNA-seq analysis showed that malignant epithelial cells and regulatory T cells with high PRG activity promoted immunosuppressive signaling through secreted phosphoprotein 1 (SPP1), transforming growth factor-β (TGF-β), and vascular endothelial growth factor (VEGF) pathways. In vitro and in vivo experiments further validated that downregulation of PSMD11 significantly inhibited tumor growth, consistent with single-cell analysis.

[0020] This invention discloses a protease-associated gene (PRG) and its applications and methods. The constructed risk assessment system based on proteasome-associated genes (PRGs) can effectively identify high-risk individuals for lung adenocarcinoma (LUAD) and possesses robust prognostic and immunophenotypic stratification capabilities. This score is closely related to immunosuppression status, driver gene alterations, and cell communication networks, and reveals its key regulatory role between malignant epithelial cells and immune cells at the single-cell level. In the context of immunotherapy, this model is expected to supplement traditional indicators such as PD-L1 to improve the accuracy of response prediction; its good clinical adaptability provides a new pathway for future precision intervention in lung adenocarcinoma.

[0021] The technical problems to be solved by this invention are the inaccurate risk stratification of lung adenocarcinoma patients, the limitations of predictive biomarkers for immunotherapy, and the insufficient capture of dynamic changes in the tumor microenvironment.

[0022] One of the technical solutions of this invention is to identify proteasome-related genes that are significantly associated with overall survival through univariate Cox proportional hazards regression analysis based on multi-cohort RNA expression profiles and clinical data, and to use functional enrichment analysis to reveal the KEGG pathway and GO function involved in them, thereby constructing a gene risk model.

[0023] The second technical solution of this invention is to adopt a comprehensive machine learning strategy, combining Lasso regression and plsRcox models for feature selection and variable optimization, to construct a prognostic risk assessment model, and to perform external validation in an independent cohort. The survival differences and prediction accuracy between the high-risk group and the low-risk group are evaluated through Kaplan-Meier survival analysis and time ROC curves.

[0024] The third technical solution of the present invention integrates and visualizes cell type distribution through single-cell transcriptome analysis to assess the differences in PRG risk assessment in different cell populations, and combines in vitro functional experiments such as cell proliferation, colony formation, invasion and apoptosis detection, as well as in vivo animal models to verify gene function, to comprehensively analyze the correlation between risk assessment and clinical molecular characteristics.

[0025] The fourth technical solution of the present invention is to use statistical analysis tools for data correction and visualization to ensure the repeatability and generalization ability of the model, thereby supporting the application of the risk assessment model in the precise stratification of lung adenocarcinoma and the prediction of immunotherapy.

[0026] The beneficial effects of this invention are: This invention significantly improves the accuracy of risk stratification in lung adenocarcinoma patients by constructing a risk assessment model based on proteasome-related genes. It demonstrates high predictive performance in multiple independent cohort validations; for example, the area under the time-reactive index (ROC) curves at 1, 3, and 5 years is superior to existing methods. This model integrates multi-omics data and functional experimental validation, comprehensively capturing dynamic changes in the tumor microenvironment and effectively identifying high-risk individuals and immunotherapy responders, thereby improving patient prognosis and reducing treatment resistance. Furthermore, this assessment model is highly correlated with clinicopathological features and molecular markers such as TMB, providing a standardized and reproducible biomarker system to support personalized medical decision-making. Finally, in vitro and in vivo experiments confirmed the key roles of core genes in cell proliferation, invasion, and apoptosis, enhancing the model's clinical applicability and translational potential.

[0027] Based on the carcinogenic effects of proteasome-related genes in various cancers, this invention systematically explored the biological functions and clinical value of proteasome-related genes in lung cancer (LUAD) through multi-omics data analysis and biological experiments. This invention innovatively combines multi-omics bioinformatics analysis with in vitro and in vivo functional experiments. The crucial role of the key proteasome gene PSMD11 in the development and progression of lung cancer was verified through in vitro functional experiments and a mouse subcutaneous tumor model. Focusing on "proteasome regulation-risk assessment and prediction," this invention, for the first time, constructed and validated a LUAD risk assessment model targeting proteasome-related genes, elucidating their key roles in tumor biology and immune regulation. This assessment model demonstrated excellent prognostic prediction and immune signature stratification capabilities in multiple cohorts, suggesting its broad clinical application prospects. Attached Figure Description

[0028] Figure 1 This is a diagram showing the characterization and prognostic assessment of proteasome-related genes (PRGs). In this diagram, A shows the differential expression analysis of PRGs in tumor tissue and paired normal lung tissue specimens; B shows the survival analysis and identification of proteasome subunits with prognostic significance. Figure 2 A diagram from the Kyoto Encyclopedia of Genes and Genomes (KEGG) illustrating the biological functions of PRGs. Figure 3 This is a graph showing the enrichment of pathways in the Gene Ontology (GO) database. Figure 4 A graph showing copy number variation (CNV) analysis; Figure 5 Prognostic C-index heatmaps generated for different algorithms; Figure 6 A circular Manhattan plot showing the chromosomal location of key proteasome-related genes and their survival-related P values; Figure 7 The image shows the prognostic results of a risk assessment model for proteasome-associated genes (PRGs) constructed and validated based on multiple machine learning algorithms. AF represents the prognostic results of this risk assessment model in six independent lung adenocarcinoma (LUAD) cohorts. Figure 8 The graphs represent the baseline validation plots of the risk model. A represents the baseline validation plot of the algorithm's C-index; B and C represent the validation plots of the algorithm's C-index on the independent datasets GSE13213 and GSE29016; D represents the validation plots ranking in the top three on the GSE30219 dataset; and E and F represent the validation plots of the C-index on the independent datasets GSE31210 and GSE42127. Figure 9 This is a graph assessing the prognostic performance of the risk model in lung adenocarcinoma (LUAD), where AF is the time-dependent receiver operating characteristic (ROC) curve and its area under the curve (AUC). Figure 10 The results are shown in the principal component analysis (PCA) diagram, where AF represent the results of different principal component analysis (PCA) methods. Figure 11 A graph showing the differential expression of core genes between the high-risk group and the low-risk group; Figure 12 This is a correlation plot between risk assessment and the expression levels of key genes, where AF is a scatter plot of the correlation between risk assessment and the expression levels of each key gene; Figure 13 This is a multi-omics atlas of alterations in the TCGA lung adenocarcinoma (LUAD) cohort based on proteasome-related genes (PRGs) risk assessment. A is an oncoprint comparing the mutation landscape between high-risk and low-risk groups; B is a copy number variation (CNV) atlas between high-risk and low-risk samples; C is a correlation analysis of risk assessment and tumor mutational burden (TMB); and D is a survival comparison between subgroups based on a combined definition of TMB and risk assessment. Figure 14 Heatmap of immune infiltration based on the abundance of different immune cell populations; Figure 15 This is a diagram illustrating the expression pattern of immune checkpoints stratified by risk components. Figure 16 A graph showing the correlation between pathway activity and tumor immune circulation status; Figure 17 The image shows a single-cell atlas of proteasome expression enrichment analysis in immunosuppressive and malignant epithelial cell populations. In this atlas, A is a bubble diagram of characteristic genes from different cell subpopulations; BC is a UMAP visualization of cell clusters; D is a diagram of the expression patterns of model genes in each cell cluster; and E is a distribution diagram of risk score across all identified cell types. Figure 18 A differential signaling pathway diagram between the high-risk assessment group and the low-risk assessment group; Figure 19 This is a comparison chart between high-risk and low-risk groups, where BC is a comparison chart of communication strength and communication quantity between the high-risk and low-risk groups, and D is a comparison chart of the bidirectional signal transduction ability of different cells in the high-risk and low-risk groups. Figure 20 A heatmap illustrating the differences in intercellular communication among lung adenocarcinoma (LUAD) patients stratified by risk status; Figure 21 The graphs show the persistently high expression of PSMD11 in lung cancer tissues and its correlation with poor patient prognosis. Specifically: A) Differential expression analysis of PSMD11 between lung cancer and normal tissues based on TCGA data; B) Comparison of PSMD11 expression in paired adjacent normal and tumor tissues; C) Kaplan-Meier survival curve analysis of the relationship between high and low PSMD11 expression levels and overall survival; D) Nomogram model assessing the role of PSMD11 in multivariate prognostic prediction; E) Immunohistochemical analysis of differential PSMD11 expression between adjacent normal and lung adenocarcinoma tissues; F) Western blot analysis of PSMD11 protein levels in paired adjacent normal and tumor tissues; G) Western blot analysis of differential PSMD11 expression between tumor and adjacent normal tissues. Figure 22 The diagrams show the effects of PSMD11 gene knockdown on the morphology of lung cancer cells. AB represents the long-term proliferation capacity of A549 and H446 cells assessed using a colony formation assay; CD represents the migration and invasion capacity of A549 and H446 cells assessed using a Transwell assay; E represents the DNA synthesis and proliferation activity of cells assessed using EdU staining; F represents the short-term proliferation curves of A549 and H446 cells analyzed using a CCK-8 assay; and G represents the apoptosis status assessed using TUNEL staining. Figure 23The diagram shows the in vivo antitumor effect of PSMD11 gene knockdown. A shows the monitoring results of tumor volume changes over time in the subcutaneous tumorigenesis experiment in nude mice; B shows the tumor difference between the control group and the shPSMD11 group (PSMD11 gene knockdown group); C shows the comparison of tumor weight between the control group and the shPSMD11 group at the experimental endpoint; D shows the immunofluorescence analysis and TUNEL staining of Ki67 expression in tumor tissue. Figure 24 A diagram illustrating antigen presentation and immune regulatory pathways such as TGF-β; Figure 25 To illustrate the prognostic relationship between high expression of core PRGs and the overall survival of LUAD patients, the following Kaplan-Meier survival curves are presented: A) High and low expression of USP5 and overall survival of LUAD patients; B) High and low expression of UBC and overall survival of LUAD patients; C) High and low expression of PSMD2 and overall survival of LUAD patients; D) High and low expression of PSMD12 and overall survival of LUAD patients; E) High and low expression of PSMD11 and overall survival of LUAD patients; and F) High and low expression of CDC20 and overall survival of LUAD patients. Figure 26 The map shows that core PRGs are significantly highly expressed in LUAD. Detailed Implementation

[0029] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0030] Specific embodiments of the present invention are as follows: Example: 1. Integration of multi-source data and identification of prognostic PRGs This invention integrates RNA expression profiles and clinical follow-up data from the TCGA-LUAD cohort and five independent GEO-GSEx cohorts (GSE13213, n=117; GSE29016, n=68; GSE30219, n=293; GSE31210, n=226; GSE42127, n=133). TCGA data were obtained from the UCSC Xena platform (https: / / xenabrowser.net). FPKM was converted to TPM and normalized using log2(TPM+1). GEO data were downloaded from the NCBI GEO database (raw CEL file or Series Matrix). Probe IDs were mapped to HGNC official gene symbols according to the platform annotation files (GPL69480, GPL6947, GPL570, GPL6884); the average expression value of multiple probes for the same gene was taken. Before cross-cohort analysis, batch effect correction was performed using the ComBat method in the R package sva (v3.44.0) with "dataset" as the batch variable, and the correction effect was visualized using principal component analysis (PCA). Subsequently, univariate Cox proportional hazards regression was performed in each cohort using the R package survival (v3.5-7) to screen proteasome-related genes (PRGs) that were significantly associated with overall survival (OS). Hazard ratios (HR), 95% confidence intervals (CI), and p-values ​​were calculated; p < 0.05 was used as the significance threshold.

[0031] 2. Functional enrichment analysis The screened prognostic-related PRGs were input into the R package clusterProfiler (v4.8.1) for KEGG pathway enrichment and GO function enrichment analysis. The annotation database was obtained from org.Hs.eg.db (v3.17.0). Multiple tests were performed with Benjamini–Hochberg correction, and the significance threshold was set to p.adjust < 0.05. The results were visualized using enrichplot (v1.20.0) and ggplot2 (v3.4.4).

[0032] 3. Construction of the PRG prognostic risk model An ensemble machine learning strategy was used to construct a prognostic risk assessment model for PRG in the TCGA cohort. Lasso feature selection was performed using glmnet (v4.1-8), and variable selection was completed using the plsRcox model; the final model was determined based on the optimal parameter combination (lambda.min). External validation was then performed in five independent GEO cohorts. Kaplan-Meier survival analysis (R packages survival v3.5-7 and survminer v0.4.9) was used to assess the survival differences between high- and low-risk groups; the area under the curve (AUC) at 1, 3, and 5 years was calculated using timeROC (v0.4).

[0033] 4. Correlation between PRG risk assessment and clinical / molecular characteristics Spearman correlation analysis was used to assess the relationship between PRG risk assessment and clinicopathological features, focusing on core genes such as USP5, UBC, and PSMD2. In the TCGA cohort, tumor mutational burden (TMB) was calculated based on MAF files; defined as the total number of nonsynonymous mutations per sample / estimated exon region size, in Mb. Patients were then divided into high and low TMB groups based on the median. Kaplan-Meier survival analysis was further performed based on the risk assessment combination.

[0034] 5. Single-cell transcriptome analysis LUAD single-cell transcriptome data were obtained from (GSE13213, n=117; GSE29016, n=68; GSE30219, n=293; GSE31210, n=226; GSE42127, n=133). Data processing and integration were performed using the R package Seurat (v4.3.0). Quality control criteria: 200–5000 genes detected per cell, with mitochondrial genes accounting for <10%. Normalization and standardization were performed using the LogNormalize method, and 2000 hypervariable genes were selected for PCA dimensionality reduction. Cross-sample integration was performed using the R package Harmony (v1.2.0) to remove batch effects, and visualization was performed using UMAP (dims=1:30). Cell type annotation was completed based on known marker genes and literature references. PRG risk assessment was calculated at the single-cell level and mapped to UMAP plots to compare distribution differences among different cell populations.

[0035] Note: The GEO number mentioned above was listed as originating from a single cell in the original text; the original statement is faithfully retained here.

[0036] 6. Cell proliferation experiment Cell proliferation was assessed using CCK-8 and EdU assays. A549 and H446 cells were cultured at 3 × 10⁻⁶ cells / year. 3Cells were seeded at a density of 100 cells / well in 96-well plates, and 100 μl of complete culture medium was added, allowing for overnight attachment. Cell proliferation was measured from day 1 to day 5 in the CCK-8 assay. For assays, 90 μl of fresh DMEM medium (Corning) and 10 μl of CCK-8 reagent (Beyotime, Shanghai, China) were added to each well, followed by incubation at 37°C for 1 hour. Afterward, absorbance (OD) at 450 nm was recorded using a multi-functional microplate reader (EnSpire®, PerkinElmer, USA) to assess cell proliferation.

[0037] In EdU experiments, A549 and H446 cells were incubated with 10 μM EdU for 2 hours using the BeyoClick™ EdU-488 Cell Proliferation Kit (Beyotime). Cells were then fixed with 4% paraformaldehyde for 30 minutes at room temperature.

[0038] After fixation, cells were treated with Beyotime immunostaining permeabilization solution for 5 minutes, and then incubated with Click reaction solution containing fluorescent azide for 30 minutes at room temperature in the dark. After washing, slides were mounted with mounting medium containing DAPI. EdU-positive cells were imaged using a confocal fluorescence microscope (LSM980, Zeiss, Germany).

[0039] 7. Cloning experiment Colony formation was assessed in exponentially growing A549 and H446 cells. Cells were grown at a rate of 3 × 10⁶. 2 Cells were seeded at a density of 1:1 in 6-well plates and cultured in complete medium for 1–2 weeks, changing the medium every 3–4 days until more than 50 cells were observed in most visible clones. Clones were fixed with 4% paraformaldehyde (Aladdin, Shanghai, China) for 20 minutes and stained with 0.1% crystal violet (Aladdin) for 30 minutes, washed, air-dried, and photographed. The number of clones was quantified using ImageJ software (NIH, Bethesda, MD, USA).

[0040] 8. Transwell invasion experiment Cell invasion was examined using Matrigel-coated Transwell chambers. Transwell inserts (Jet Biofil, Guangzhou, China) were pre-coated with Matrigel (Corning, New York, USA) according to the manufacturer's instructions. Exponentially growing A549 or H446 cells were resuspended in serum-free DMEM at 1 × 10⁻⁶ cells per chamber. 3Cells were seeded at a density per well into the upper chamber. DMEM containing 20% ​​FBS was added to the lower chamber as a chemical inducer. After incubation at 37°C for 48 hours, non-invasive cells were gently removed from the membrane with a cotton swab. Cells that had invaded the lower surface were fixed with 4% paraformaldehyde for 20 minutes, stained with 0.1% crystal violet for 30 minutes, washed, and air-dried. Invading cells were counted in multiple random fields using ImageJ.

[0041] 9. TUNEL apoptosis assay Apoptosis in cultured cells and paraffin-embedded tissue sections was assessed using the One-Step TUNEL Apoptosis Detection Kit (Beyotime, Shanghai, China), with slight adjustments to the manufacturer's instructions. Briefly, cultured cells or paraffin sections were fixed with 4% paraformaldehyde for 30 minutes at room temperature, followed by treatment with the immunostaining permeabilization solution (Beyotime) provided in the kit for 5 minutes. After washing with PBS, samples were incubated at 37°C for 1 hour with a TUNEL reaction mixture containing terminal deoxynucleotidyl transferase (TdT) and Cy3-labeled dUTP. Cell nuclei were counterstained with mounting medium containing DAPI, and TUNEL-positive cells were visualized using a confocal fluorescence microscope (LSM980, Zeiss, Germany).

[0042] 10. Western Imprint (WB) Cell and tissue samples were lysed in ice-cold RIPA lysis buffer with protease inhibitors added. The lysis buffer was centrifuged at 12000 × g for 15 min at 4°C to collect the supernatant. Total protein concentration was determined using a BCA protein quantification kit (Thermo Scientific, Waltham, MA, USA). Equal volumes of protein (approximately 30 µg each) were separated on 10–12% SDS-PAGE gels and transferred to PVDF membranes (Millipore, Bilerica, MA, USA).

[0043] The membrane was blocked with 5% skim milk powder (BD Biosciences) at room temperature for 1 hour, then incubated overnight at 4°C with primary antibodies against PSMD11 (1:1000, ABclonal), CDK4 (1:1000, ABclonal), and microtubules (1:5000, ABclonal). After washing three times with TBST, the membrane was incubated with HRP-labeled secondary antibodies (1:4000, ABclonal) at room temperature for 1 hour. Protein bands were detected using an enhanced chemiluminescence (ECL) kit (ABclonal) and visualized using a chemiluminescence imaging system (Multi-5200, Tanon, Shanghai, China).

[0044] 11. Quantitative Real-Time PCR (qRT-PCR) Total RNA was extracted from cells using the TransZol Up Plus RNA Kit (TransGen Biotech, Beijing, China). RNA concentration and purity were determined spectrophotometer-wise, and then 1 μg of total RNA was reverse transcribed into cDNA using the PrimeScript™ RT Kit (Takara, Dalian, China). qRT-PCR was performed using the Green qPCR SuperMix (TransGen Biotech) on a CFX96 Touch Real-Time PCR System (Bio-Rad, Hercules, California, USA). Relative mRNA expression was measured using 2... - The ΔΔCt method was used for calculation, with GAPDH as an internal reference. The primers for PSMD11 were: the forward F primer was 5'-ACGAAGCTGCTCTGGAAACA-3' (as shown in SEQ ID No. 1); the reverse R primer was 5'-GTCTCGACACGCAGAAGACA-3' (as shown in SEQ ID No. 2).

[0045] 12. Immunofluorescence staining Paraffin-embedded tissue sections were dewaxed, rehydrated, and subjected to heat-induced antigen retrieval in sodium citrate buffer (Beyotime). Sections were permeabilized with 0.1% Triton X-100 (Beyotime) and blocked with blocking solution (Beyotime) for 30 minutes at room temperature. After blocking, sections were incubated overnight at 4°C with a primary antibody against Ki67 (1:200, ABclonal), followed by incubation in the dark at room temperature for 1 hour with a secondary antibody labeled with Alexa Fluor 488 (1:500, Abcam). Sections were mounted with DAPI-containing mounting medium (Abcam). Fluorescence images were acquired using a confocal fluorescence microscope (LSM980, Zeiss).

[0046] 13. Subcutaneous xenograft model A xenograft model was established by subcutaneously injecting A549 cells into the backs of BALB / c nude mice (GemPharmatech, China). Twenty male mice aged 4-6 weeks and of similar weight were randomly divided into two groups. Normal A549 cells (in logarithmic growth phase) and PSMD11 knockout A549 cells were collected, washed, and resuspended in sterile phosphate-buffered saline (PBS). The cell suspension was injected into the mice every two days. Tumor length (L) and width (W) were measured every other day using calipers, and the result was calculated using the formula V = 0.5 × L × W. 2 Tumor volume was calculated. Mice were sacrificed on day 16 after cell transplantation (or when the tumor volume approached 1000 mm). 3The xenograft tumors were euthanized in advance, dissected, and weighed. Immunofluorescence analysis and TUNEL staining were performed on xenograft tumor sections to assess proliferation and apoptosis. All procedures complied with institutional guidelines and were approved by the institutional ethics committee (ethics approval number: SRRZJU202402010).

[0047] 14. Statistical Analysis Statistical analysis was performed using R software (v4.3.2). Unless otherwise specified, data are presented as mean ± standard deviation (SD). Comparisons were two-tailed, and p < 0.05 was considered statistically significant. False discovery rates were controlled for using the Benjamini-Hochberg method where appropriate. Graphs were generated using the R packages ggplot2 (v3.4.4), survminer (v0.4.9), and enrichplot (v1.20.0).

[0048] The following results were obtained from the experimental process of the above embodiments: 1. Proteasome-associated genes (PRGs) are widely involved in pan-oncology. To investigate the role of PRGs in tumorigenesis and development, this invention conducted functional enrichment analysis based on the TCGA pan-cancer dataset. The results showed that PRGs were enriched in key oncogenic and tumor microenvironment-related pathways across multiple cancer types, such as the cell cycle, oxidative phosphorylation, and MYC signaling pathways; they were also involved in immune regulatory pathways such as antigen presentation and TGF-β. Figure 24 As shown, this suggests that PRGs have broad biological regulatory functions in a pan-cancer context and may participate in immune escape by regulating the tumor immune microenvironment, laying a theoretical foundation for further research on the oncogenic function of PRGs in LUAD.

[0049] 2. Identification of key PRGs driving tumor progression To identify key genes regulating tumorigenesis and progression from PRGs, this invention first performs batch effect correction on GEO and TCGA data. Differential expression analysis initially yielded a group of PRGs that were significantly differentially expressed in tumor tissues, such as... Figure 1 As shown in Figure A, it suggests that it may be involved in the tumor pathological process. Further analysis using Cox proportional hazards regression identified 13 prognostic genes significantly associated with the proteasome pathway (P < 0.05), whose expression levels could independently predict patient survival. Figure 1 As shown in B in the figure. KEGG and GO annotations show that these genes are significantly enriched in pathways such as DNA replication, mismatch repair, cell cycle, and viral oncology. Figure 2 and Figure 3 As shown, this suggests its role in maintaining genome stability and promoting tumor cell proliferation. Figure 4The genomic loci and copy number variation status of PRGs were labeled. Copy number variation (CNV) analysis revealed widespread copy number alterations in PRGs, with the expression levels of PSMD11 and PSMD12 highly consistent with their amplification status, highlighting their potential as driver genes for LUAD. These findings provide a basis for further elucidating the mechanisms of action of key PRGs in lung cancer and their feasibility as potential therapeutic targets.

[0050] 3. Construction of a LUAD prognostic risk model based on PRGs This invention employs an ensemble machine learning approach to construct a signature of 13 survival-related proteasome pathway genes, and based on this signature, establishes various risk assessment models (including Lasso+plsRcox, CoxBoost+plsRcox, and CoxBoost+SuperPC). The predictive performance of the models is evaluated using TCGA as the training set and five GEO cohorts as the validation set.

[0051] The proteasome-associated gene (PRG) risk assessment model, constructed and validated based on multiple machine learning algorithms, prioritizes algorithms with the best average C-index on the validation set. Results show that the model constructed using Lasso+plsRcox is optimal. Figure 5 As shown. Figure 6 As shown, red dots represent genes significantly associated with prognosis; black dots represent core genes with the smallest P-values. Chromosomal localization analysis further demonstrates the distribution and expression patterns of PRGs across the genome. Figure 7 As shown in AF, this risk assessment model demonstrated stable prognostic stratification in six independent lung adenocarcinoma (LUAD) cohorts, with the low-risk group consistently showing better prognostic outcomes. In the validation of these six independent LUAD cohorts, the high-risk group defined by the model showed significantly shorter overall survival than the low-risk group. Figure 25 As shown, its reliability in LUAD prognostic assessment is confirmed. To systematically evaluate the predictive performance of the algorithm of this invention, this invention retrieved previous prognostic models based on TCGA-LUAD from PubMed for comparison; as shown... Figure 8 As shown, the performance of this model outperforms several previously published prognostic features for lung adenocarcinoma (LUAD), highlighting its superior predictive ability. The results show that the C-index of this algorithm ranks 4th among all reported models. Figure 8 As shown in A, the algorithm performs excellently. Further validation shows that the proposed algorithm significantly outperforms other comparative models in C-index on four independent datasets: GSE13213, GSE29016, GSE31210, and GSE42127. Figure 8As shown in B, C, D, E, and F in the dataset, it ranks among the top three in the GSE30219 dataset. These results fully demonstrate the effectiveness and clinical translational value of this risk assessment model in personalized prognostic prediction of LUAD.

[0052] 4. Performance validation of the PRG-based prognostic model like Figure 9 As shown, the predictive accuracy of the model on six independent datasets was quantified. Time-dependent ROC analysis showed that the PRGs risk model achieved high AUC in the 1-, 3-, and 5-year survival predictions for each cohort, indicating excellent survival prediction capabilities. Figure 9 As shown in Figure A, in the TCGA queue, the AUCs for 1, 3, and 5 years are 0.64, 0.63, and 0.62, respectively; Figure 9 As shown in B, the values ​​in GSE13213 are 0.83, 0.72, and 0.76 respectively; as... Figure 9 As shown in C, GSE29016 is 0.84, 0.86, and 0.81; as Figure 9 As shown in D, GSE30219 is 0.57, 0.67, and 0.64; as Figure 9 As shown in Figure E, the AUCs of GSE31210 are 0.72 and 0.74 in 3 and 5 years, respectively; Figure 9 As shown in Figure F, GSE42127's returns for years 1, 3, and 5 were 0.76, 0.70, and 0.70, respectively. Figure 10 As shown in the AF diagram, in each dataset, based on the expression levels of model-related genes, a clear cluster separation can be formed between the high-risk and low-risk groups (green: low-risk group; red: high-risk group). PCA further confirms the clear spatial separation between high- and low-risk groups based on gene expression profiles in the TCGA cohort. This phenomenon is also highly reproducible in the five GEO datasets, further supporting the model's effective differentiation between high- and low-risk populations. This provides a reliable tool for individualized survival prediction of lung cancer patients and highlights its clinical advantages in LUAD prognostic assessment.

[0053] 5. Association between PRGs risk assessment and clinical characteristics of LUAD Given the significant impact of clinical TNM stage and age on the prognosis of LUAD, this invention further evaluates the relationship between PRG risk assessment and these key clinical characteristics. Results showed that after risk model stratification, core PRGs (USP5, UBC, PSMD2, PSMD12, PSMD11, CDC20) were more prominently expressed in the high-risk group, and high-risk assessment was significantly correlated with T stage, tumor stage, and patient prognosis. Figure 11 As shown, Figure 11 and Figure 12The correlation analysis shows the relationship between risk assessment and the expression levels of key prognostic genes. Further correlation analysis revealed the following correlation coefficients between the core PRGs and risk assessment: USP5 r = 0.74, UBC r = 0.62, PSMD2 r = 0.86, PSMD12 r = 0.76, PSMD11 r = 0.86, and CDC20 r = 0.72. Figure 12 As shown in AF. Figure 25 and Figure 26 As shown, at the prognostic level, the high expression of these genes can effectively distinguish the survival outcome of LUAD patients. These core PRGs are significantly highly expressed in LUAD and also show high expression characteristics in other tumors, further emphasizing the effectiveness of this risk assessment model in predicting the prognosis of LUAD, while also suggesting that core PRGs have the potential for pan-cancer prognostic prediction.

[0054] 6. Genomic instability and TMB characteristics in high-risk LUAD patients like Figure 13 As shown in Figure A, somatic mutation analysis revealed significantly elevated TMB in high-risk PRGs patients; the mutation frequency of common driver genes (TP53, TTN, CSMD3, ZFHX4) was also significantly increased. These genes are involved in key processes such as the cell cycle and matrix remodeling, further highlighting their importance in tumorigenesis. Figure 13 As shown in Figure B, CND analysis further revealed more frequent chromosomal alterations in the high-risk group, significantly exacerbating genomic instability and suggesting a more aggressive tumor biological behavior. This invention stratifies patients into four groups (high / low TMB × high / low risk) for prognostic comparison. Notably, although risk assessment is positively correlated with TMB (R = 0.3, p = 7e),... -11 However, patients with low TMB / high risk had the worst outcomes, such as... Figure 13 As shown in C and D, patients with high TMB / low risk had the best outcomes, which is not entirely consistent with traditional TMB assessment models. This finding suggests that PRG scores not only reflect TMB levels but can also serve as an independent indicator for more refined patient stratification management; as a complementary tool to TMB, PRG scores can provide more targeted guidance for individualized treatment.

[0055] 7. High-risk LUADs exhibit an immunosuppressive tumor microenvironment. The characteristics of the immune microenvironment are crucial for tumor progression and immunotherapy response. Immunological differences exist between high- and low-risk groups in the tumor microenvironment, immune checkpoint mapping, and functional pathways. Figures 14-16 As shown in the image. A heatmap of differential immune infiltration is presented. Figure 14The results showed that multiple immune deconvolution algorithms were used to quantify the abundance of different immune cell populations in high- and low-risk tumors, with red representing high infiltration and white representing low infiltration. In the high-risk group, multiple immune cells were significantly upregulated, including regulatory T cells (Tregs), M2 macrophages, and myeloid suppressor cells (MDSCs), suggesting that they have significant immunosuppressive characteristics, constitute an "immune cold" microenvironment, and may inhibit anti-tumor immune responses, promote tumor progression, and weaken the therapeutic effect.

[0056] In contrast, the low-risk group was enriched with multiple anti-tumor immune cell types, such as B cells and CD8 cells. + The presence of T cells and activated dendritic cells suggests a stronger immune response. For example... Figure 15 As shown, further expression analysis of immune checkpoint molecules supports the above conclusions: the low-risk group exhibited higher immunotherapy response potential across multiple immune subtypes, including CTLA-4. + / PD-1 + The most significant differences were observed in the double-positive population (this subgroup is often considered a key target for the efficacy of immune checkpoint inhibitors). Furthermore, enrichment analysis of immune function-related pathways, such as... Figure 16 The chart shows the correlation between oncogenic pathways and immune-related pathways. The left heatmap represents the correlation between these pathways (color gradients represent positive or negative correlation coefficients). The right heatmap shows the activity levels of high- and low-risk tumors in various steps of the cancer immune cycle (darker colors indicate higher activity; green and purple lines represent positive and negative regulatory relationships, respectively). The low-risk group showed higher activity in multiple key immune pathways, including antigen presentation, T cell activation, and interferon response. In summary, PRG risk assessment is closely related to the characteristics of the immune microenvironment: high-risk groups often exhibit an immunosuppressive "cold tumor" phenotype with poor immunotherapy response; low-risk groups have "hot tumor" characteristics and are more likely to benefit from immunotherapy. Therefore, PRG risk assessment has both prognostic value and potential application value in immunotherapy stratification, providing a potential basis for screening patients for immunotherapy.

[0057] 8. Single-cell atlas reveals cellular heterogeneity and immune complexity in LUAD. like Figure 17 As shown in Figure A, the bubble chart illustrates the average expression levels and positive rates of proteasome-related genes (especially PSMD11) in different immune and tumor cells, confirming that these genes are elevated in both immune and malignant cells, suggesting their potential key role in tumor immune regulation and tumor progression. Single-cell RNA sequencing analysis of LUAD tumor tissue identified 40,218 cells, divided into 30 different subpopulations. This analysis revealed the cellular heterogeneity of the tumor samples, providing high-resolution cell type information for subsequent studies, such as... Figure 17As shown in B and C. Based on cell subtype annotation, this invention further identifies 14 major cell lineages that play important roles in tumor biology and the immune microenvironment. Within the framework of a risk assessment model, this invention visualizes core genes and risk assessment at the single-cell level using UMAP, such as... Figure 17 As shown in Figure D, the results indicate that these core genes are mainly enriched in cell types associated with immunosuppression (such as Tregs) and malignant cells. Figure 17 As shown in Figure E, the violin plot further illustrates the risk assessment distribution of each cell type: the risk assessment of malignant cells and immunosuppressive populations such as Tregs is significantly higher than that of other cell types. These findings support the crucial role of these cell populations in immune escape, immunosuppression, and tumor progression, reveal the complex relationship between the intrinsic cellular heterogeneity of LUAD and the immune microenvironment, and further confirm the functional roles of core genes in tumor and immune cells.

[0058] 9. High-risk cell populations exhibit enhanced immunosuppressive signaling and intercellular communication. Intercellular communication analysis revealed that immunosuppressive pathways were significantly upregulated in high-risk cell populations, and single-cell analysis showed how model risk assessment affects immune infiltration and intercellular communication. Figures 18-20 As shown. Its signal network activity is enhanced, such as Figure 18 and Figure 19 As shown in A and B, these high-risk cells exhibit prominent immune escape characteristics and demonstrate strong bidirectional signal transduction capabilities, manifested as enhanced signal transmission and reception, such as... Figure 19 As shown in C, this indicates that high-risk cells have a stronger immunosuppressive effect in the tumor microenvironment and regulate the immune response through effective intercellular communication. Figure 20 As shown, the left panel represents the frequency of communication between sender (row) and receiver (column) cell types, with red indicating increased interaction in the high-risk group and blue indicating decreased interaction. The right panel represents the communication strength of the same cell pair, with red indicating stronger signaling in the high-risk group and blue indicating weaker signaling. The heatmap further compares the quantity (left) and intensity (right) of communication between different cell sources in the high / low-risk groups: in the high-risk group, malignant cells, acting as signal transducers, communicate more frequently and strongly with various immune cells, such as T cells, macrophages, and dendritic cells. This suggests that high-risk cell populations not only promote tumor progression but may also exacerbate the formation of an immunosuppressive environment through more frequent communication with immune cells, thereby promoting tumor immune escape. These findings indicate that high-risk assessment not only reflects the degree of tumor cell progression but also reveals their central role in regulating intercellular communication and inducing immunosuppression. Further analysis provides new theoretical support for the development of combined immunotherapy strategies, particularly interventions targeting the immunosuppressive microenvironment, which have significant clinical implications.

[0059] 10. High expression of PSMD11 in LUAD tissues and its prognostic correlation Based on the above cohort study results, this invention further validated the findings at the animal and histological levels regarding the key proteasome gene PSMD11. At the histological and molecular levels, box plots and paired sample analysis from the TCGA cohort showed that PSMD11 expression in tumor tissues was significantly higher than in normal or adjacent tissues, indicating a clear overexpression trend of PSMD11 in lung cancer. Figure 21 As shown in A and B. Furthermore, as... Figure 21 The survival curves of C in the study showed that patients with high PSMD11 expression had significantly lower overall survival, indicating that high PSMD11 levels are closely associated with poor prognosis.

[0060] In addition, such as Figure 21 Nonoline analysis of D in the model indicates that PSMD11, as an independent variable, significantly contributes to survival prediction, suggesting its potential application in clinical prediction models within a multivariate prediction model integrating clinical features. Furthermore, as... Figure 21 Immunohistochemical staining of E in lung adenocarcinoma tissue showed a significant increase in PSMD11 staining intensity, while the signal was weaker in adjacent tissues, suggesting histological upregulation. Figure 21 Western blot analysis of F in the sample further confirmed this trend, indicating that molecular validation is consistent with histological findings. Protein quantification also clearly showed a significant increase in PSMD11 expression in the tumor group, demonstrating the statistical robustness of the experimental data. Figure 21 As shown in G in the figure. Taken together, these findings systematically demonstrate that PSMD11 is widely upregulated in lung cancer and is closely associated with poor patient prognosis, providing multifaceted evidence for its tumor-promoting function and prognostic value. Furthermore, this also confirms the immunoassay results showing abnormal PSMD11 expression in tumor cells.

[0061] 11. In vitro experiments showed that PSMD11 knockdown inhibited proliferation and promoted apoptosis. In in vitro experiments, PSMD11 knockdown significantly inhibited the growth and invasive potential of lung cancer cells and induced apoptosis. Figure 22 As shown in A and B, clonogenic assays revealed that PSMD11 knockdown significantly reduced the number of clones in A549 and H446 cells, suggesting a decline in long-term viability. Transwell assays further demonstrated that PSMD11 knockdown effectively inhibited the migration and invasion abilities of both cell lines. Figure 22 As shown in C and D in the diagram. EdU fluorescent staining revealed a significant reduction in DNA synthesis and inhibition of cell proliferation, as shown in... Figure 22As shown in E in the figure. The CCK-8 experiment showed that the growth curve of the shPSMD11 group shifted downwards overall within 1 to 5 days, and the proliferation rate decreased significantly, as shown in E. Figure 22 The F in the image. Meanwhile, TUNEL imaging showed a significant increase in apoptosis rates in A549 and H446 cells, such as... Figure 22 The results comprehensively demonstrate the proteasome gene PSMD11's proliferative, invasive, and anti-apoptotic effects in lung cancer cells, covering multiple aspects such as colony formation, migration / invasion, proliferation, and apoptosis.

[0062] 12. In vivo validation of the tumor-suppressive effect of PSMD11 knockdown In in vivo validation experiments, PSMD11 knockdown also exhibited a significant anti-tumor effect. For example... Figure 23 As shown in Figure A, the volume curves in the subcutaneous tumor formation experiment (Figure C) show that tumor growth in the shPSMD11 group was significantly slowed. Figure 23 As shown in B and C, tumor appearance and final weight further confirm that PSMD11 knockdown mice have smaller tumors and lighter weights. Figure 23 The D in the assay was used to assess the proliferation and apoptosis of tumor cells. Immunofluorescence analysis showed a significant decrease in Ki-67 expression, indicating a decline in tumor cell proliferation activity; TUNEL staining showed an increased proportion of apoptotic cells, indicating enhanced apoptosis. Overall, in vivo knockdown of PSMD11 slowed tumor growth, reduced proliferation, and promoted apoptosis, further validating its pro-tumor role in the development and progression of lung cancer. Immunofluorescence analysis showed that PSMD11 knockdown significantly reduced Ki67 expression, indicating reduced tumor cell proliferation.

[0063] Lung cancer is one of the most common and deadliest cancers worldwide. Due to its high metastasis and recurrence rates, patients typically have poor prognoses, and there are currently no safe and effective treatments. Proteasome-associated genes (PRGs) are key molecular regulators of intracellular protein homeostasis, involved in regulating and driving the development and progression of various malignant tumors. Based on this, this invention utilizes machine learning to develop a risk model for predicting the prognosis of patients with lung adenocarcinoma (LUAD). This invention identifies PSMD11 and PSMD12 as key genes driving LUAD progression and assesses the risk of death in LUAD patients. The results show that the model performs well and has clinical applicability. Kaplan-Meier survival analysis and receiver operating characteristic (ROC) analysis confirmed the model's stratification ability in high-risk patients, with significantly lower survival rates in the high-risk group compared to the low-risk group. Furthermore, comprehensive analysis of TCGA, GEO, and internal single-cell datasets indicates that the model accurately reflects the characteristics of the tumor immune microenvironment, potential oncogene mutations, and the effectiveness of immunotherapy. Therefore, this invention suggests that the PRG risk prediction model has the potential to guide personalized stratification of early-stage LUAD and provide clinical decision support.

[0064] Proteasome-related genes (PRGs) are a group of structural and regulatory genes of the 26S proteasome. These genes collectively support the normal function of the ubiquitin-proteasome system (UPS) in substrate recognition, unfolding, transport, and degradation, forming the basis of cellular protein homeostasis and signal transduction. Using a univariate Cox proportional hazards regression model, this invention identified 13 PRGs significantly associated with the prognosis of LUAD patients, among which PSMD11, PSMB7, and PSMA3 were identified as key molecular signatures. Existing research indicates that PRGs and their pathways can serve as risk biomarkers for various cancers. PSMD11 is crucial for maintaining protein homeostasis and regulating the cell cycle and chromatin remodeling. Its high expression is associated with poor prognosis in breast and pancreatic cancer. PSMB7 can affect the regulation of proliferation and migration in multiple myeloma and breast cancer. PSMA3 is associated with MYC pathway activation and tumor progression in malignant solid tumors. The 19S subunit can predict the prognosis of acute myeloid leukemia.

[0065] The mechanism of action of PRGs in tumor development and progression can be summarized as follows: ① UPS mediates the binding of cyclin / CDK inhibitors to p53 pathway proteins, promoting the continuous proliferation and anti-apoptosis of tumor cells.

[0066] ② By degrading IkB, NF-κB is released, which exacerbates the release of inflammatory factors and the sustained activation of inflammatory pathways.

[0067] ③ Upregulation of PRGs can reduce the protein toxicity of cancer cells, thereby providing them with a growth advantage.

[0068] ④ Immunoplasminosomal inhibition of MHC-I antigen processing and presentation affects immunogenicity and immunotherapy response. Furthermore, key features of PRGs are significantly correlated with the clinical characteristics of LUAD. At the single-gene level, PSMD11 is upregulated in LUAD and is associated with poor prognosis, lymph node metastasis, and immune escape phenotype. High expression of immunoproteasomes is associated with better survival and immunotherapy benefit in LUAD patients, confirming the close link between PRGs and the tumor immune microenvironment. Therefore, this invention suggests that a PRG risk assessment model can effectively identify specific prognostic features of LUAD. PRGs can reflect antigen processing, MHC-I presentation, and CD8+. +PSMD11, a 19S regulatory subunit of the 26S proteasome, plays a crucial role in T-cell recognition and other immune processes, and can serve as a complementary immune biomarker, providing new insights for early stratification of high-risk LUAD patients and assisting in clinical management. It is a key component of proteasome assembly and stability. Vilchez et al. found that upregulation of PSMD11 enhances the assembly efficiency and degradation activity of the 26S proteasome, thereby maintaining intracellular protein homeostasis. This property allows tumor cells to better buffer protein toxicity under high-synthetic and high-stress environments, and promotes sustained proliferation by accelerating the turnover of cell cycle regulatory molecules and inhibiting apoptosis signals, indicating that PSMD11 supports tumor cell survival. Huang et al. found that PSMD11 is abnormally overexpressed in various malignant tumors, including breast cancer, pancreatic cancer, and lung adenocarcinoma, and is closely associated with poor prognosis. In lung adenocarcinoma, Huang et al. found that PSMD11 overexpression not only promotes tumor cell proliferation and migration but is also associated with immune escape phenotype and poor survival outcomes, suggesting it as a potential prognostic biomarker and therapeutic target. Furthermore, analysis of the Human Protein Atlas (HPA) database revealed that PSMD11 expression patterns were significantly associated with poor survival across multiple cancer types, further supporting its potential as a pan-cancer prognostic factor. In this study, in vitro experiments showed that knockdown of PSMD11 significantly inhibited LUAD cell proliferation and induced apoptosis. In a nude mouse subcutaneous tumor model, downregulation of PSMD11 significantly slowed tumor growth and reduced Ki-67 expression, confirming the previously demonstrated pro-tumor effect of PSMD11. Therefore, this invention suggests that PSMD11 is not only a key driver of LUAD progression but also a novel stratification predictor. LUAD is the most common type of non-small cell lung cancer (NSCLC), primarily occurring in the bronchial mucosa or bronchial glands, and is a malignant lung tumor.

[0069] The etiology of lung cancer (LUAD) is complex and diverse, with smoking, air pollution, prior lung disease, and germline pathogenic variations all contributing to its development. It results from genomic damage caused by the combined effects of exogenous carcinogens and genetic factors. The pathological progression of LUAD is driven by stimuli that induce driver mutations in oncogenes such as EGFR, KRAS, and BRAF, and inactivate tumor suppressor factors such as TP53, STK11, and NF1. These changes promote tumor cell proliferation and immune escape by persistently activating the RTK–RAS–MAPK and PI3K–AKT–mTOR pathways, while also impairing DNA damage responses, metabolic homeostasis, and anti-tumor immunity, thus exacerbating the severity, metastasis, and recurrence of LUAD. Consequently, patients exhibit immunocold tumor characteristics, leading to poor prognosis and resistance to immunotherapy. These findings result in significant genomic instability in LUAD, not only exacerbating cancer progression and treatment resistance but also accelerating tumor heterogeneity. The loss of tumor suppressor factors impairs G1 / S and metabolic checkpoints, promoting individual resistance to single inhibitors, leading to widespread drug resistance and a tendency for tumor recurrence. Abnormal oxidative stress leads to sustained upregulation of NRF2, enhancing the antioxidant activity and metabolic reprogramming of cancer cells, helping tumors survive in a highly reactive oxygen species environment, and promoting gene mutations and immune escape. Furthermore, a significant dose-response relationship exists between smoking exposure and tumor microbiome size (TMB), and mutations in driver genes such as EGFR and KRAS are also associated with elevated TMB. These changes not only promote genomic instability and tumor heterogeneity but also facilitate immune metastasis. This is the reason for elevated TMB in high-risk LUAD patients. Therefore, this invention combines TMB prognostic assessment with PRG scores to predict the prognosis of LUAD patients. This comprehensive indicator may improve the accuracy of risk assessment and clinical decision-making.

[0070] Previous studies have shown that high-risk LUAD patients exhibit a heterogeneous and immunosuppressive tumor microenvironment. At the single-cell level, this invention validates that high PRG scores are primarily concentrated in immunosuppressive cell subsets. High-risk patients are enriched in immunosuppressive cell populations such as regulatory T cells (Tregs), M2 macrophages (TAMs), and myeloid-derived suppressor cells (MDSCs). These cell populations not only weaken the activity of the immune microenvironment but also impair antitumor responses, leading to immunotherapy failure. In LUAD, Tregs inhibit CD8 by overexpressing co-inhibitory receptors such as CTLA-4 and TIGIT, and by secreting inflammatory cytokines such as IL-10 and TGF-β. + T cells and activated dendritic cells enhance the function of immune-stimulating cells. Tregs also continuously promote the upregulation and expression of CD39 / CD73, thereby activating the A2A receptor and generating an immunosuppressive response.

[0071] Tumor-associated tumor cells (TAMs) overexpress PD-L1 and secrete cytokines to suppress T cell recruitment. SPP1 macrophages enhance tumor cell signaling by binding to CD44 and integrins, thereby suppressing T cell function. Myeloid suppressor cells can secrete tumor-derived IL-6, directly inhibiting T cell receptor signaling, proliferation, and differentiation, and promoting immune escape. Furthermore, high-risk cell populations can enhance the transmission of immunosuppressive signals and intercellular communication. Deletion of tumor suppressor genes and activation of KEAP1 / NRF2 inhibit antigen presentation, leading to persistent activation of myeloid-related pathways and amplification of cellular immunosuppressive communication networks. These mechanisms collectively explain the cellular heterogeneity and immune complexity of high-risk LUADs, indicating that PRG risk assessment can effectively reflect changes in the tumor immune microenvironment.

[0072] While PRG risk assessment models have proven reliable for predicting LUAD, single predictive systems still cannot fully reflect changes in patient physical characteristics in clinical practice. Therefore, this invention aims to combine PRGs with existing indicators such as PD-L1 and TMB in the future. By integrating multiple technologies such as proteomics, metabolomics, and spatial transcriptomics, this invention can further explore the signaling network by which PRGs regulate tumor immunity and provide personalized immunotherapy strategies for LUAD patients. Furthermore, in terms of therapeutic translation, drugs that inhibit 19S deubiquitinase and NRF1 can be developed to reduce the high heterogeneity of LUAD and mitigate damage to the immune microenvironment. This invention also considers drug-induced enhancement of MHC-I antigen presentation responses, combined with immune checkpoint inhibitors. Additionally, PSMD11 and PSMB7 can be used for network pharmacology studies and trial enrollment to prospectively validate combination strategies and conduct clinical trials.

[0073] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

[0074] The amino acid sequence involved in this invention is as follows: SEQ ID No. 1: Name: PSMD11-F primer gene sequence Sequence type: other DNA Biological origin: synthetic construct 5'-ACGAAGCTGCTCTGGAAACA-3'.

[0075] SEQ ID No. 2: Name: PSMD11-R primer gene sequence Sequence type: other DNA Source: synthetic construct 5'-GTCTCGACACGCAGAAGACA-3'.

Claims

1. An application of a proteasome-related gene, characterized in that, Application in gene prognostic risk assessment and construction of a gene prognostic risk assessment model for lung adenocarcinoma.

2. The application according to claim 1, characterized in that, The proteasome-related genes mentioned are those that are significantly associated with overall survival. The association features are those that change in the risk assessment of proteasome-related genes due to changes in clinical characteristics, tumor mutation burden, and immune characteristics in the lung adenocarcinoma cohort.

3. A method for constructing a gene prognostic risk assessment model based on proteasome-related genes, characterized in that, S1. Integrate the data from the existing lung adenocarcinoma genome atlas cohort and the data from five independent gene expression comprehensive database cohorts to form a gene dataset; S2. Cox regression analysis was performed on the gene dataset to identify proteasome-related genes. Functional enrichment analysis was performed on the proteasome-related genes, and the association features between the risk assessment of proteasome-related genes and the lung adenocarcinoma genome map cohort data were obtained. The association features were used as key molecular features. S3. Utilizing key molecular features, various algorithms are used to establish corresponding risk assessment models through machine learning. Data from the Cancer Genome Atlas cohort is used as the training set, and data from the Gene Expression Comprehensive Database cohort is used as the validation set. The established risk assessment models are trained using the training set and then validated using the validation set. The model with the best consistency index is selected as the optimal model, which is the final lung adenocarcinoma gene prognostic risk assessment model.

4. The method for constructing a gene prognostic risk assessment model based on proteasome-related genes according to claim 3, characterized in that, Step S1 involves: acquiring RNA expression profile data from the existing platform's genomic atlas cohort for lung adenocarcinoma and five independent gene expression comprehensive database cohorts; normalizing the RNA expression profile data to obtain the final gene dataset.

5. The method for constructing a gene prognostic risk assessment model based on proteasome-related genes according to claim 3, characterized in that, The genomic atlas cohort data for lung adenocarcinoma used data from the UCSC Xena platform; the five independent gene expression comprehensive database cohorts used five different gene datasets from the GEO-GSEx cohort, and the gene dataset numbers GSE and sample numbers n in the GEO-GSEx cohorts used were as follows: GSE13213, n=117; GSE29016, n=68; GSE30219, n=293; GSE31210, n=226; GSE42127, n=133.

6. The method for constructing a gene prognostic risk assessment model based on proteasome-related genes according to claim 3, characterized in that, In step S2, the proteasome-related genes are those that have a significant correlation with overall survival. The correlation features are those that change in the risk assessment of proteasome-related genes due to changes in clinical characteristics, tumor mutation burden, and immune characteristics in the lung adenocarcinoma cohort.

7. The method for constructing a gene prognostic risk assessment model based on proteasome-related genes according to claim 6, characterized in that, The proteasome-related genes mentioned are mainly one or more combinations of USP5, UBC, PSMD2, PSMD12, PSMD11, and CDC20.

8. The method for constructing a gene prognostic risk assessment model based on proteasome-related genes according to claim 3, characterized in that, Step S3 is as follows: S3.1 Utilizing key molecular features, using data from the Cancer Genome Atlas as the training set, and employing multiple ensemble machine learning methods in the Cancer Genome Atlas cohort to establish various risk assessment models; S3.

2. Data from the gene expression comprehensive database is used as a validation set for external validation, and the risk assessment model is evaluated for its risk prediction performance to obtain risk samples. S3.3 Calculate the risk assessment of the risk sample, and divide the sample into high-risk group and low-risk group according to the median of the risk assessment; Kaplan-Meier survival analysis was used to assess the survival differences between the high-risk and low-risk groups in order to determine the predictive accuracy of the risk assessment model and obtain the risk assessment model with the optimal consistency index. S3.4 Select the risk assessment model with the best consistency index as the optimal model, which is the final prognostic risk assessment model.

9. The application of the gene risk assessment model constructed by the method for constructing a gene prognostic risk assessment model based on proteasome-related genes as described in any one of claims 3-8, characterized in that, Applications in assisting risk stratification, prognosis assessment, and immune prediction in patients with lung adenocarcinoma.

10. A machine learning apparatus for implementing the method of any one of claims 3-8, characterized in that, Processor, used for regression analysis and machine learning model training; Storage device used to store gene datasets; A display screen is used to present the modeling results of the lung adenocarcinoma gene prognostic risk assessment model; A computer program, which is stored in a memory and executed by a processor, wherein when the computer program is executed by the processor, it implements the method for constructing a gene prognostic risk assessment model based on proteasome-related genes as described in any one of claims 3-8.