A method, device, medium, and procedure for predicting molecular subtypes of lung adenocarcinoma.
By integrating multi-omics datasets and using machine learning models to identify molecular subtypes of lung adenocarcinoma, the problem of treatment ineffectiveness caused by molecular heterogeneity of lung adenocarcinoma was solved, and accurate prediction and targeted therapy of LUAD-C2 subtype were achieved.
Patent Information
- Application Number
- CN202411873116.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-12-18
AI Technical Summary
The molecular heterogeneity of lung adenocarcinoma renders traditional treatments ineffective, creating an urgent need to identify different molecular subtypes and potential therapeutic targets through multi-omics data integration.
By integrating genomics, transcriptomics, proteomics, and epigenomics data, machine learning models are used to identify LUAD-C2 subtypes and non-LUAD-C2 subtypes. The subtypes are classified using key characteristic genes such as NDNF, CCNA2, and RNASE1, and the apoptosis status is output to guide targeted therapy.
It has enabled accurate prediction of molecular subtypes of lung adenocarcinoma, revealed the high apoptosis characteristics of LUAD-C2 subtype tumor cells, provided targeted therapy strategies, and improved treatment efficacy.
Smart Images

Figure CN119811501B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent healthcare, and more specifically, to a method, device, medium, and program product for predicting molecular subtypes of lung adenocarcinoma. Background Technology
[0002] Lung adenocarcinoma (LUAD) is a common form of non-small cell lung cancer (NSCLC), accounting for a large proportion of all lung cancer cases. Originating from glandular cells in the lungs responsible for secreting mucus and other fluids, LUAD has a complex molecular structure and diverse clinical presentations. This complexity presents challenges for its diagnosis and treatment. Despite advances in radiotherapy, chemotherapy, and surgical interventions, the complexity of LUAD necessitates a deeper understanding to develop effective treatment strategies, thus requiring comprehensive research and analysis in this field. The molecular landscape of LUAD is highly complex, characterized by widespread gene mutations, gene expression variations, and multiple epigenetic alterations. This heterogeneity not only complicates our understanding of the disease's pathogenesis but also poses significant challenges to developing effective targeted therapies. Traditional "one-size-fits-all" treatments often fail due to the unique molecular characteristics of each patient's tumor. Therefore, a comprehensive analysis of the molecular landscape of LUAD is urgently needed. Such analysis will involve integrating multi-omics data—genomics, transcriptomics, proteomics, and epigenomics—to obtain a holistic view of the disease. This comprehensive approach is crucial for identifying specific molecular subtypes of LUAD, understanding the underlying mechanisms driving each subtype, and discovering potential therapeutic targets.
[0003] Recent advances in multi-omics technologies have ushered in a new era in our understanding of diseases such as LUAD. Multi-omics refers to the collective analysis and integration of datasets from various biological levels, including genomics, transcriptomics, proteomics, and epigenomics. Each "omics" layer provides unique insights into the cellular and molecular mechanisms of cancer. Genomics reveals variations in DNA sequence and structure, providing a blueprint for gene mutations and alterations. Transcriptomics examines RNA and gene expression patterns, providing a dynamic view of gene activity and regulation in response to various conditions. Proteomics delves into the proteome—the entire set of proteins that produce it—revealing their functions, modifications, and interactions. Finally, epigenomics focuses on epigenetic changes, such as DNA methylation and histone modifications, which affect gene expression without altering the DNA sequence. Integrating these diverse datasets through multi-omics analysis provides a comprehensive understanding of LUAD, enabling researchers to link genomic alterations with changes in gene expression, protein function, and epigenetic modifications, thus providing a holistic view of cancer biology. This comprehensive approach is crucial for dissecting the heterogeneity of LUAD because it reveals the multifaceted interactions and regulatory mechanisms driving the disease. By integrating multi-omics data, researchers can identify different molecular subtypes of LUAD, explore the pathogenesis of the disease, and discover new therapeutic targets. Summary of the Invention
[0004] In view of the above problems, this invention provides a method for predicting molecular subtypes of lung adenocarcinoma (LUAD). By integrating various multi-assembly datasets, it dissects the complex molecular landscape of LUAD and facilitates the identification of different molecular subtypes. This comprehensive approach utilizes complementary insights from genomics, transcriptomics, proteomics, and epigenomics to gain a more detailed understanding of the heterogeneity of LUAD, elucidate the pathogenesis of the disease, and improve its prognosis and treatment.
[0005] This application (first aspect) discloses a method for predicting molecular subtypes of lung adenocarcinoma, comprising:
[0006] Obtain the genetic data of the sample to be tested;
[0007] The gene data of the sample to be tested are processed to obtain key feature genes; the key feature genes include: NDNF, CCNA2 or RNASE1;
[0008] The key characteristic genes are input into the classifier to obtain classification results of LUAD-C2 subtype or non-LUAD-C2 subtype; the LUAD-C2 subtype tumor cells have high apoptosis; the non-LUAD-C2 subtype tumor cells have low apoptosis.
[0009] Furthermore, the classifier is obtained in the following ways:
[0010] Obtain the gene dataset for the training set;
[0011] Clustering the gene dataset yields the LUAD-C2 subtype and non-LUAD-C2 subtype;
[0012] Differential gene expression was identified by differential gene expression analysis between the LUAD-C2 subtype and non-LUAD-C2 subtype.
[0013] Feature selection was performed on the differentially expressed genes to obtain key feature genes;
[0014] The key feature genes are input into a machine learning model to obtain a predicted classification result, which is then compared with the classification label. The model is optimized based on the comparison result to obtain the classifier.
[0015] Furthermore, based on the classification results of the LUAD-C2 subtype, the results of the TME with high immune activity of the test sample cells are output; based on the classification results of the non-LUAD-C2 subtype, the results of the TME with low immune activity of the test sample cells are output.
[0016] Furthermore, the method also includes: the expression level of NDNF is high and the expression level of CCNA2 is low in the LUAD-C2 subtype.
[0017] Furthermore, the key characteristic gene also includes CIP2A; the expression level of CIP2A is low in the LUAD-C2 subtype;
[0018] Optionally, the key characteristic gene further includes: MYBL2; the expression level of MYBL2 is low in the LUAD-C2 subtype;
[0019] Optionally, the key characteristic gene also includes MELK; the expression level of MELK is low in the LUAD-C2 subtype.
[0020] Furthermore, the key characteristic gene also includes CASP-3; the expression level of CASP-3 in the LUAD-C2 subtype is low;
[0021] Furthermore, the key characteristic gene also includes RNASE1, and the expression level of RNASE1 is high in the LUAD-C2 subtype.
[0022] The second aspect of this application discloses a system for predicting molecular subtypes of lung adenocarcinoma, comprising:
[0023] Acquisition module: Acquires the gene data of the sample to be tested;
[0024] Feature extraction module: processes the sequencing data of the sample to be tested to obtain key feature genes; the key feature genes include: NDNF, CCNA2 or RNASE1;
[0025] Prediction module: used to input the key feature genes into the classifier to obtain the classification result of LUAD-C2 subtype or non-LUAD-C2 subtype; the LUAD-C2 subtype has high apoptosis; the non-LUAD-C2 subtype has low apoptosis.
[0026] A third aspect of this application discloses a computer device, the device comprising: a memory and a processor; the memory being used to store program instructions; the processor being used to invoke the program instructions, which, when executed, are used to perform the steps of the method described above.
[0027] The fourth aspect of this application discloses a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.
[0028] The fifth aspect of this application discloses a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0029] This application has the following beneficial effects:
[0030] (1) This application obtained two different molecular subtypes through the study of lung adenocarcinoma and predicted the subtypes by gene expression levels;
[0031] (2) This application demonstrated through molecular experiments that the LUAD-C2 subtype tumor cells have low viability and high apoptosis by inducing overexpression of key characteristic genes of the subtype;
[0032] (3) Targeted therapy for the subtype identified in this application is a potential treatment for the subtype of lung adenocarcinoma. Attached Figure Description
[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0034] Figure 1 This is a schematic diagram of the method flow provided in the first aspect of the present invention;
[0035] Figure 2 This is a schematic diagram of a program product provided in the second aspect of the present invention;
[0036] Figure 3 This is a schematic diagram of a computer device provided in an embodiment of the present invention;
[0037] Figure 4 This is a schematic diagram of the architecture of an exemplary computing device provided in an embodiment of the present invention;
[0038] Figure 5 This is a schematic diagram of the storage medium provided in an embodiment of the present invention;
[0039] Figure 6 This invention provides a multi-omics integrated consensus subtype for LUAD: (A) Plotting the clustering prediction index (CPI, blue dashed line) and gap statistics (red solid line) to identify the optimal number of clusters from multi-omics data. The elbow method indicates the optimal number of clusters based on the trends observed in these indices; (B) A comprehensive heatmap of the consensus set subtypes, including mRNA, lncRNA, miRNA, DNA CpG methylation sites, and mutated genes; (C) A silhouette plot showing the two clusters (C1 and C2) and their silhouette widths, reflecting cluster confidence; an average silhouette width of 0.6 indicates a reasonable centralized structure of the data; (D) Clustering of MUC patients using 10 cutting-edge multi-omics clustering methods; (E) Consistent clustering matrices of three novel prognostic subtypes based on 10 algorithms; (F) Different survival outcomes between the two subtypes;
[0040] Figure 7 This invention provides a comprehensive molecular characterization and validation of a consensus subtype (CSs) of LUAD: (A) A heatmap showing the enrichment differences of key pathways between the two identified subtypes (CS1 and CS2); (B) A heatmap showing immune and matrix enrichment scores, as well as immune checkpoint gene expression in the two subtypes; (C) Recent Template Prediction (NTP) analysis confirming that patients in the META-LUAD cohort are divided into CS1 and CS2; (D) Consistency assessment of CMOIC and NTP classification; (E) Kaplan-Meier survival curves showing a significant difference in overall survival between CS1 and CS2; (FG) High consistency of CMOIC, NTP, and PAM classification methods; (H) Bar charts of biological process enrichment analysis associated with each subtype.
[0041] Figure 8This invention provides a comprehensive immune landscape analysis of a cancer patient cohort: a comprehensive bioinformatics analysis and prediction model of gene importance under multiple conditions; 22(A) PCA plot illustrating the distribution of individuals along the first two principal components in datasets GSE31210, GSE50081, GSE68465, and TCGA; (B) PCA scatter plot visualizing gene expression data, with ellipses representing confidence intervals revealing sample grouping and expression profile overlap; (C) Scatter plot of the negative logarithm of the Cox coefficient relative to various factors; dashed lines represent the threshold of statistical significance; (D) Feature selection frequency waterfall plot by selection frequency. Genes are ranked, and genes that are consistently important throughout the analysis are highlighted using color gradients; (E) Classifier stability plots show feature importance scores across multiple runs; (F) Model accuracy scatter plots combine bee colony plots and box plots to compare the predictive accuracy (C-index) of prognostic models; (G) Forest plots show the hazard ratios and 95% confidence intervals for each gene in the survival analysis of the TCGA-LUAD dataset; (H) Box plots illustrate the C-index distributions for various models, including CoxBoost and XGBoost; (I) Box plots illustrate the IBS for various models, including CoxBoost and XGBoost.
[0042] Figure 9 This invention provides a comprehensive immune landscape analysis of a cancer patient cohort: (A) Box plots illustrating the relative abundance of infiltrative immune cell types in patients belonging to high-risk and low-risk groups; (BC) Box plots of relative expression levels of 27 immune checkpoint profiles between high-risk and low-risk patients; (D) Color-coded scatter plots and box plots showing the risk score distribution of treatment responses (PR, PD, CR, SD); (E) Bar plots of immune checkpoint inhibitor responses; (F) Box plots comparing TIDE scores, IFNG scores, functional impairment scores, exclusion scores, TAM / M2 scores, and MDSC scores between high-risk and low-risk groups.
[0043] Figure 10This invention provides a comprehensive analysis comparing high-risk and low-risk LUAD patient groups: (A) Volcano plot and scatter plot of the correlation coefficient between 3S-MMR score and drug-available mRNA expression in the TCGA-LUAD cohort, obtained from Spearman's rank correlation analysis; red dots indicate significant positive correlation (p<0.05, Spearman's r>0.2); (B) Volcano plot and scatter plot of the correlation and significance between 3S-MMR score and drug target CERES score; green dots indicate significant negative correlation (p<0.05, Spearman's r>0.2). (r<-0.2); (C) Comparison of IC50 values between high and low 3S-MMR score groups for paclitaxel, gemcitabine, and cisplatin; (D) Spearman correlation analysis results of CTP-derived compounds and prism-derived compounds; Differential drug response analysis results of CTP-derived compounds and prism-derived compounds show that the lower the y-axis value of the box plot, the higher the drug sensitivity; Abbreviations: *p<0.05; **p<0.01; ***p<0.001;
[0044] Figure 11 This invention provides a comprehensive analysis of LUAD molecular subtypes and immune landscape: (A) a radar map comparing the molecular characteristics of high-risk (HR) and low-risk (LR) LUAD patients; (B) a heatmap showing the expression patterns of numerous genes in multiple patient samples, with hierarchical clustering used to group similar profiles; and (C) a heatmap of immune-related gene expression, showing that HR patients have higher levels of inhibitory checkpoints and lower HLA gene expression.
[0045] Figure 12 This invention provides a multi-faceted analysis of gene networks, immune cell interactions, and expression variability in LUAD: (A) GeneMANIA network visualization, depicting the functional network of selected genes; (B) Correlation matrix describing the relationship between various immune cells and their functional states; (C) Correlation matrix of key LUAD-related genes; (D) Correlation diagram of gene expression levels and immune infiltration scores; (E) Mutation frequency diagram of LUAD samples with key gene mutations.
[0046] Figure 13This invention provides an insight into single-cell data and ground truth (GT) activity: (A) Cell populations are divided into 32 clusters; (B) Samples are labeled with 12 unique cell types in the TME using UMAP; (C) Expression profiles of cell typing marker genes; (D) Expression levels of various genes (AURKB, CDK1, CDKN3, CDT1, CHEK1, DLGAP5, EXO1, FOSL1, KRT6A) in different cell types, including T cells, NK cells, epithelial cells, macrophages, monocytes, fibroblasts, mast cells, endothelial cells, plasma cells, and plasmacytoid dendritic cells (PDCs).
[0047] Figure 14 This invention provides an embodiment of the role of NDNF in the progression of LUAD and its potential as a therapeutic target: (A) Quantitative PCR analysis of the expression levels of CDT1, PLK1, RNASE1 and NDNF in normal and LUAD samples; (B) Western blot analysis of the NDNF protein level in normal and LUAD tissue samples; (C) Quantitative determination of the NDNF protein level in normal and LUAD tissue samples; (D) Western blot analysis confirms that the expression of NDNF in LUAD cells is reduced compared with normal lung cells; (E) Quantitative determination of the NDNF protein level in normal and LUAD tissue samples; (F) Evaluation of the effect of nndnf overexpression on LUAD cell viability using the CCK-8 assay; (G) Comparison of apoptosis rates between control (NC) and LUAD cells overexpressing nndnf (OE); (H) Western blot analysis of the cleaved Caspase-3, Caspase-3 and NDNF protein levels in control (CT) and LUAD cells overexpressing NDNF (OE); (1) Cleaved Quantitative analysis of Caspase-3 protein levels showed a significant increase in apoptosis markers;
[0048] Figure 15 This invention provides an embodiment of the differences in gene expression between the LUAD-C2 subtype and non-LUAD-C2 subtype; AF shows the expression differences of NDNF, MYBL2, MELK, CIP2A, CCNA2, and CASP3, respectively;
[0049] Figure 16The following are ROC curves of several LUAD-C2 subtype prediction models provided in the embodiments of the present invention: Figure A shows the ROC curve of the training set using NDNF and CCNA2 prediction, with the AUC values from top to bottom representing the accuracy of NDNF and CCNA2 prediction alone and jointly; Figure B shows the joint prediction curve and AUC value of NDNF+CCNA2+MYBL2 in the training set; Figure C shows the joint prediction curve and AUC value of NDNF+MYBL2+CCNA2+MELK in the training set; Figure D shows the ROC curve of the validation set using NDNF and CCNA2 prediction, with the AUC values from top to bottom representing the accuracy of NDNF and CCNA2 prediction alone and jointly; Figure E shows the joint prediction curve and AUC value of NDNF+CCNA2+MYBL2 in the validation set; Figure F shows the joint prediction curve and AUC value of NDNF+MYBL2+CCNA2+MELK in the validation set. Detailed Implementation
[0050] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0051] In some of the processes described in the specification, claims, and accompanying drawings of this invention, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as S101, S102, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0053] Figure 1 This is a schematic flowchart of a prediction method based on assessing molecular subtypes of lung adenocarcinoma provided by an embodiment of the present invention. Specifically, the method includes the following steps:
[0054] S101: Obtain the genetic data of the sample to be tested;
[0055] S102: Process the sequencing data of the sample to be tested to obtain key characteristic genes; the key characteristic genes include: NDNF, CCNA2, or RNASE1.
[0056] S103: Input the key feature gene into the classifier to obtain the classification result of LUAD-C2 subtype or non-LUAD-C2 subtype; the LUAD-C2 subtype has high apoptosis; the non-LUAD-C2 subtype has low apoptosis.
[0057] In some embodiments, we combined mRNA, long non-coding RNA (lncRNA) and microRNA (miRNA) expression profiles, genomic mutations and epigenomic DNA methylation data to develop an integrated consensus subtype of LUAD using 10 multi-omics integration strategies.
[0058] In this study, we used the Multi-omics Integration and Visualization (MOVICS) software package for cancer subtype analysis to screen gene signatures using the "getElites" function. For continuous variables (mRNA, lncRNA, miRNA, and methylation), we set the "method" parameter of the "getElites" function to "mad" to screen the top 1500 genes with the highest degree of variation.
[0059] Then, seven machine learning algorithms were benchmarked: CoxBoost, Elastic Network (Enet), Lasso, Cox Partial Least Squares Regression (plsRcox), Ridge, Supervised Principal Component Analysis (SuperPC), and eXtreme Gradient Boosting Survival (XGBoost). The hyperparameters of each algorithm were tuned using an internal 5x CV, and the performance of the best-tuned model was evaluated using an external 10x CV. Finally, the optimal model algorithm was determined by comparing the average of the Harrell's Concordance Index (C-index) and the Integrated Brier Score (IBS) after 10 test folds.
[0060] Immune Cell Infiltration Analysis and Immunotherapy Response Prediction: We used seven algorithms (“CIBERSORT”, “MCPcounter”, “EPIC”, “ESTIMATE”, “TIMER”, “quanTIseq”, and “IPS”) to estimate the infiltration level of immune cells in the tumor microenvironment and visualized it through a “complex heatmap”. Specifically, we used the R package CIBERSORT and the single-sample gene set enrichment analysis (ssGSEA) algorithm to determine the distribution of immune cell types and the enrichment score of each immune cell type between the high-risk and low-risk groups. Furthermore, we compared the expression levels of several important immune checkpoint genes between the two different scoring groups. The Tumor Immune Dysfunction and Exclusion (TIDE) algorithm was initially used to assess the potential differences in treatment response between the high-risk and low-risk groups. A higher TIDE score indicates reduced efficacy, emphasizing the negative correlation between the TIDE score and treatment effect.
[0061] Functional enrichment analysis was performed to identify specific biological pathways enriched in different clusters. We used the R package CIBERSORT and the single-sample gene set enrichment analysis (ssGSEA) algorithm to determine the distribution of immune cell types and the enrichment score of each immune cell type between the high and low CMLS groups. Furthermore, the expression levels of several important immune checkpoint genes were compared between the two different risk groups. Kyoto Genome Encyclopedia (KEGG) analysis was also performed to elucidate the functions of selected core genes.
[0062] Cluster Analysis: The clustering process is guided by two main metrics: the Cluster Predictive Index (CPI) and Gap Statistics. CPI is used to assess the internal consistency and segregation among potential clusters, providing a quantifiable measure of cluster quality based on within-cluster variance relative to between-cluster variance. On the other hand, Gap Statistics compares the logarithmically transformed within-cluster sums of squares for different k values with the expected value under a null reference distribution of the data.
[0063] Determining the Number of Clusters: The optimal number of clusters was determined using silhouette plots and Gap Statistics. Silhouette analysis involves calculating a score for each sample within a cluster, which measures how similar that sample is to samples in other clusters compared to samples in its own cluster. This metric ranges from -1 to 1, with high values indicating a good match with its own cluster and a poor match with neighboring clusters. Silhouette plots provide a visual representation of how each cluster distinguishes itself from other clusters, which helps determine the most coherent and unique clustering solutions. Gap Statistics further confirms these findings, providing a statistical basis for selecting the number of clusters that best captures the inherent groupings in the data.
[0064] Clustering Algorithms: To achieve robust clustering of multiple datasets, ten state-of-the-art clustering methods were employed, each effectively handling the complexity and dimensionality of the data. These methods include K-means, hierarchical clustering, Gaussian Mixture Model (GMM), DBSCAN, spectral clustering, affinity propagation, clustering, BIRCH, Mean Shift, and CURE. Each algorithm employs a unique approach to grouping the data: K-means divides the data into K clusters based on centroids; hierarchical clustering constructs nested clusters through stepwise merging or splitting; GMM models clustering as a mixture of Gaussian distributions; and DBSCAN groups dense points and identifies outliers in low-density regions. The selection was based on each method's ability to discover nonlinear patterns and handle varying cluster sizes, shapes, and densities. These methods were evaluated using performance metrics such as profile scores and the Davies-Bouldin index to ensure that the final clustering results are statistically valid and meaningful for subsequent biological interpretation.
[0065] Subtype Identification and Validation: Consensus clustering is employed to integrate the results of multiple clustering algorithms, enhancing the robustness and reproducibility of subtype identification. This method uses a consensus clustering matrix to aggregate clustering results from different methods, where each entry in the matrix represents the proportion of algorithms that assigned the same pairs of samples to the same cluster. Therefore, the consensus matrix provides an intuitive and quantitative measure of consistency across various methods, highlighting consistent grouping. The final subtype is determined based on the stability of cross-algorithm clustering, ensuring that the selected clusters represent true underlying biological patterns, rather than the product of any single clustering technique, as evidenced by high consensus values and low variability.
[0066] Validation Methods: The identified LUAD subtypes were validated using external cohorts (META-LUAD and IMvigor-LUAD) to confirm their generality and robustness. Two main methods were employed: Nearest Template Prediction (NTP) and Microarray Predictive Analysis (PAM). NTP involves comparing the molecular characteristics of new samples with predefined template characteristics representing each subtype to assign the most similar template to each sample. PAM, on the other hand, uses a class prediction strategy based on gene expression data to classify new samples into one of the established subtypes. These methods are crucial for demonstrating that the subtypes not only hold true in the original TCGA cohorts but also retain their prognostic relevance and unique molecular characteristics across different and independent patient populations. Statistical metrics such as classification accuracy, sensitivity, and specificity were calculated to evaluate the performance of subtype assignment in these validation cohorts.
[0067] Immune Microenvironment Analysis: To assess immune cell infiltration in the tumor microenvironment, several computational tools were used, each providing unique insights into the composition and function of immune cells in the LUAD context. CIBERSORT was used to deconvolve the cellular composition of complex tissues based on gene expression data to estimate the relative proportions of immune cell types within the tumor. MCPcounter utilized a set of marker genes specific to each cell type to estimate the abundance of various immune cell and stromal cell populations, thus providing quantification of these cells in the tumor sample. TIMER employed a similar approach but emphasized the influence of tumor purity, thereby adjusting the infiltration score to more accurately reflect the true immune status. Finally, quantISEQ was used to specifically quantify tumor-infiltrating lymphocytes using a machine learning model trained on digital cytology data.
[0068] Differential Expression Analysis of Immune Checkpoint Genes: Researchers analyzed the expression of immune checkpoint genes to understand their differential regulation among identified molecular subtypes. Immune checkpoint genes play a crucial role in regulating immune responses. Using RNA-Seq data, the expression levels of key checkpoint genes such as PD-1, CTLA-4, and LAG-3 were quantified. Statistical tests, including ANOVA and post-hoc Tukey tests, were used to detect significant differences in expression among subtypes. This analysis complements related studies that link expression patterns to clinical data such as treatment response and survival outcomes. Furthermore, the use of box plots and heatmaps for data visualization helps to visually understand the expression changes of checkpoint genes in different subtypes. This analysis not only highlights the potential mechanisms of immune evasion but also proposes targeted therapy strategies based on different LUAD subtype-specific immune profiles to improve the efficacy of immunotherapy.
[0069] Single-cell RNA sequencing and data analysis: LUAD and adjacent normal tissue samples were collected from surgical patients. Fresh tissue samples were immediately dissociated into single cells and digested with enzymes according to a pre-defined protocol. Single-cell suspensions were filtered to remove cell debris, and cell viability was assessed using trypan blue staining. Cells with a viability >85% were used for subsequent analysis.
[0070] Single-cell RNA sequencing libraries were prepared using a 10x Genomics Chromium platform according to the manufacturer's instructions. Single cells were captured in microdroplets, and barcoded mRNA was reverse transcribed to generate a cDNA library. Sequencing was performed on an Illumina platform, producing paired-end reads with an average depth of 50,000 reads per cell.
[0071] Raw sequencing data were processed using the Cell Ranger pipeline (v.3.1, 10x Genomics), aligning reads to the human reference genome (GRCh38) and generating a gene expression matrix. Low-quality cells with fewer than 200 detected genes or cells with mitochondrial gene content >10% were excluded from the analysis. Genes expressed in fewer than three cells were also removed.
[0072] After normalizing and logarithmically transforming the expression data, highly variable genes were selected for downstream analysis. Principal component analysis (PCA) was used for dimensionality reduction, and the optimal principal components were selected for clustering. Based on the Louvain algorithm, the Seurat R package (v.3.1) was used to cluster the cells. Uniform manifold approximation and projection (UMAP) was used to visualize the clustering results.
[0073] Clusters were annotated by comparing differentially expressed genes (DEGs) in each cluster with known cell type markers in published literature. Identified cell types included T cells, B cells, NK cells, epithelial cells (Epi), macrophages (Mac), monocytes (Mon), fibroblasts (Fib), dendritic cells (MDC), mast cells, plasma cells, endothelial cells (Endo), and plasmacytoid dendritic cells (PDC).
[0074] Cell culture: Normal lung cells and the LUAD cell line A549 were cultured in Dulbecco's Modified Eagle medium (DMEM) supplemented with 10% fetal bovine serum and 1% penicillin-streptomycin. Cells were maintained at 37°C under a humidified atmosphere containing 5% CO2. The medium was changed every 2-3 days, and cells were separated by co-passage with trypsin-edta solution at approximately 80-90%.
[0075] qPCR analysis: Total RNA was extracted from cell samples using an RNA extraction kit. The RNA was then reverse transcribed into cDNA using a reverse transcription kit. Quantitative PCR (qPCR) was performed using specific primers for NDNF and a SYBR Green master mixture. The relative expression level of NDNF was calculated using the 2^-ΔΔCt method, with GAPDH as an internal control.
[0076] Western Blot Analysis: Proteins were extracted from cell samples using lysis buffer. The extracted proteins were separated by SDS-PAGE and transferred to a PVDF membrane. The membrane was then probed with anti-NDNF primary antibody and incubated with enzyme-labeled secondary antibody. Protein bands were visualized using an enhanced chemiluminescence (ECL) detection system, and relative protein levels were quantified using image analysis software.
[0077] Adenoviral Vector Construction: An adenoviral vector containing full-length NDNF cDNA was constructed using standard cloning techniques. Under the control of a strong promoter, the NDNF cDNA was inserted into the adenoviral vector. A549 cells were then transduced using the adenoviral vector with appropriate transduction reagents. Transduction efficiency was monitored by the expression of the reporter gene contained in the vector.
[0078] CCK8 Assay: A549 cells overexpressing nndnf and those in control were seeded at a density of 5000 cells per well in 96-well plates. After allowing cells to adhere and grow for 24 hours, Cell Counting Kit-8 (CCK8) reagent was added to each well according to the manufacturer's instructions. The plates were incubated at 37°C for 2 hours. The absorbance of each well was measured at 450 nm using a microplate reader. Cell viability was assessed by comparing the absorbance values of nndnf-overexpressing cells with those of control cells.
[0079] Normal lung cell samples and the LUAD cell line A549 were cultured in Dulbecco modified Eagle medium (DMEM) supplemented with 10% fetal bovine serum (FBS) and 1% penicillin-streptomycin. Cells were maintained at 37°C in a humid atmosphere containing 5% CO2. The medium was changed every 2-3 days, and cells were separated using trypsin-EDTA solution, passaged when approximately 80-90% confluence was reached.
[0080] qPCR analysis: Total RNA was extracted from cell samples using an RNA extraction kit. The RNA was then reverse transcribed into cDNA using a reverse transcription kit. Quantitative PCR (qPCR) was performed using NDNF-specific primers and the SYBR Green master mix. The relative expression level of NDNF was calculated using the 2^-ΔΔCt method, with GAPDH as an internal control.
[0081] Western blot analysis: Proteins were extracted from cell samples using lysis buffer. The extracted proteins were separated by SDS-PAGE and transferred to a PVDF membrane. The membrane was then probed with an anti-NDNF primary antibody and incubated with an HRP-conjugated secondary antibody. Protein bands were visualized using an enhanced chemiluminescence (ECL) detection system, and relative protein levels were quantified using image analysis software. 13. Adenovirus vector construction: Adenovirus vectors containing full-length NDNF cDNA were constructed using standard cloning techniques. The NDNF cDNA was inserted into the adenovirus vector under the control of a strong promoter. A549 cells were then transduced with the adenovirus vector using appropriate transduction reagents. Transduction efficiency was monitored by the expression of the reporter gene contained in the vector.
[0082] CCK8 assay: NDNF-overexpressing and control A549 cells were seeded at a density of 5,000 cells per well in 96-well plates. After allowing cells to adhere and grow for 24 hours, Cell Counting Kit 8 (CCK8) reagent was added to each well according to the manufacturer's instructions. The plates were incubated at 37°C for 2 hours. The absorbance of each well was measured at 450 nm using a microplate reader. Cell viability was assessed by comparing the absorbance values of NDNF-overexpressing cells with those of control cells. Data were analyzed using GraphPad Prism or similar statistical software. Unpaired Student's t-tests were performed to compare between two groups. For multiple comparisons, one-way ANOVA was used, followed by post-hoc tests such as Tukey's or Bonferroni's multiple comparison tests. A p-value less than 0.05 was considered statistically significant. All experiments were repeated at least three times to ensure reproducibility, and data are expressed as mean ± standard error (SEM).
[0083] The analyses described in this study were supported by a comprehensive suite of software tools and packages, primarily within the R and Python ecosystems. Various statistical tests were applied to ensure the validity of the results from different types of data analyses: Kaplan-Meier analysis and Cox proportional hazards models: these methods were used to assess survival differences, with the log-rank test used to compare Kaplan-Meier curves to determine statistical significance. Analysis of variance and t-tests: used to compare gene expression among multiple groups; when significant ANOVA results were found, a post-hoc Tukey test was used to adjust for multiple comparisons. Mann-Whitney U and Wilcoxon sign-rank tests: these nonparametric tests are used when data are not normally distributed, applicable to comparing two independent or paired groups, respectively. The Benjamin-Hochberg procedure: this false discovery rate (FDR) control procedure was used to adjust p-values in the context of multiple hypothesis testing, particularly in gene expression and biomarker analyses, ensuring that the probability of Type I errors is minimized.
[0084] Results (1): Discovery of multi-omics consensus prognostic molecular subtypes of LUAD: In our comprehensive analysis of LUAD, we integrated multiple omics data types to identify different molecular subtypes. We independently identified two subtypes from the Cancer Molecular Understanding (MUC) TCGA LUAD cohort using 10 multi-omics ensemble clustering algorithms. The optimal number of clusters was determined using the Cluster Predictive Index (CPI). Figure 6 A), the contour plot shows that the clusters are structurally robust and well-separated. Figure 6C). Then, the clustering results were further combined with different molecular patterns using a consensus integration method, including transcriptome expression (mRNA, lncRNA, and miRNA), epigenetic methylation, and somatic mutations. Figure 6 B. Figure 6 D-6E).
[0085] Survival analysis revealed a significant difference in overall survival between the two subtypes, with the C1 (CS1) subtype associated with a significantly worse prognosis (p = 0.001). This finding highlights the prognostic significance of molecular subtypes, suggesting that patients with the CS1 subtype may require more aggressive or targeted treatment interventions compared to those with the CS2 subtype. Figure 6 F).
[0086] LUAD Integration Consensus Molecular Subtype Classification: Currently, molecular subtypes are mainly classified based on molecular expression levels, which may be related to specific biological functions. We used the Single Sample Gene Set Enrichment Analysis (ssGSEA) algorithm to explore the different molecular characteristics of these two CSs. The results showed that CS1 was significantly enriched in the cell cycle and DNA replication pathways, while CS2 was significantly enriched in FGFR3 co-expressed genes, PPARG, and WNTβ catenin pathways. Figure 7 A). Tumor immunity plays a crucial role in the occurrence and development of tumors. Therefore, we quantified the infiltration level of immune cells in the microenvironment and found that the infiltration of immune cells was relatively high in CS2 ( Figure 7 B).
[0087] Based on differential expression analysis between subtypes, the genes upregulated in C1 and downregulated in C2 include: MYBL2, UBE2C, TPX2, SLC2A1, CDC20, BIRC5, TOP2A, ANLN, KIF2C, TNNT1, MELK, CCNA2, CDCA5, FOXM1, TRIP13, HJURP, CEP55, RRM2, TK1, KIF4A, DLGAP5, KIFC1, CCNB1, and MMP12. , AURKB, CENPA, CDCA8, MAGEA3, LYPD3, NEK2, CCNB2, KRT6A, CDKN3, CDC6, CDK1, UCHL1, MKI67, COL11A1, AUR KA, CENPW, GJB2, DSP, HMGA1, PRR11, NDC80, CDC45, NUF2, KIF18B, NCAPG, GCLC, FAM83D, KPNA2, PLOD2, RPL3 9L, PLK1, TROAP, BUB1, UBE2S, TTK, FOSL1, SPAG5, EIF4EBP1, UBE2T, TXNRD1, ZWINT, CENPF, CDT1, PBK, CTHR C1, FAM83A, KIF11, PFN2, MAD2L1, CTSV, KIF23, BUB1B, CXCL10, PIMREG, MAGEA6, ASF1B, RAD51AP1, NCAPH, A RNTL2, MCM2, SERPINB5, KIF20A, CDA, GZMB, CCNE1, RACGAP1, EXO1, SLC16A1, MCM4, SKA1, SKA3, CHEK1, UHRF 1. GJB3, C15orf48, MCM10, GREM1, FSCN1, ORC1, GTSE1, ECT2, STEAP1, IQGAP3, PRAME, HILPDA, HMMR, CIP2A;
[0088] The C2 promoter genes include C16orf89, CACNA2D2, TMPRSS2, CYP4B1, NAPSA, and CFAP221. SELENBP1, ADGRF5, SCNN1B, DLC1, TMEM163, CTSH, SCTR, TMEM125, NFIX, IN MT, SFTPB, SUSD2, SLC22A3, PRDM16, HLF, PLA2G1B, ADH1B, VSIG2, GGTLC1. FCER1A, RNASE1, SCGB3A2, RHOBTB2, PLA2G4F, ABCA3, PIGR, C4BPA, SLC26A 9. SFTPD, B3GNT8, SLC22A31, SFTA2, ST3GAL5, GPRC5C, KCNQ1, C1orf116, V.S EGFD, PGC, PEBP4, ALPL, MFSD2A, FOLR1, RAP1GAP, NPC2, VWA2, PARM1, CAPN 8. LRRK2, C7, FOXA2, CD207, ATP13A4, CDKL2, C9orf152, SLC34A2, BTG2, CX CL17, ROS1, MGP, NKX2.1, SHE, KCNJ15, CPAMD8, GGT6, HSD17B6, MALL, DUOX 1, IRX3, IRX5, IRX2, DRAM1, CRYM, SFTPA1, LMO3, HABP2, CHRDL1, ELN, SFTP A2, C5orf38, LPL, WFDC2, FMO5, WIF1, MFAP4, AQP3, MAOA, CLU, HOPX, DMBT1 AQP4, ZNF750, SCGB3A1, PLA2G10, SLC44A4, GKN2, ACKR1, LAMP3, AHCYL2 MUC1, NDNF, PTPN13, AQP1, SCNN1G, CLDN2, HAS3, CD1A, PPP1R1B, SFTPC, AL OX15B, CYP4X1, ZNF385B, AGER, MLPH, ACSL5, FCGBP, PRR15L, CRTAC1, GJB1 、AGR3、CLIC6、NRGN、CLDN18、SPINK5、FOS、SERPIND1、TPPP3、APOD、AQP5、S T6GALNAC1, ELAPOR1, PCP4L1, GDF15, CREB3L1, CTSE, HLA.DQB2, HPGD, GFR A3, C3, MS4A15, SCGB1A1, CXCL14, SLPI, SPINK1, CEACAM6, HP, CRLF1, MSLN.
[0089] A classifier was constructed based on the above genes and validated in multiple external cohorts to further verify the stability of the subtypes. Nearest Template Prediction (NTP) classified each sample in the external cohorts as one of the identified CS types. Consistent with this, CS2 showed a better prognosis than CS1 in the META LUAD cohort (p<0.005), with similar results in other external cohorts. Figure 7 E). The consistency of CS with NTP, partitioned surrounding media (PAM), and COMIC algorithms was also evaluated (p<0.005; Figure 2 D、 Figure 2 F-2G). Furthermore, we performed disease ontology (DO) enrichment analysis on 200 taxonomic genes, finding that they were primarily enriched in lung diseases and non-small cell lung cancer, indirectly validating the effectiveness of the classification. Figure 7 H).
[0090] To continue selecting key characteristic genes from differentially expressed genes between subtypes, the following steps were performed:
[0091] A model was constructed based on the differentially expressed genes between the screened subtypes to distinguish between two different molecular subtypes.
[0092] Due to severe batch effects in different queues ( Figure 8 A) We first perform batch effect removal ( Figure 8 B); Next, we performed univariate Cox regression analysis in the training cohort and identified 129 genes significantly associated with the prognosis of lung adenocarcinoma (p<0.05, Figure 8 C). To ensure the robustness of the selection, we first adopted the bootstrap method. All patients underwent 1000 samplings, and 88 genes with p<0.05 and more than 500 samplings were retrieved. Figure 8 D). Then, I Based on the Boruta algorithm, and by comparing the importance of selected and random features, as well as previous literature research, we retained 18 features. Genes considered more relevant to prognosis ( Figure 8 E-8G (as shown in Table 1) constructs a predictive model for molecular typing. To build accurate and stable models, we benchmarked seven machine learning algorithms in the TCGA-LUAD queue using nested cross-validation (CV). Specifically, we tuned the hyperparameters of each algorithm using internal 5x CV and evaluated the performance of the best-tuned model using external 10x CV. By comparing the average of the Harrell chord exponent (C-index) and the comprehensive Brier Score (IBS) across 10 test folds, we determined the model's performance. Ridge Algorithm The best performance was achieved by (the following) Figure 8 H-8I).
[0093] Analysis shows that the model is in Ridge Algorithm The reason for the best performance is that there is a synergistic effect in gene expression, leading to multicollinearity in gene expression characteristics. In order to accurately screen key characteristic genes that distinguish molecular subtypes, based on Figure 9 Based on the results in D, we examine the multicollinearity of these 18 features (Table 1):
[0094] Table 1. Multicollinearity between genes and genotypes
[0095] Gene KRT6A FOSL1 NDNF LYPD3 RNASE1 SLC34A2 UBE2S PBK CDT1 VIF 1.602 1.692 1.767 1.848 1.991 2.004 3.052 3.997 4.373 Gene AURKB CHEK1 CDK1 RRM2 PLK1 EXO1 NEK2 CDKN3 DLGAP5 VIF 5.43 5.614 6.212 6.669 6.717 7.15 8.452 9.156 11.739
[0096] Finally, based on the results in Table 1, we selected KRT6A, FOSL1, NDNF, LYPD3, RNASE1, SLC34A2, UBE2S, PBK, and CDT1 to construct a predictive model for molecular typing, based on the strict requirement of VIF < 5.
[0097] Ultimately, a robust consensus machine learning driven signature / scoring (CMLS) was built.
[0098] The ability of CMLS scores to predict the efficacy of immunotherapy: Using the Immunotumor Biology Research (IOBR) software package, we comprehensively analyzed the tumor microenvironment (TME) of LUAD, observing the level of immune cell infiltration, including CD4 T cells, DC cells, NK cells, and CAF. Patients with low CMLS showed significantly higher levels of these cells than those with high CMLS, indicating an immune-activated state. Patients with high CMLS primarily accumulated bone marrow-driven suppressor cells (MDSCs), exhibiting an immunosuppressive state. Figure 9 A). We further calculated the tracking tumor immune phenotype (TIP) to explore the potential biological mechanisms associated with LUAD, and the results showed that the low CMLS group exhibited significant differences mainly in step 4 (CD4 T cell recruitment) and step 5 (immune cell infiltration), consistent with the results of the above analysis. Figure 9 B). Comparing the relative expression levels of immune checkpoints in the high and low groups, it was found that CD274, CD276, CD70, and IDO1 were significantly upregulated in the high group, while CD27, CD40LG, CD28, and HHLA2 were significantly upregulated in the low group. Figure 9 C). These findings suggest that LUAD characterized by low CMLS levels is more likely to be classified as a “hot tumor,” while LUAD characterized by high CMLS levels is more likely to be classified as a “cold tumor.” The distribution of CMLS in patients with different remission levels also showed that the CMLS score in the remission group (complete remission [CR] / partial remission [PR]) was significantly lower than that in the non-remission group (progressive disease [PD] / stable disease [SD]) (p<0.05; Figure 8 D). Furthermore, using the Tumor Immune Dysfunction and Rejection (TIDE) algorithm to assess patient response to immunotherapy, the low CMLS group showed better responsiveness (P [Fisher's exact test] < -0.2, p < 0.05).
[0099] Characteristic gene mutation status and immune-related analysis: We further analyzed proteins and immune cells associated with the 18 characteristic genes involved in model construction. Figure 8These 18 genes are associated with a variety of immune cells and may play important roles in the immune response. NDNF, RNASE1, and SLC34A2 are positively correlated with immune cells, while the other genes are negatively correlated with most immune cells. Immune scores also showed the same results. Figure 9 D) indicates that NDNF, RNASE1, and SLC34A2 may be key genes promoting immune responses. Tumor mutation burden (TMB) is a powerful molecular marker that can be used to quantitatively assess mutations carried by tumor cells. The waterfall plot shows that the four genes with the highest mutation frequency are EXO1, KRT6A, SLC34A2, and NDNF (…). Figure 8 E).
[0100] By comparing LUAD with normal tissues, we found that 14 of the 18 significantly differentially expressed characteristic genes (AURKB, CDK1, CDKN3, CDT1, CHEK1, DLGAP5, EXO1, KRT6A, LYPD3, NEK2, PBK, PLK1, PRM2, and UBE2S) were significantly upregulated in lung adenocarcinoma tissues, suggesting their potential role in promoting lung adenocarcinoma progression. Normal tissues showed higher expression levels of FOSL1, NDNF, RNASE1, and SLC34A2. Figure 9 A). Figure 10 B shows paired boxplots comparing gene expression between matched normal tissue and LUAD tissue from the same patient. The upward trend in gene expression from normal tissue to LUAD tissue was consistent across all analyzed genes, including AURKB, CDK1, CDKN3, and CDT1, further confirming their significant upregulation in cancerous tissue. The visualization of paired data highlights the robustness of these genes as potential biomarkers distinguishing LUAD from normal lung tissue.
[0101] We further analyzed the single-cell sequencing data from LUAD. We first performed quality control on the single-cell sequencing data, and then used t-SNE to divide the cells into 32 distinct subsets through dimensionality reduction. Figure 10 A). Figure 10 C indicates the expression of cell type marker genes. Figure 11 B shows the t-SNE distribution for each cell type. A total of 12 cell types were identified, including plasma cells, mast cells, fibroblasts, macrophages, and tumor cells. We presented the expression of 18 characteristic genes from these 12 cell types, finding that most genes were primarily expressed in epithelial cells and macrophages. Figure 11 D).
[0102] The role of NDNF in the progression of LUAD and its potential as a therapeutic target:
[0103] Based on the previous research, after removing collinear gene expression, we selected KRT6A, FOSL1, NDNF, LYPD3, RNASE1, SLC34A2, UBE2S, PBK, and CDT1 to construct a predictive model for molecular typing.
[0104] In the analysis of characteristic gene mutation status and immune-related factors, we found that NDNF, RNASE1, and SLC34A2 were positively correlated with immune cells, while other genes were negatively correlated with most immune cells. This prompted us to further investigate why NDNF, RNASE1, and SLC34A2 are key characteristic genes for molecular typing.
[0105] Figure 14 A and 13C further indicated that NDNF expression was significantly reduced in LUAD samples compared to normal lung tissue. To confirm this result, we analyzed NDNF expression in normal lung cells and LUAD cells at the gene and protein levels, revealing a significant downregulation of NDNF in LUAD cells.
[0106] Therefore, we wanted to further explore the functional role of NDNF in LUAD, and we constructed a LUAD cell model (A549) overexpressing NDNF. qPCR results confirmed successful NDNF overexpression. Figure 12 D-11E). Using CCK-8 assays, we observed a significant decrease in cell viability in NDNF-overexpressing cells, indicating that NDNF inhibits cell proliferation (D-11E). Figure 13 F). Furthermore, flow cytometry and Western blot analysis showed that apoptosis was significantly increased in A549 cells overexpressing NDNF. Figure 14 G-13I). Western blot results also showed that in NDNF-overexpressing cells, the level of cleaved caspase-3 (a key marker of apoptosis) was elevated. Figure 14 HI).
[0107] Therefore, we revised our conclusion that subtypes were related to prognosis, and now we believe that subtypes are related to apoptosis, and differences in apoptosis affect different prognoses.
[0108] We established an A549 LUAD cell line model overexpressing NDNF. This model showed that NDNF overexpression inhibited LUAD cell viability and promoted apoptosis.
[0109] In the study described above, we retained genes related to prognosis for constructing a molecular subtype classifier and obtained the following prediction accuracy.
[0110] Table 2 shows the predicted feature selection and corresponding AUC for LUAD-C2.
[0111]
[0112]
[0113] Molecular experiments confirmed the relationship between NDNF and tumor cell apoptosis. We discarded the conventional thinking that it was related to prognosis and directly screened genes based on differentially expressed C1 and C2 genes. During the screening process, we used a machine learning feature selection algorithm, abandoning the prognostic correlation test, and identified six genes more relevant to the molecular subtype classification, such as... Figure 15 As shown in Table 3:
[0114] As can be seen, the predictions in Table 3 can obtain more accurate results using fewer genes, and the C1 and C2 subtypes can predict different levels of apoptosis (see partial ROC figures). Figure 16 ).
[0115] Table 3. Feature selection and corresponding AUC for predicted LUAD-C2 (Part 2)
[0116]
[0117]
[0118] - indicates that information is missing from the test set.
[0119] Figure 3 This is a schematic diagram of a computer device provided in an embodiment of the present invention, such as... Figure 3 As shown, the device may include: one or more processors and one or more memories; wherein the memories store computer-readable code that, when run by the one or more processors, can perform the methods described above.
[0120] The processor in this embodiment can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, operations, and logic block diagrams disclosed in this embodiment. The general-purpose processor can be a microprocessor or any conventional processor, and can be based on an x86 or ARM architecture.
[0121] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0122] For example, the method or apparatus according to embodiments of this disclosure can also be used by means of Figure 4 The architecture of the computing device 3000 shown is used for implementation. For example... Figure 4 As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage devices in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for processing and / or communication of the methods provided in this disclosure, as well as program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 4 The architecture shown is merely exemplary and can be omitted as needed when implementing different devices. Figure 4 One or more components in the computing device shown.
[0123] This invention also includes a computer-readable storage medium, such as... Figure 5The diagram illustrates a storage medium provided in an embodiment of the present invention. The computer storage medium 4020 stores computer-readable instructions 4010. When the computer-readable instructions 4010 are executed by a processor, the method described above according to embodiments of the present disclosure can be performed. The computer-readable storage medium in the embodiments of the present disclosure can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Synchronous Link Dynamic Random Access Memory (SLDRAM), and Direct Memory Bus Random Access Memory (DR RAM). It should be noted that the memory used in the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0124] This disclosure also provides a computer program product or computer program that, when executed by a processor, implements the steps of the above-described method, such as... Figure 2 As shown, the computer program product or computer program includes:
[0125] Acquisition Module 201: Acquires the gene data of the sample to be tested;
[0126] Feature extraction module 202: processes the sequencing data of the sample to be tested to obtain key feature genes; the key feature genes include: NDNF, CCNA2 or RNASE1;
[0127] Prediction module 203: Input the key feature genes into the classifier to obtain the classification results of LUAD-C2 subtype or non-LUAD-C2 subtype; the LUAD-C2 subtype has high apoptosis; the non-LUAD-C2 subtype has low apoptosis.
[0128] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0129] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0130] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0131] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0132] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0133] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0134] The exemplary embodiments of this disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art will understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of this disclosure, and such modifications should fall within the scope of this disclosure.
Claims
1. A method for predicting molecular subtypes of lung adenocarcinoma, characterized in that, The method includes: Obtain the genetic data of the sample to be tested; The gene data of the sample to be tested are processed to obtain key feature genes; the key feature genes include NDNF, and the key feature genes also include CCNA2 or RNASE1; The key characteristic genes are input into a classifier to obtain classification results for LUAD-C2 subtype or non-LUAD-C2 subtype; the LUAD-C2 subtype has high tumor cell apoptosis; the non-LUAD-C2 subtype has low tumor cell apoptosis; the LUAD-C2 subtype has high expression levels of NDNF and RNASE1, and low expression levels of CCNA2; the classifier is obtained through the following methods: Obtain the gene dataset for the training set; Clustering the gene dataset yields the LUAD-C2 subtype and non-LUAD-C2 subtype; Differential gene expression was identified by differential gene expression analysis between the LUAD-C2 subtype and non-LUAD-C2 subtype. Feature selection was performed on the differentially expressed genes to obtain key feature genes; The key feature genes are input into a machine learning model to obtain a predicted classification result, which is then compared with the classification label. The model is optimized based on the comparison result to obtain the classifier.
2. The method for predicting molecular subtypes of lung adenocarcinoma according to claim 1, characterized in that, Based on the classification results of the LUAD-C2 subtype, the result of high cellular immune activity of the test sample is output; based on the classification results of the non-LUAD-C2 subtype, the result of low cellular immune activity of the test sample is output.
3. The method for predicting molecular subtypes of lung adenocarcinoma according to claim 1, characterized in that, When the key characteristic genes include NDNF and CCNA2, the key characteristic genes also include CIP2A; the expression level of CIP2A is low in the LUAD-C2 subtype.
4. The method for predicting molecular subtypes of lung adenocarcinoma according to claim 1, characterized in that, When the key characteristic genes include NDNF and CCNA2, the key characteristic genes also include MYBL2; the expression level of MYBL2 in the LUAD-C2 subtype is low.
5. The method for predicting molecular subtypes of lung adenocarcinoma according to claim 1, characterized in that, When the key characteristic genes include NDNF and CCNA2, the key characteristic genes also include MELK; the expression level of MELK is low in the LUAD-C2 subtype.
6. The method for predicting molecular subtypes of lung adenocarcinoma according to claim 1, characterized in that, When the key characteristic genes include NDNF and CCNA2, the key characteristic genes also include CASP-3; the expression level of CASP-3 in the LUAD-C2 subtype is low.
7. The method for predicting molecular subtypes of lung adenocarcinoma according to claim 1, characterized in that, When the key characteristic genes include NDNF and RNASE1, the key characteristic genes also include SLC34A2; the expression level of SLC34A2 is high in the LUAD-C2 subtype.
8. The method for predicting molecular subtypes of lung adenocarcinoma according to claim 1, characterized in that, When the key characteristic genes include NDNF and RNASE1, the key characteristic genes also include any one or more of the following: KRT6A, FOSL1, LYPD3, UBE2S, PBK, CDT1; the expression levels of KRT6A, FOSL1, LYPD3, UBE2S, PBK, and CDT1 are low in the LUAD-C2 subtype.
9. A computer device, characterized in that, The device includes: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method according to any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1-8.
11. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method described in any one of claims 1-8.
Citation Information
Patent Citations
Lung adenocarcinoma molecular typing and survival risk gene group, diagnosis product and application
CN114341367A
Molecular typing of lung adenocarcinoma based on metabolic genes
CN114480644A