M6a cluster differential gene set for evaluating prognosis survival risk of non-small cell lung cancer, screening method and prognosis survival risk scoring model
By constructing a differential gene set of the m6A cluster for prognostic survival risk in non-small cell lung cancer and its scoring model, the problem of lack of effective prognostic assessment in existing technologies has been solved, realizing a new approach to accurate prognosis assessment and immunotherapy targets for NSCLC patients, and improving the predictive effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANDONG UNIV
- Filing Date
- 2022-06-08
- Publication Date
- 2026-05-05
AI Technical Summary
The lack of effective models for assessing the prognosis of non-small cell lung cancer (NSCLC) with m6A modification in current technologies leads to a low five-year survival rate for NSCLC patients, highlighting the urgent need to improve the prediction of prognosis and the effectiveness of immunotherapy.
A set of differentially expressed genes in the m6A cluster was constructed to assess the prognostic survival risk of non-small cell lung cancer. These genes included S100A10, IGFBP1, SLC52A1, KREMEN2, SCPEP1, COL4A3, GSTM2, STRAP, PCDH7, PAK2, CYP4B1, MTUS1, ESPL1, PRRX2, TTK, COL4A4, CCR7, NPC2, and SKA1. Survival risk scores were calculated based on expression levels, and a prognostic survival risk score model was established using univariate and multivariate Cox regression analysis.
This study provides a novel method for assessing the prognostic survival risk of non-small cell lung cancer (NSCLC). A scoring model constructed using the m6A cluster differential gene set can more accurately predict the prognosis of NSCLC patients, revealing the m6A gene clusters associated with prognosis and providing new treatment ideas, particularly regarding their relationship with TMB, immune cell infiltration, immunotherapy scores, and anticancer drug sensitivity.
Smart Images

Figure CN116230217B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of tumor prognostic assessment technology, and in particular relates to a set of differentially expressed genes of the m6A cluster for assessing the prognostic survival risk of non-small cell lung cancer, as well as screening methods and prognostic survival risk scoring models. Background Technology
[0002] Lung cancer is a leading cause of cancer-related deaths worldwide. Non-small cell lung cancer (NSCLC) accounts for approximately 80–85% of all lung cancer cases. Although molecular immunotherapy and targeted therapy are common in cancer research, the five-year survival rate for NSCLC patients remains low (18%). Therefore, there is an urgent need to improve the prognosis of NSCLC and the prediction of immunotherapy efficacy. Techniques exploring RNA modifications have revealed that some of these modifications exist in certain messenger RNAs (mRNAs). m6A modifications are involved in multiple functions, including tumorigenesis. Among RNA modifications, m6A modifications are the most common, participating in various biological processes such as the formation of cancer stem cells, tumorigenesis, and cell death.
[0003] However, there are currently no reports on models related to the prognostic assessment of m6A modification in non-small cell lung cancer. Summary of the Invention
[0004] In view of this, the purpose of this invention is to provide a set of differentially expressed genes in the m6A cluster for assessing the prognostic survival risk of non-small cell lung cancer, as well as a screening method and a prognostic survival risk scoring model.
[0005] This invention provides a set of differentially expressed genes in the m6A cluster for assessing prognostic survival risk in non-small cell lung cancer, characterized by including the following genes: S100A10, IGFBP1, SLC52A1, KREMEN2, SCPEP1, COL4A3, GSTM2, STRAP, PCDH7, PAK2, CYP4B1, MTUS1, ESPL1, PRRX2, TTK, COL4A4, CCR7, NPC2, and SKA1.
[0006] This invention provides a prognostic survival risk scoring model for non-small cell lung cancer (NSCLC) constructed using the m6A cluster differentially expressed gene set for assessing prognostic survival risk. The model is defined as: Survival Risk Score = (0.1543) × S100A10 expression level + (0.0926) × IGFBP1 expression level + (-0.1585) × SLC52A1 expression level + (0.0512) × KREMEN2 expression level + (-0.0787) × SCPEP1 expression level + (0.1195) × COL4A3 expression level + (-0.0915) × GSTM2 expression level + (0.1032) × STR The expression level of AP is calculated as follows: (0.0693) × PCDH7 expression level + (0.1389) × PAK2 expression level + (0.0437) × CYP4B1 expression level + (-0.0950) × MTUS1 expression level + (0.1204) × ESPL1 expression level + (0.0690) × PRRX2 expression level + (-0.1307) × TTK expression level + (-0.1181) × COL4A4 expression level + (-0.0965) × CCR7 expression level + (-0.0891) × NPC2 expression level + (0.0806) × SKA1 expression level.
[0007] This invention provides a kit for detecting the expression of the m6A cluster differentially expressed gene set, which is used to assess the prognostic survival risk of non-small cell lung cancer, and its application in the preparation of diagnostic or auxiliary diagnostic products for overall survival of non-small cell lung cancer patients.
[0008] This invention also provides a method for screening the m6A cluster differentially expressed gene set for assessing prognostic survival risk in non-small cell lung cancer, comprising the following steps:
[0009] 1) Download clinical data of non-small cell lung cancer patients, as well as RNA-Seq transcriptome data, somatic mutation data, and CNV data of cancerous tissue and adjacent normal tissue from the database;
[0010] 2) Extract m6A-related regulatory factors from the transcriptome data obtained in step (1), perform unsupervised clustering analysis on the extracted m6A-related regulatory factors to obtain the genotypes of the m6A-related regulatory factors; perform intersection genotyping on the genotyped m6A-related regulatory factors again;
[0011] 3) Analyze the differences in the expression of m6A-related regulatory factors in cancerous tissues and adjacent normal tissues of non-small cell lung cancer patients under different subtypes;
[0012] 4) Based on the expression differences of different subtypes of m6A-related regulatory factors in cancerous tissues and adjacent normal tissues of non-small cell lung cancer patients as determined in step 3), and the complete survival information data of patients, univariate and multivariate Cox regression analysis was performed to obtain the m6A cluster differential gene set for assessing the prognostic survival risk of non-small cell lung cancer.
[0013] Preferably, the database in step 1) is the GEO database, TCGA and UCSC-Xena; two NSCLC queues are collected: (GSE50081, GSE68465) and (TCGA-LUSC, TCGA-LUAD).
[0014] Preferably, the unsupervised clustering analysis in step 2) is performed using the ConsensuClusterPlus package.
[0015] Preferably, the expression differences of m6A-related regulatory factors of different subtypes in step 3) are analyzed using the "limma" package.
[0016] Compared with existing technologies, this invention has the following beneficial effects: This invention provides a set of differentially expressed m6A cluster genes for assessing prognostic survival risk in non-small cell lung cancer (NSCLC), along with its screening method and prognostic risk assessment model. This invention explores prognostic-related m6A gene clusters in NSCLC patients and obtains a survival risk scoring model composed of 19 m6A cluster genes using univariate and multivariate Cox multivariate regression analysis. Firstly, m6A cluster analysis was performed using the TCGA project and the GSE50081 and GSE68465 projects, providing new insights for predicting NSCLC prognosis. Furthermore, this invention analyzes the relationship between the risk score of prognostic m6A cluster differentially expressed genes and TMB, immune cell infiltration, immune subtype, immunotherapy score, and anticancer drug sensitivity, providing new treatment strategies for revealing immunotherapy targets in NSCLC patients. Attached Figure Description
[0017] Figure 1 This is a technical roadmap for the entire research of this invention.
[0018] Figure 2 The following data represent the copy number alterations, somatic mutational burden, and differential expression of 23 m6A-related regulators in NSCLC: A. Maftools representation of the somatic mutational burden frequency of the 23 m6A-related regulators in NSCLC; B. Copy number alteration frequency of the 23 m6A-related regulators in NSCLC; C. Copy number pie chart of the 23 m6A-related regulators across the 23 human chromosomes; D. Differential expression levels of the 23 m6A-related regulators in NSCLC and normal tissues.
[0019] Figure 3Survival analysis of m6A regulators in 1622 NSCLC patients; red and blue lines show high and low expression of m6A regulators, respectively, where A to M are FMR1, HNRNPA2B1, HNRNPC, IGFBP1, IGFBP2, IGFBP3, LRPPRC, METTL3, RBM15, WTAP, YTHDC2, YTHDF2, and ZC3H13, respectively.
[0020] Figure 4 The prognostic network for m6A-related modulators was constructed, and unsupervised clustering and survival analysis of 23 m6A modulators were performed. A represents the prognostic network construction, where circles represent m6A, the left half represents the type of m6A, and the right half represents risk factors; red and blue lines represent positive and negative correlations, respectively. B represents the unsupervised clustering of 23 m6A-related modulators in NSCLC. C represents the survival analysis of three m6A clusters in NSCLC. The numbers 1, 2, and 3 in B and C correspond to M6A clusters A, B, and C, respectively.
[0021] Figure 5 Heatmap analysis of the M6A cluster, GSVA, ssGSEA, and differential expression analysis of the m6A cluster in NSCLC; correlation analysis of the expression of the three m6A clusters in NSCLC with clinical characteristics. GSVA analysis of the B-D m6A clusters (A, B), (A, C), and (B, C); E shows the infiltration degree of the three m6A clusters in 23 immune cell types; F is the Venn diagram of differentially expressed genes in m6A clusters A, B, and C.
[0022] Figure 6 Enrichment analysis of differentially expressed genes in the m6A cluster; Boxplot and barplot show the most abundant Go terms and KEGG pathways for differentially expressed genes in the m6A cluster.
[0023] Figure 7 The analyses included univariate and multivariate Cox regression analysis, ROC analysis, and survival analysis. A and B were univariate and multivariate Cox multiple regression analyses of the m6A cluster differentially expressed genes; C and D were univariate and multivariate independent prognostic analyses of the m6A cluster differentially expressed genes; E was ROC analysis of the m6A cluster differentially expressed genes; and F was survival analysis of the m6A cluster differentially expressed genes.
[0024] Figure 8 This study presents the following: Genotyping analysis in NSCLC; A represents gene cluster analysis of differentially expressed genes in the prognostic m6A cluster; B represents survival analysis of the three gene clusters in NSCLC; C represents correlation analysis of m6A cluster and gene cluster expression with clinicopathological features in NSCLC patients; and D represents differential expression analysis of m6A among the three gene clusters.
[0025] Figure 9The correlation analysis of TMB survival rate and risk score with immune infiltration is as follows: A represents the risk score of TMB combined with the predictive model survival analysis; B represents the TMB survival analysis; C represents the correlation analysis of TMB and risk score in NSCLC; D represents the correlation analysis of survival status and risk score in NSCLC patients; E represents the correlation analysis of pathological stage N in NSCLC patients with risk score; F represents the correlation analysis of pathological stage T in NSCLC patients with risk score; and G represents the correlation analysis of immune infiltration and risk score in NSCLC.
[0026] Figure 10 This is a clinical subgroup survival analysis, where AH represents the survival analysis of clinical characteristics in low-risk and high-risk cohorts of NSCLC patients; the red and blue lines represent high-risk and low-risk groups.
[0027] Figure 11 Association analysis between IPS and non-small cell lung cancer risk scores; violin plots showing the correlation between IPS and risk scores in predictive models for NSCLC patients, with AD representing PD-1 and CTLA-4 negative; PD-1 negative and CTLA-4 positive; PD-1 positive and CTLA-4 negative; and correlation analysis of the differences between high- and low-risk groups for IPS with both PD-1 and CTLA-4 positive, with red and blue violin plots representing high- and low-risk groups, respectively.
[0028] Figure 12 This study analyzed the expression, correlation, and immune subtypes of differentially expressed genes in the prognostic m6A cluster. AB shows box plots of 19 differentially expressed genes in the prognostic m6A cluster from LUAD and LUSC tissues from UCSC-Xena; CD shows the correlation analysis of the expression of the 19 differentially expressed genes in the prognostic m6A cluster; and EF shows the correlation analysis of the expression of the 19 differentially expressed genes in the m6A cluster with immune subtypes.
[0029] Figure 13 The correlation analysis of 19 prognostic m6A cluster differentially expressed genes with clinical features of NSCLC was performed; the correlation analysis of m6A cluster differentially expressed genes with clinicopathological features in LUAD patients was performed in AD patients; and the correlation analysis of m6A cluster differentially expressed genes with clinicopathological features in LUSC patients was performed in EH patients.
[0030] Figure 14 Correlation analysis of m6A cluster differential gene expression with stemness score and immune microenvironment in non-small cell lung cancer (NSCLC); A is the correlation analysis of m6A cluster differential gene expression with stemness score and immune microenvironment in LUAD; B is the correlation analysis of m6A cluster differential gene expression with stemness score and immune microenvironment in LUSC.
[0031] Figure 15Survival analysis of differentially expressed genes in the m6A cluster in NSCLC; AP represents the survival analysis of differentially expressed genes in the m6A cluster in LUAD; QR represents the survival analysis of differentially expressed genes in the m6A cluster in LUSC.
[0032] Figure 16 The differential expression of 19 prognostic m6A cluster differentially expressed genes in adjacent non-LUAD and LUAD tissues is shown; red and blue box plots represent LUAD and non-LUAD tissues in the TCGA database; * indicates P<0.05, ** indicates P<0.01, *** indicates P<0.001, and **** indicates P<0.001.
[0033] Figure 17 Differential analysis of 19 prognostic m6A cluster differentially expressed genes in LUSC and adjacent non-LUSC tissues; red and blue box plots represent LUSC and non-LUSC tissues from the TCGA database. * indicates P<0.05, ** indicates P<0.01, *** indicates P<0.001, **** indicates P<0.001.
[0034] Figure 18 This study analyzed the association between differentially expressed genes in the m6A cluster in CellMiner and the sensitivity of NSCLC to anticancer drugs; the top 20 statistically significant associations between differentially expressed genes in the m6A cluster and anticancer drugs were analyzed.
[0035] Figure 19 Differential gene expression of the m6A cluster in the HPA database.
[0036] Figure 20 To verify differential gene expression of the 19m6A cluster in NSCLC and Beas-2B cell lines; * indicates P<0.05, ** indicates P<0.01, *** indicates P<0.001, **** indicates P<0.001. Detailed Implementation
[0037] The technical solutions provided by the present invention will be described in detail below with reference to the embodiments, but they should not be construed as limiting the scope of protection of the present invention.
[0038] Example 1
[0039] 1) Download clinical data of non-small cell lung cancer patients, as well as RNA-Seq transcriptome data, somatic mutation data, and CNV data of cancerous tissue and adjacent normal tissue from the database;
[0040] 2) Extract m6A-related regulatory factors from the transcriptome data obtained in step (1), perform unsupervised clustering analysis on the extracted m6A-related regulatory factors to obtain the genotypes of the m6A-related regulatory factors; perform intersection genotyping on the genotyped m6A-related regulatory factors again;
[0041] 3) Analyze the differences in the expression of m6A-related regulatory factors in cancerous tissues and adjacent normal tissues of non-small cell lung cancer patients under different subtypes;
[0042] 4) Based on the expression differences of different subtypes of m6A-related regulatory factors in cancerous tissues and adjacent normal tissues of non-small cell lung cancer patients as determined in step 3), and the complete survival information data of patients, univariate and multivariate Cox regression analysis was performed to obtain the m6A cluster differential gene set for assessing the prognostic survival risk of non-small cell lung cancer.
[0043] 5) Based on multivariate regression analysis, 19 m6A-related differentially expressed genes were obtained. Subsequently, the risk scores of these 19 m6A-related differentially expressed genes were correlated with tumor mutation burden, immune checkpoint inhibitors, gene expression, immunophenotyping, stem cell scores, and resistance to antitumor drugs. Finally, the expression levels of the 19 m6A-related differentially expressed genes were verified at the mRNA and protein levels using qRT-PCR and the immunohistochemistry database of the HPA database.
[0044] The specific steps and results are as follows:
[0045] Data collection
[0046] RNA-Seq and clinicopathological data were extracted from the GEO database, TCGA, and UCSC-Xena. Patients with complete survival information were extracted for further analysis. Two NSCLC cohorts (GSE50081, GSE68465) and (TCGA-LUSC, TCGA-LUAD) datasets were collected for further analysis; somatic mutation data for TCGA-LUSC and TCGA-LUAD from the TCGA database and copy number (gene-level) data for TCGA-LUSC and TCGA-LUAD from the UCSC-Xena database were downloaded. For the TCGA dataset, RNA sequencing gene expression data (FPKM values) were converted to TPM (ten-kilobase million transcripts) values. R version 4.0.3 was used for data analysis.
[0047] Copy numbers of 23 RNA modification regulators in NSCLC were downloaded. Somatic CNAs of these m6A regulators were evaluated. m6Amaftools showed that ZC3H13, FMR1, RBM15, YTHDC2, LRPPRC, FTO, YTHDC1, YTHDF1, IGFBP3, YTHDF3, HRNPC, RNPA2B1, METTL16, RBM15B, YTHDF1, RBMX, METTL14, and WTAP were altered in 215 NSCLC patients from 1052 samples. Figure 2(A in the text). CNV frequencies show that most m6A-related regulatory factors are elevated ( Figure 2 (B in the diagram). The copy number pie chart shows the CNV frequencies of 23 m6A-related regulatory factors across the 23 human chromosomes. As shown in the pie chart, FMR1, LRPPRC, YTHDC1, HNRNPA2B1, WTAP, IGFBP1, YTHDF3, IGFBP3, VIRMA, FTO, ZC3H13, METTL3, HNRNPC, METTL16, and ALKBH5 are higher in samples with increased copy numbers than in samples with decreased copy numbers. Conversely, in samples with decreased copy numbers, RBM15, YTHDF1, YTHDF2, IGFBP2, RBM15B, METTL14, YTHDC2, and RBMX are higher than in samples with increased copy numbers. Figure 2 (C) Box plots showed differential expression of 23 m6A-related modulators in TCGA-LUAD and LUSC patients. METTL3, LRPPRC, RBM15, VIRMA, YTHDF1, HNRNPC, YTHDF2, RBMX, HNRNPA2B1, IGFBP2, and IGFBP3 were expressed higher in NSCLC tissues than in normal tissues, while METTL14, METTL16, and ZC3H13 were expressed significantly lower in NSCLC tissues than in normal tissues. In normal tissues ( Figure 2 Survival analysis showed that high expression of FMR1, HNRNPC, IGFBP1, IGFBP3, LRPPRC, RBM15, WTAP, and ZC3H13 was associated with a worse prognosis in NSCLC patients than low expression. High expression of IGFBP2, METTL3, YTHDC2, and YTHDF2 was associated with a better prognosis in NSCLC patients than low expression of these markers. Figure 3 )
[0048] Construction and unsupervised clustering analysis of the prognostic network of 23 m6A-related regulatory factors
[0049] Twenty-three m6A-related regulatory factors (selected from published articles) were extracted from two GEO datasets (GSE50081, GSE68465) and (TCGA-LUAD, TCGA-LUSC) to identify different m6A regulation patterns mediating m6A RNA modification. These 23 m6A-related regulatory factors include 13 readers (YTHDC1, YTHDF3, HNRNPA2B1, YTHDC2, HNRNPC, YTHDF1, IGFBP1, YTHDF2, IGFBP2, FMR1, LRPPRC, IGFBP3, RBMX), 8 writers (METTL3, RBM15, VIRMA, WTAP, ZC3H13, METTL14, RBM15B, METTL16), and 2 erasers (FTO, ALKBH5). Unsupervised clustering analysis was applied to explore different m6A RNA methylation modification patterns in relation to the expression of these 23 m6A-related regulatory factors. Consensus clustering methods are used to explore stability and the number of clusters. The ConsensuClusterPlus package is used to execute the clustering method steps and 1000 repetitions to evaluate classification stability.
[0050] The m6A regulatory factor prognostic network showed that IGFBP1, HNRNPA2B1, LRPPRC, HNRNPC, YTHDF3, YTHDF1, RBM15B, RBM15, ZC3H13, WTAP, and IGFBP3 were risk factors for predicting clinical outcomes in NSCLC. FMR1, YTHDF2, YTHDC2, YTHDC1, FTO, METTL3, and IGFBP2 were favorable factors for predicting prognosis in NSCLC patients. The m6A regulatory factor prognostic network indicated that the expression levels of FMR1, YTHDC1, and IGFBP2 were negatively correlated with IGFBP1 expression. The expression levels of YTHDF2, YTHDC2, and METTL3 were negatively correlated with IGFBP3 expression. Figure 4 A)
[0051] The NSCLC samples (TCGA-LUAD, TCGA-LUSC, GSE50081, GSE68465) were divided into three clusters: cluster A (1), cluster B (2), and cluster C (3). Figure 4 Then, survival analysis was performed in the three clusters. Survival analysis showed that the prognosis of NSCLC patients in cluster B of m6A was significantly better than that of NSCLC patients in clusters C and A of m6A (P = 0.018). Figure 4 (C in the middle).
[0052] Genomic Variation Analysis (GSVA)
[0053] GSVA was performed using the "GSVA" R package to identify differences among m6A-related regulators based on biological processes. GSVA is an unsupervised and nonparametric algorithm commonly used to assess changes in the activity and pathways of biological processes in samples of expression data. The "c2.cp.kegg.v7.2" package was used to extract gene sets from the MSigDB database. An adjusted p-value <0.05 was considered statistically significant. Genomic variation analysis showed that m6A clusters A, B, and C were primarily enriched in immune-related pathways and functions.
[0054] To further investigate the expression of m6A-related regulatory factors in clusters A, B, and C, a heatmap analysis of m6A clusters was performed. The heatmap analysis showed that, compared to clusters B and C, the expression of m6A-related regulatory factors was downregulated in cluster A. Compared to clusters A and C, the expression of m6A-related regulatory factors was upregulated in cluster B. Furthermore, compared to clusters A and B, IGFBP2 and IGFBP3 were upregulated in cluster C. Figure 5 (A to D in the original text).
[0055] TME cell infiltration assessment
[0056] The relative abundance of each cell infiltrating the tumor microenvironment in NSCLC was calculated using the single-sample gene set enrichment analysis (ssGSEA) algorithm. The enrichment score was used to indicate the relative abundance of TME infiltration in the sample.
[0057] Immune cell infiltration in 1622 NSCLC patients was assessed using ssGSEA based on three m6A clusters. M6A clusters A, B, and C were significantly different in 21 immune cell types (P<0.05). Figure 5 (E in the middle).
[0058] Identifying differential expression of m6A genotype gene in NSCLC
[0059] The "limma" package was used to analyze differential expression of m6A genotype genes. First, the m6A gene expression data and m6A cluster data were merged, and then the crossover point was taken. Differential expression analysis was performed based on an adjusted p-value < 0.001 to obtain the differentially expressed m6A genotype genes in m6A clusters C and B. Then, the "VennDiagram" R package was applied to find the intersection of differentially expressed m6A genotype genes in m6A clusters B, C, and A.
[0060] Regarding m6A-related gene expression, NSCLC patients were divided into three m6A clusters. Differential expression analyses were then performed within m6A clusters A and B; B and C; and A and C; p < 0.001 was considered statistically significant. Next, the intersection of the differentially expressed genes from the three m6A clusters was taken to obtain the final m6A cluster differentially expressed genes. Finally, 982 m6A cluster differentially expressed genes were obtained. Figure 5 (F in the text)
[0061] Enrichment analysis of genes among m6A subtypes
[0062] The R packages “ClusterProfiler”, “enrichplot”, “org.Hs.eg.db”, and “ggplot2” were used for enrichment analysis using the Kyoto Encyclopedia of Genes and Genomes (KEGG) and Gene Ontology (GO). A p-value < 0.05 was considered statistically significant. GO analysis showed that the m6A cluster differentially expressed genes were most abundant in glandular and epidermal development. KEGG analysis showed that the m6A cluster differentially expressed genes were most enriched in phagosomes. Figure 6 ).
[0063] Based on univariate and multivariate Cox regression analysis of m6A genotypes
[0064] The “Survival” package is used for univariate and multivariate Cox regression analysis. Univariate and multivariate Cox regression analysis is performed by combining m6A subtype gene expression data with complete survival information data of NSCLC patients in GSE50081, GSE68465, TCGA-LUSC, and TCGA-LUAD.
[0065] The packages "survival", "timeROC", and "survminer" were used to perform ROC (receiver operating characteristic) and survival analyses. First, the survival differences between the low-risk and high-risk cohorts were compared, yielding significant p-values (P < 0.001), and survival curves for the low-risk and high-risk cohorts were obtained. Next, the "timeROC" package was used to obtain the area under the curve (AUC) of the ROC curve at 1, 3, and 5 years. A p-value < 0.05 was considered statistically significant.
[0066] Based on p-values < 0.001, 57 differentially expressed genes from the m6A cluster were selected for multivariate Cox regression analysis. Figure 7 (A in the middle). Finally, 19 differentially expressed genes from the m6A cluster were selected to establish a prediction model ( Figure 7(B in the original text). Survival risk score = (0.1543) * expression of S100A10 + (0.0926) * expression of IGFBP1 + (-0.1585) * expression of SLC52A1 + (0.0512) * expression of KREMEN2 + (-0.0787) * expression of SCPEP1 + (0.1195) * expression of COL4A3 + (-0.0915) * expression of GSTM2 + (0.1032) * expression of STRAP + (0.0693) * expression of PCDH7 +(0.1389)*PAK2 expression +(0.0437)*CYP4B1 expression +(-0.0950)*MTUS1 expression +(0.1204)*ESPL1 expression +(0.0690)*PRRX2 expression +(-0.1307)*TTK expression +(-0.1181)*COL4A4 expression +(-0.0965)*CCR7 expression +(-0.0891)*NPC2 expression +(0.0806)*SKA1 expression.
[0067] Clinical information and risk data were downloaded from GSE50081, GSE68465, TCGA-LUAD, and TCGA-LUSC. Independent prognostic analyses were performed using Perl and the survivalR package. Univariate independent prognostic analysis showed that age, sex, pathological T and N stages, and risk scores were independent prognostic factors for NSCLC patients (P<0.001). Figure 7 (C) Multivariate independent prognostic analysis showed that age, pathological T and N stage, and risk score characteristics of NSCLC patients could be considered as high-risk factors for predicting the prognosis of NSCLC patients (P<0.001). Figure 7 (D in the original text). ROC curves were performed to test the impact of the 19-m6A cluster differentially expressed prognostic model on overall survival in NSCLC. Areas under the ROC curve (AUC) at 1, 3, and 5 years showed that the predictive features of the 19-m6A cluster differentially expressed genes better predicted the prognosis of NSCLC patients (AUC = 0.69, 0.673, and 0.653, respectively). Figure 7 In terms of median risk score, NSCLC patients were divided into two cohorts: high-risk (n=811) and low-risk (n=811). Survival analysis showed that the clinical outcomes of the low-risk cohort were significantly better than those of the high-risk cohort (P<0.001). Figure 7 (F in the text)
[0068] Gene clustering analysis
[0069] Gene clustering analysis was performed using the expression of prognostic-related genes. The "limma" and "ConsensusClusterPlus" packages were used for gene clustering analysis. Survival analysis was then performed using the "survminer" and "survival" packages. Finally, the "pheatmap" package was used to analyze the differences between gene clusters in NSCLC by combining m6A cluster data, gene cluster data, prognostic-related gene data, and clinical characteristic data from the GEO and TCGA datasets.
[0070] Genotyping m6A differential analysis
[0071] To identify differences in m6A-related genes across different gene clusters, we performed a genotyping-based m6A differential analysis. The packages “limma”, “reshape2”, and “ggpubr” were used for this analysis. A p-value <0.05 was considered statistically significant.
[0072] Results: Genotyping analysis was performed using differential gene expression in the prognostic m6A cluster. The "ConsensusClusterPlus" package was used for gene clustering analysis. Ultimately, three gene clusters were obtained: cluster A (n=179), cluster B (n=273), and cluster C (n=179). Figure 8 (A) Genotyping survival analysis showed that the prognosis of gene cluster C was significantly better than that of gene clusters B and A (P = 0.007). Figure 8 (B). Genotyping heatmap analysis showed that differentially expressed genes in the prognostic m6A cluster were significantly overexpressed in cluster A and significantly underexpressed in cluster C. Figure 8 Box plots showed significant differences in the expression of METTL3, WTAP, ZC3H13, YTHDC1, IGFBP2, YTHDC2, YTHDF2, and IGFBP3 within gene cluster AC (P<0.05). Figure 8 (D in the middle).
[0073] Correlation analysis between risk score and TMB (tumor mutation burden)
[0074] Based on the median risk score of NSCLC patients from the TCGA and GEO datasets, NSCLC patients were divided into low-risk and high-risk cohorts, and correlation analysis was performed on the risk score and TMB. The TMB of NSCLC patients was obtained using Perl. Then, the "ggpubr", "reshape2" packages, and Spearman association analysis were used to obtain the association between risk score and TMB.
[0075] TMB and TMB Joint Risk Score Survival Analysis
[0076] TMB data and risk score cohorts of NSCLC patients were integrated and cross-referenced to obtain mean cutoff values for classifying NSCLC patients into low-TMB and high-TMB cohorts, followed by TMB survival analysis. Combined TMB-risk score survival analysis was then performed, combining TMB data with NSCLC patient risk score cohorts. Based on the mean cutoff values, NSCLC patients were divided into low-TMB and high-risk, high-TMB and high-risk, low-TMB and low-risk, and high-TMB and low-risk groups.
[0077] result:
[0078] Survival analysis using TMB combined with risk scores showed that high TMB and low risk scores had better clinical outcomes than low risk scores, low TMB, high TMB and high risk scores, and high risk scores and low TMB. Figure 9 (A) TMB survival analysis showed that patients with low TMB had significantly lower clinical outcomes than patients with high TMB (P = 0.039). Figure 9 (B in the middle).
[0079] Differences between TMB and Risk Score in NSCLC
[0080] To explore the differences in median total molecular weight (TMB) among NSCLC patients, a differential analysis of median TMB between high-TMB and low-TMB cohorts was performed. To further investigate the differences in risk scores between surviving and deceased NSCLC patients, a differential analysis of risk scores between surviving and deceased NSCLC patients was also conducted.
[0081] Immune correlation analysis
[0082] The “corrplot” package is used to perform Pearson correlation analysis. First, correlation spectra are obtained using riskscore data and ssGSEA immune cell data. Then, the “corrplot” package is used to obtain the association between riskscore and 23 immune cells.
[0083] result:
[0084] Correlation analysis between prognostic m6A cluster differential gene risk scores and total mesotherapy (TMB) showed that patients with high-risk scores had higher TMB (P = 4.4e-8). Figure 9 (C) The risk score of deceased NSCLC patients was higher than that of living patients (P<2.22e-16). Figure 9 (D) Patients with NSCLC pathological stage N1-3 had a higher risk score than patients with pathological stage N0 (P = 0.0023). Figure 9 E in the middle). The risk score of T1-2 NSCLC patients was lower than that of T3-4 patients (P = 0.00018) (E). Figure 9The association analysis between risk score and immune infiltration showed that immune cells were significantly positively correlated with risk score (F in the original text). Figure 9 (G in the middle).
[0085] Clinical subgroup survival analysis
[0086] Clinical subgroup survival analysis was performed using the "survival" and "survminer" software packages. First, clinical information and riskcore group data from NSCLC patients were integrated. Then, the clinical and riskcore group data were cross-referenced to extract clinical characteristics of NSCLC patients. Next, the categories within each clinical characteristic were iterated over, and clinical subgroup survival analysis was performed.
[0087] Results: Clinical subgroup survival analysis showed that, compared with the low-risk cohort, patients in the high-risk cohort who were older than or equal to 65 years, female, male, N0, N1-3, or T1-2 had a worse prognosis (P<0.001). Figure 10 In the AF (affected follicle syndrome) assessment, patients with T3-4 scores in the low-risk cohort had better clinical outcomes than those in the high-risk cohort (P = 0.05). Figure 10 (GH in the middle).
[0088] Immunotherapy Analysis
[0089] Immunotherapy scores for LUSC and LUAD were downloaded from the Cancer Immunogastronomical Atlas (TCIA) database. The immunotherapy score comprised four components: negative for both CTLA4 and PD1, positive for CTLA4 and negative for PD1, negative for both CTLA4 and PD1, and positive for both CTLA4 and PD1. NSCLC patients were then divided into high-risk and low-risk cohorts based on their median risk score. Immunotherapy analysis was performed between the low-risk and high-risk populations using the "ggpubr" package.
[0090] Results: In patients with CTLA4 and PD1 negative NSCLC, there was a significant difference in IPS between the low-risk and high-risk score cohorts (P = 1.6e-09). Figure 11 (A) In the analysis of IPS, NSCLC patients with CTLA4-negative and PD1-positive IPS in the low-risk cohort had better immunotherapy response than those in the high-risk cohort (P = 6.1e-06). Figure 11 (B) In CTLA4-positive and PD1-negative NSCLC patients, iPS in the low-risk cohort had better efficacy with immunotherapy than in the high-risk cohort (P = 6.6e-11). Figure 11 (C) In CTLA4 and PD1 positive NSCLC patients, iPS in the low-risk cohort had better efficacy than in the high-risk cohort (P = 9.6e-07). Figure 11 (D in the middle).
[0091] Differences in risk scores among different clinical characteristics
[0092] To explore the differences in clinical characteristics among risk scores, a differential analysis was performed. Clinical characteristics of LUAD and LUSC patients were downloaded from the TCIA database and then combined with risk score groups of NSCLC patients for differential analysis using the "ggplot2" and "ggpubr" packages.
[0093] Analysis of gene expression, correlation, clinical characteristics, immune subtypes, NSCLC microenvironment, survival rate, and drug sensitivity among prognostic m6A subtypes
[0094] Transcriptional expression data, clinical information, immunophenotypic scores, and stemness scores (RNA methylation (RNAs) and DNA expression (DNAss)) for LUAD and LUSC were downloaded from the UCSCXena website. Differential analysis was performed using the "ggpubr" R package. Next, Spearman association analysis was used to perform gene association analysis between m6A subtypes. Kruskal's test was used to analyze differences in pathological N, T, and stages between LUAD and LUSC, while Wilcox's test was used to analyze differences in pathological M stages. Estimated score data for LUAD and LUSC were extracted using the "estimate" R package. Spearman association analysis was used to analyze the association between gene expression in m6A subtypes and matrix scores, NSCLC immunophenotypic scores, estimated RNAss scores, and DNAss scores. Survival analysis was performed using the "survminer" and "survival" software packages. The CellMiner database (https: / / discover.nci.nih.gov / cellminer / home.do) was used to extract the same drug sensitivity and gene expression samples. Then, the anticancer drug sensitivity data were validated by clinical laboratories and filtered according to FDA standards. Finally, Pearson correlation analysis was performed to obtain the correlation between genes in the m6A subtype and the drug sensitivity of LUAD and LUSC.
[0095] result:
[0096] The UCSC-Xena database was used to download expression, immune subtypes, clinical characteristics, LUSC, and LUAD dryness scores for TCGA-LUAD and TCGA-LUSC. Box plot analysis showed that S100A10, SCPEP1, STRAP, PAK2, and NPC2 were upregulated in LUAD and LUSC patients. Figure 12 Correlation analysis showed that the differentially expressed genes in the prognostic m6A cluster were positively correlated with each gene (AB in the data). Figure 12(CD in the text). Immunotype analysis of LUAD and LUSC showed that the 19 m6A cluster differentially expressed genes were different in immunotypes C1, C6, C2, C4, and C3 (P<0.05). Figure 12 EF in LUAD). ESPL1 expression showed a statistically significant difference in pathological M1 and M0 stages in LUAD patients (P<0.05). Figure 13 (A) SKA1 and CCR7 expression showed statistically significant differences in pathological M1 and M0 stages in LUSC patients (P<0.05). Figure 13 The expression of S100A10, IGFBP1, GSTM2, STRAP, CYP4B1, PRRX2, TTK, and SKA1 differed in the pathological N1-3 and N0 stages of LUAD patients (P<0.05). Figure 13 (B in the text). The expression of CYP4B1, ESPL1, SKA1, and NPC2 differed between pathological stages N0 and N1-3 in LUSC patients (P<0.05). Figure 13 The expression of SCPEP1, COL4A3, CYPB1, MTUS1, COL4A4, and CCR7 differed in the pathological T1-4 stages of LUSC patients (P<0.05). Figure 13 The expression of SCPEP1, GSTM2, MTUS1, and SKA1 differed in pathological stages I-IV of LUAD patients (P<0.05). Figure 13 The expression of KREMEN2, CYP4B1, ESPL1, SKA1, and TTK showed statistically significant differences in pathological stages I-IV of LUSC patients (P<0.05). Figure 13 The expression of KREMEN2, STRAP, PAK2, ESPL1, TTK, and SKA1 was positively correlated with RNAss and DNAss in LUAD patients (P<0.05). Figure 14 The expression of S100A10, SCPEP1, COL4A3, PCDH7, PAK2, CYP4B1, PRRX2, NPC2, and CCR7 was positively correlated with the immune matrix score of LUAD patients (P<0.05). Figure 14 (A) The expression of SLC52A1, GSTM2, STRAP, PAK2, ESPL1, TTK, and SKA1 was positively correlated with RNAss in LUSC patients (P<0.05). Figure 14 (B in the text) The expression of S100A1, IGFBP1, COL4A3, COL4A4, NPC2, and CCR7 was negatively correlated with DNAss in LUSC patients (P<0.05). Figure 14(B) The expression of COL4A3, PCDH7, CYP4B1, COL4A1, NPC2, and CCR7 was positively correlated with the immune matrix score of LUSC patients (P<0.05). Figure 14 (B) Expression of SLC52A1, GSTM2, STRAP, PAK2, ESPL1, TTK, and SKA1 was negatively correlated with immune and matrix scores in LUSC patients (P<0.05). Figure 14 (B) Survival analysis showed that high expression of ESPL1, IGFBP1, KREMEN2, PAK2, PCDH7, PRRX2, S100A10, STRAP, and TTK in LUAD patients indicated poor prognosis (P<0.05). Figure 15 The C, E, F, I, J, K, L, O, and P components were mentioned. Higher CYP4B1 expression in LUSC patients indicates a poor prognosis (P = 0.020). Figure 15 (Q). In LUAD patients, higher expression of CCR7, CYP4B1, GSTM2, MTUS1, NPC2, SCPEP1, and SLC52A1 was associated with a better prognosis (P<0.05). Figure 15 The A, B, D, G, H, M, and N in the spectrum. Higher expression of SLC52A1 in LUSC patients was associated with better clinical outcomes (P = 0.004). Figure 15 (R). Differential analysis of 19 m6A cluster genes showed statistically significant differences in LUAD, LUSC, and normal tissues (P<0.05). Figure 16 , Figure 17 Association analysis of anticancer drug sensitivity with m6A cluster gene expression showed that CCR7 expression was positively correlated with drug sensitivity to nerabine, fluphenazine, dexamethasone decavalent, PX-136, and chelidonine (P<0.05). S100A10 expression was positively correlated with kahalideff and irrofulven (P<0.05). SLC52A1 expression was positively correlated with fulvestrant and negatively correlated with vincristine and acetylcholine (P<0.05). KREMEN2 expression was positively correlated with fulvestrant (P<0.001). Figure 18 ).
[0097] Expression validation of 19 differentially expressed genes in the m6A cluster
[0098] Cell culture
[0099] Human NSCLC cells (H1975, A549, H1299) and Beas-2B (human bronchial epithelial-like cells) were purchased from ProcellLife Sciences & Technology Co., Ltd. and Saibai Kang Biotechnology Co., Ltd., respectively. The short tandem repeat (STR) algorithm was used to validate NSCLC and Beas-2B cells. Beas-2B and NSCLC cells were cultured in DMEM medium (Gibco, Invitrogen, Carlsbad CA, USA) and containing 10% fetal bovine serum (Lonsera, Shanghai Shuangru Biotechnology Co., Ltd.). NSCLC and Beas-2B cells were cultured at 37°C in a sterile humidified incubator containing 5% carbon dioxide.
[0100] Quantitative real-time polymerase chain reaction (qRT-PCR)
[0101] Total RNA was isolated from Beas-2B and NSCLC cells (A549, H1975, H1299) using Fastgen reagents (Shanghai Feijie Biotechnology Co., Ltd.). Reverse transcription was then performed using the EvoM-MLVRT Kit (Accurate Biotechnology (Hunan) Co., Ltd.), and qRT-PCR was performed using the SYBR Green Premix ProTaq HS qPCR Kit (Accurate Biology). GAPDH was used as a reference gene for 19 differentially expressed genes in the prognostic m6A cluster. Primer sequences synthesized from BioSune Co., Ltd. (Shanghai, China) are detailed below. Relative expression values were calculated using the 2-ΔΔCT algorithm.
[0102] Primer sequence (5'->3')
[0103]
[0104]
[0105] Statistical analysis
[0106] All experiments were repeated three times. Data were analyzed using GraphPad Prism 7.0 and R4.0.3 software. Significant differences between Beas-2B and NSCLC cells were determined using one-way ANOVA. A p-value < 0.05 was considered statistically significant. Results: The HPA database was used to validate differential gene expression in the prognostic m6A cluster in NSCLC tissues. Statistically significant differences were found in the expression levels of S100A10, SCPEP1, KREMEN2, GSTM2, STRAP, PCDH7, PAK2, ESPL1, TTK, NPC2, and SKA1 between LUAD and LUSCLC tissues. Figure 19 ).
[0107] qRT-PCR was used to validate differentially expressed genes in the prognostic m6A cluster in normal human bronchial epithelial cells Beas-2B and human NSCLC cells.
[0108] The expression levels of CCR7, COL4A3, COL4A4, CYP4B1, ESPL1, IGFBP1, MTUS1, PRRX2, SCPEP1, SKA1, SLC52A1, and TTK in A549 and Beas-2B cells differed significantly (P<0.05). The expression levels of ESPL1, GSTM2, KREMEN2, NPC2, PRRX2, SCPEP1, SKA1, SLC52A1, STRAP, and TTK in Beas-2B cells were higher than those in the H1299 cell line (P<0.05). The expression levels of ESPL1, GSTM2, MTUS1, NPC2, PCDH7, S100A10, and SCPEP1 differed significantly between the H1975 and Beas-2B cell lines (P<0.05). Figure 20 ).
[0109] As demonstrated by the above embodiments, this invention explored prognostic m6A-related gene clusters in NSCLC. Univariate and multivariate Cox multivariate regression analyses revealed that m6A cluster genes (S100A10, IGFBP1, SLC52A1, GSTM2, PCDH7, MTUS1, PRRX2, TTK, COL4A4, CCR7) can serve as independent prognostic factors for predicting clinical outcomes in NSCLC. Firstly, m6A cluster analysis was performed using the TCGA project and the GSE50081 and GSE68465 projects, providing new insights into predicting NSCLC prognosis. Furthermore, the relationship between risk scores of differentially expressed genes in the prognostic m6A cluster and TMB, immune cell infiltration, immune subtype, immunotherapy score, and anticancer drug sensitivity was analyzed, providing new insights into revealing immunotherapy targets for NSCLC patients.
[0110] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention. sequence list <110> Shandong University <120> differentially expressed gene set of m6A cluster for assessing prognostic survival risk in non-small cell lung cancer, screening methods, and prognostic survival risk scoring model. <160> 40 <170> SIPOSequenceListing 1.0 <210> 1 <211> 20 <212> DNA <213> Artificial Sequence <400> 1 tcgctgggga taaaggctac 20 <210> 2 <211> 20 <212> DNA <213> Artificial Sequence <400> 2 aagaagctct ggaagcccac 20 <210> 3 <211> twenty one <212> DNA <213> Artificial Sequence <400> 3 gcacggagat aactgaggag g 21 <210> 4 <211> 20 <212> DNA <213> Artificial Sequence <400> 4 ctgatggcgt cccaaaggat 20 <210> 5 <211> 20 <212> DNA <213> Artificial Sequence <400> 5 ttgctgttgc catcactacc 20 <210> 6 <211> twenty two <212> DNA <213> Artificial Sequence <400> 6 caaagcctct tcttcctcct tc 22 <210> 7 <211> 20 <212> DNA <213> Artificial Sequence <400> 7 gtgggcactg ggttcagtta 20 <210> 8 <211> 20 <212> DNA <213> Artificial Sequence <400> 8 agctctagac caatgccagc 20 <210> 9 <211> 20 <212> DNA <213> Artificial Sequence <400> 9 cctgcctgga agaattccga 20 <210> 10 <211> 20 <212> DNA <213> Artificial Sequence <400> 10 tccccagctt tcacagttga 20 <210> 11 <211> 20 <212> DNA <213> Artificial Sequence <400> 11 tgcggggaat cagaaaagga 20 <210> 12 <211> twenty one <212> DNA <213> Artificial Sequence <400> 12 gctgcatacg gctgtccata a 21 <210> 13 <211> 20 <212> DNA <213> Artificial Sequence <400> 13 tgctacgcca gggagataca 20 <210> 14 <211> 20 <212> DNA <213> Artificial Sequence <400> 14 cctgagacag catcccacac 20 <210> 15 <211> 20 <212> DNA <213> Artificial Sequence <400> 15 cctactcgct ggactcctct 20 <210> 16 <211> 20 <212> DNA <213> Artificial Sequence <400> 16 gtccagcacg gtattgacca 20 <210> 17 <211> 20 <212> DNA <213> Artificial Sequence <400> 17 ttgatggtgc tgccaagtct 20 <210> 18 <211> 20 <212> DNA <213> Artificial Sequence <400> 18 ccagaagccc cttgtccaat 20 <210> 19 <211> 19 <212> DNA <213> Artificial Sequence <400> 19 tgagcaccag catcgttgt 19 <210> 20 <211> twenty one <212> DNA <213> Artificial Sequence <400> 20 cccagatcat cccactggaa g 21 <210> twenty one <211> twenty one <212> DNA <213> Artificial Sequence <400> twenty one gcaatctcaa ggcagctttc c 21 <210> twenty two <211> 20 <212> DNA <213> Artificial Sequence <400> twenty two gagcccgaat tccttggtga 20 <210> twenty three <211> 20 <212> DNA <213> Artificial Sequence <400> twenty three tgctgtttgg ctgtagcagt 20 <210> twenty four <211> 20 <212> DNA <213> Artificial Sequence <400> twenty four tccgtgtagc ggtcaatgtc 20 <210> 25 <211> twenty four <212> DNA <213> Artificial Sequence <400> 25 tgctgatact acagataact cggg 24 <210> 26 <211> twenty one <212> DNA <213> Artificial Sequence <400> 26 agcatcactt agcggaacac t 21 <210> 27 <211> 20 <212> DNA <213> Artificial Sequence <400> 27 gcagggaact tgccactttt 20 <210> 28 <211> 20 <212> DNA <213> Artificial Sequence <400> 28 ggctgatttt ctggcgttgg 20 <210> 29 <211> 20 <212> DNA <213> Artificial Sequence <400> 29 ggctcaagac catgaccgat 20 <210> 30 <211> 20 <212> DNA <213> Artificial Sequence <400> 30 aaagtggaca ccgaagaccc 20 <210> 31 <211> 20 <212> DNA <213> Artificial Sequence <400> 31 cccgtaaaga agcctcccaa 20 <210> 32 <211> twenty one <212> DNA <213> Artificial Sequence <400> 32 gcgggatttc atgtacgaag g 21 <210> 33 <211> 20 <212> DNA <213> Artificial Sequence <400> 33 ggctcaagac catgaccgat 20 <210> 34 <211> 20 <212> DNA <213> Artificial Sequence <400> 34 aaagtggaca ccgaagaccc 20 <210> 35 <211> 20 <212> DNA <213> Artificial Sequence <400> 35 gcggttctgt ggatggagtt 20 <210> 36 <211> 20 <212> DNA <213> Artificial Sequence <400> 36 cggccttgct gcttttagac 20 <210> 37 <211> twenty one <212> DNA <213> Artificial Sequence <400> 37 gcgcacaact tctgccgtaa c 21 <210> 38 <211> twenty one <212> DNA <213> Artificial Sequence <400> 38 gtgcccctga gtccacaaag c 21 <210> 39 <211> 20 <212> DNA <213> Artificial Sequence <400> 39 gcaccgtcaa ggctgagaac 20 <210> 40 <211> 19 <212> DNA <213> Artificial Sequence <400> 40 tggtgaagac gccagtgga 19
Claims
1. A set of differentially expressed genes in the m6A cluster for assessing prognostic survival risk in non-small cell lung cancer, characterized in that, The following genes are included: S100A10, IGFBP1, SLC52A1, KREMEN2, SCPEP1, COL4A3, GSTM2, STRAP, PCDH7, PAK2, CYP4B1, MTUS1, ESPL1, PRRX2, TTK, COL4A4, CCR7, NPC2, and SKA1.
2. A prognostic survival risk scoring model for non-small cell lung cancer constructed using the m6A cluster differentially expressed gene set for assessing prognostic survival risk as described in claim 1, characterized in that... The model is: Survival Risk Score = (0.1543) × S100A10 expression level + (0.0926) × IGFBP1 expression level + (-0.1585) × SLC52A1 expression level + (0.0512) × KREMEN2 expression level + (-0.0787) × SCPEP1 expression level + (0.1195) × COL4A3 expression level + (-0.0915) × GSTM2 expression level + (0.1032) × STRAP expression level + (0.0693) × PCDH7 expression level +(0.1389)×PAK2 expression level +(0.0437)×CYP4B1 expression level +(-0.0950)×MTUS1 expression level +(0.1204)×ESPL1 expression level +(0.0690)×PRRX2 expression level +(-0.1307)×TTK expression level +(-0.1181)×COL4A4 expression level +(-0.0965)×CCR7 expression level +(-0.0891)×NPC2 expression level +(0.0806)×SKA1 expression level 3. The use of the kit for detecting the expression of the m6A cluster differentially expressed gene set for assessing the prognostic survival risk of non-small cell lung cancer as described in claim 1 in the preparation of diagnostic or auxiliary diagnostic products for overall survival of non-small cell lung cancer patients.
4. The method for screening the m6A cluster differentially expressed gene set for assessing prognostic survival risk in non-small cell lung cancer as described in claim 1, characterized in that, Includes the following steps: 1) Download clinical data of non-small cell lung cancer patients, as well as RNA-Seq transcriptome data, somatic mutation data, and CNV data of cancerous tissue and adjacent normal tissue from the database; 2) Extract m6A-related regulatory factors from the transcriptome data obtained in step (1), perform unsupervised clustering analysis on the extracted m6A-related regulatory factors to obtain the genotypes of the m6A-related regulatory factors; perform intersection genotyping on the genotyped m6A-related regulatory factors again; 3) Analyze the differences in the expression of m6A-related regulatory factors of different genotypes in cancerous tissues and adjacent normal tissues of non-small cell lung cancer patients; 4) Based on the expression differences of different subtypes of m6A-related regulatory factors in cancerous tissues and adjacent normal tissues of non-small cell lung cancer patients as determined in step 3), and the complete survival information data of patients, univariate and multivariate Cox regression analysis was performed to obtain the m6A cluster differential gene set for assessing the prognostic survival risk of non-small cell lung cancer.
5. The screening method according to claim 4, characterized in that, Step 1) The databases mentioned are GEO database, TCGA and UCSC-Xena; two NSCLC queues are collected: (GSE50081, GSE68465) and (TCGA-LUSC, TCGA-LUAD).
6. The screening method according to claim 5, characterized in that, Step 2) describes unsupervised clustering analysis performed using the ConsensuClusterPlus package.
7. The screening method according to claim 4, characterized in that, The expression differences of m6A-related regulatory factors in different subtypes in step 3) were analyzed using the "limma" package.
Citation Information
Patent Citations
Gastric cancer prognosis model construction based on m6A-related IncRNA network and clinical application
CN112048559A
Combined marker for predicting prognosis of liver cancer and application of combined marker
CN114592065A