A method for constructing a non-small cell lung cancer prognosis marker and a prognosis model and application thereof
Patent Information
- Application Number
- CN202111514088.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2041-12-10
AI Technical Summary
因此,亟需一种精准的非小细胞肺癌预后模型,申请号为“201911140625.7”名称为“一种肺癌预后标志物、肺癌预后分型模型及其应用”的发明专利提供了一种肺癌预后标志物、肺癌预后分型模型及其应用,根据四种印记基因Den、Peg10、Snrpn/Snurf和Trappca9的总表达量与拷贝数异常表达量的乘积,构建肺癌分型模型,个体化预测患者的五年生存期,其选择了四种肺癌预后标志物,而影响非小细胞肺癌预后的基因作用及表达各不相同且相互影响,故不适用于精确评估非小细胞肺癌的预后
1.本发明提供的m6A相关基因,包括甲基化基因METTL3、METTL14、METTL15、WTAP、VIRMA、RBM15、RBM15B、KIAA1429、ZC3H13,去甲基化基因FTO、ALKBH5, 以及m6A结合、效应蛋白RBMX、YTHDC1、YTHDC2、IGF2BP1、 IGF2BP2、IGF2BP3、YTHDF1、YTHDF2、YTHDF3、HNRNPA2B1、HNRNPC等与非小细胞肺癌预后具有显著的相关性,可作为评估非小细胞肺癌预后情况的灵敏分子标记物;
Smart Images

Figure CN115572762B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of molecular biology and biomedical technology, and in particular relates to a method for constructing and applying a prognostic marker and prognostic model for non-small cell lung cancer. Background Technology
[0002] Lung cancer is one of the fastest-growing malignant tumors in terms of incidence and mortality, posing a significant threat to public health and life. In recent years, the global incidence and mortality rates of lung cancer have been on the rise, particularly in developing countries. Non-small cell lung cancer (NSCLC) accounts for over 80% of all lung cancers, with squamous cell carcinomas (SCC) and adenocarcinomas (AC) being the most prevalent pathological types. Currently, with multidisciplinary comprehensive treatment, the prognosis of NSCLC has greatly improved. Tumor invasion and metastasis are the most significant factors affecting prognosis. Patients with poor prognoses are often those with the fastest and most frequent invasion and metastasis; therefore, early detection, early diagnosis, and early treatment are crucial for improving prognosis.
[0003] The ability to accurately assess the prognosis of non-small cell lung cancer (NSCLC) is crucial for selecting the most timely and effective treatment, directly impacting the chances of saving a patient's life. Therefore, a precise prognostic model for NSCLC is urgently needed. The invention patent application number "201911140625.7," entitled "A Lung Cancer Prognostic Biomarker, a Lung Cancer Prognostic Subtyping Model and Its Application," provides a lung cancer prognostic biomarker, a lung cancer prognostic subtyping model, and its application. It constructs a lung cancer subtyping model based on the product of the total expression level and abnormal copy number expression level of four imprinted genes: Den, Peg10, Snrpn / Snurf, and Trappca9, to individually predict the five-year survival of patients. However, this model selects four lung cancer prognostic biomarkers. Since the functions and expression of genes affecting the prognosis of NSCLC are different and mutually influential, it is not suitable for accurately assessing the prognosis of NSCLC.
[0004] m 6 RNA methylation is a newly emerging type of RNA methylation modification, specifically the methylation of the first nitrogen atom of adenine in the RNA molecule (N1-methyladenosine, m...). 6 A). Studies have shown that m 6 A is a post-transcriptional modification that is highly abundant in eukaryotic tRNA and rRNA. 6 A modification can regulate mRNA translation, and in the development and progression of lung cancer, mRNA... 6 The modifier A plays a crucial role, m 6Abnormal expression of A-related proteins can lead to malignant proliferation, migration, invasion, metastasis, and drug resistance in non-small cell lung cancer. Therefore, m 6 A-modification can serve as a molecular marker for diagnosis, drug treatment, and prognostic assessment, providing a precise and effective assessment for the diagnosis of early-stage lung cancer and the prognosis of advanced-stage patients. Currently, there are no existing technologies for non-small cell lung cancer m-modification. 6 Therefore, providing a prognostic marker for non-small cell lung cancer (NSCLC) and establishing a prognostic model for NSCLC are urgent problems to be solved in the biomedical field. Summary of the Invention
[0005] The purpose of this invention is to provide a method and application for constructing prognostic markers and prognostic models for non-small cell lung cancer. This method can accurately and sensitively predict the prognosis of non-small cell lung cancer and can fully and effectively predict the risk of poor prognosis in patients with non-small cell lung cancer in the early stage, which is conducive to doctors taking the most reasonable intervention measures, saving patients' lives in time and improving their prognosis.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A prognostic marker for non-small cell lung cancer, said marker is m 6 A-related genes, including methylation genes, demethylation genes, and m 6 A binds to at least one of the effector proteins.
[0007] Furthermore, the methylation genes include METTL3, METTL14, METTL15, WTAP, VIRMA, RBM15, RBM15B, KIAA1429, and ZC3H13, and the demethylation genes include FTO, ALKBH5, and m 6 A-binding effector proteins include RBMX, YTHDC1, YTHDC2, IGF2BP1, IGF2BP2, IGF2BP3, YTHDF1, YTHDF2, YTHDF3, HNRNPA2B1, and HNRNPC.
[0008] A method for constructing a prognostic model for non-small cell lung cancer, using the aforementioned markers, is described below. Step 1: Download LUAD and LUSC gene expression level data from the TCGA database, retain LUAD and LUSC samples with survival prognostic information, and obtain NSCLC tumor samples and normal control samples. Use the R3.6.1 language sva package version 3.38.0 to remove batch effects from the LUAD and LUSC expression profile data to obtain the expression profile data of the merged samples. Step 2: Extract m from NSCLC tumor samples in the prepared TCGA dataset 6 The expression levels of related genes were measured; then, the intergroup t-test in R language 3.6.1 was used to compare the expression levels of m. 6 The expression levels of gene A in TCGA NSCLC tumor samples and control samples were analyzed. A p-value less than 0.05 was selected as the screening threshold, and samples with significant differences were retained. 6 Gene A was used to perform hierarchical clustering of expression values based on the centered pearson correlation algorithm using pheatmap Version 1.0.8 in R3.6.1. Step 3: For the m molecules that showed significantly different expression levels in TCGA NSCLC tumor samples and control samples as described in Step 2... 6 The expression levels of A-related genes were analyzed, and tumor subtype analysis was performed on the samples using the ConsensusClusterPlus package Version 1.54.0 in R3.6.1. Based on the obtained disease subtypes, the Kaplan-Meier curve method in the survival package Version 2.41-1 in R3.6.1 was used to evaluate the survival prognostic correlation between different disease subtype sample groups. Then, the clinical information of different subtype sample groups was statistically analyzed and compared. Step 4: Use the CIBERSORT algorithm to calculate the proportion of various immune cells, and then use the intergroup t-test function in R3.6.1 to compare the differences in the proportion of immune cells in different tumor subtypes; Step 5: Calculate the Immune Score using the estimate package in R3.6.1, and then use the t-test function in R3.6.1 to compare the distribution differences of Immune Score among different subtypes. Define the subtype with a high Immune Score as High TIME and the subtype with a low Immune Score as Low TIME. Finally, based on the gene expression levels detected in all samples, use GSEA to screen for KEGG signaling pathways that are significantly associated with different TIME environments. In the GSEA analysis results, select the corrected p-value, i.e., FDR less than 0.05, as the threshold for screening significantly associated KEGG signaling pathways. Step 6: Obtain the expression level of the PD-L1 gene in NSCLC tumor samples. Then, use the intergroup t-test function in R3.6.1 to compare the differences in expression levels of this gene in NSCLC and control samples, and in different subtypes. Finally, use the cor function in R3.6.1 to calculate the m-value of the PD-L1 gene and the significantly differentially expressed genes identified in Step 2. 6 The expression correlation between gene A expression levels; Step 7: Based on the m samples with significantly different expression levels in TCGA NSCLC tumor samples and control samples identified in Step 2. 6 A-related genes were analyzed using the LASSO regression algorithm in the lars package Version 1.2 of R3.6.1 on the TCGA NSCLC tumor dataset, based on the preserved m... 6 Survival regression analysis was performed on genes related to A to screen for the optimal m. 6 Gene combination A is then used to construct a Risk score model (RS model) based on the LASSO coefficients of each element in the optimized gene combination and the gene expression level in the TCGA dataset. The RS calculation formula is as follows: RS = ∑Coefgenes ×Exp genes Coefgenes represents the LASSO coefficient of the target gene, and Exp genes represents the expression level of the target gene in the TCGA dataset.
[0009] Furthermore, in step 1, the LUAD and LUSC gene expression level data contained 585 and 550 samples, respectively, and the detection platform was Illumina HiSeq 2000 RNA Sequencing. The LUAD and LUSC samples with survival prognostic information contained a total of 994 NSCLC tumor samples and 107 normal control samples, and the expression profile data obtained were 994 samples.
[0010] The application of a prognostic marker for non-small cell lung cancer, m 6 At least one of the A-related genes may be used as a molecular marker to assess the prognosis of non-small cell lung cancer, or to be used to formulate a diagnostic reagent for the prognosis of non-small cell lung cancer.
[0011] Application of a prognostic model for non-small cell lung cancer: This model is used to assess the prognosis and improve the prognosis of patients with non-small cell lung cancer.
[0012] The advantages of this invention are: 1. The m provided by this invention 6A-related genes include methylation genes METTL3, METTL14, METTL15, WTAP, VIRMA, RBM15, RBM15B, KIAA1429, and ZC3H13, demethylation genes FTO and ALKBH5, and m 6 A-binding proteins, effector proteins RBMX, YTHDC1, YTHDC2, IGF2BP1, IGF2BP2, IGF2BP3, YTHDF1, YTHDF2, YTHDF3, HNRNPA2B1, and HNRNPC are significantly correlated with the prognosis of non-small cell lung cancer and can be used as sensitive molecular markers for assessing the prognosis of non-small cell lung cancer. 2. The non-small cell lung cancer prognostic model provided by this invention retains significantly different m 6 By analyzing tumor subtypes using gene A and comprehensively comparing the differences in the proportion of immune cells in different tumor subtypes, the prognostic assessment results become more sensitive, accurate, and reliable. This enables timely monitoring of the prognosis of non-small cell lung cancer patients and early and effective prediction of adverse prognostic risks in non-small cell lung cancer patients. This allows doctors to take the most appropriate intervention measures to save patients' lives and improve their prognosis in a timely manner. 3. This invention studied m in non-small cell lung cancer samples and control samples. 6 Expression profiles of A-related genes were analyzed, and m was evaluated. 6 The correlation between A-related genes and NSCLC prognosis, PD-L1, and TIME provides new insights for further research on the regulatory mechanism of methylation-related TIME and immunotherapy for non-small cell lung cancer. 4. The m provided by this invention 6 A-related genes are closely associated with the clinical characteristics of tumor recurrence and survival time, among which five m 6 RS of A-related genes HNRNPA2B1, HNRNPC, IGF2BP1, METTL3, and RBM15B are independent risk factors for the prognosis of non-small cell lung cancer (NSCLC) patients. They can be used to establish a prognostic model for NSCLC and can accurately and sensitively assess the prognosis of NSCLC patients. Attached Figure Description
[0013] Figure 1 This is a graph showing the sample relationship before and after SVA batch effect removal. A represents before removal, and B represents after removal.
[0014] Figure 2 m, which is significantly expressed in NSCLC tumor and control samples 6 A is a heatmap showing the expression level of related genes; B is a distribution map of expression levels in NSCLC tumor and control samples.
[0015] Figure 3In the diagram, A is a clustering diagram of NSCLC tumor sample subtypes, and B is a KM survival curve for different subtypes.
[0016] Figure 4 This is a distribution map of the proportions of 22 immune cells in different NSCLC tumor subtypes and a comparison p-value.
[0017] Figure 5 These are score charts for different NSCLC tumor subtypes: A is the Stromal score chart; B is the Immune score chart; and C is the ESTIMATE score chart.
[0018] Figure 6 This is an ES-plot of the KEGG signaling pathway that is significantly related to the TIME environment.
[0019] Figure 7 In Figure A, the expression level distribution of the PD-L1 gene in NSCLC, controls, and different subtypes is shown; in Figure B, the correlation between the expression levels of the PD-L1 gene and differentially expressed genes is shown.
[0020] Figure 8 These are the RS score distribution and clinical survival information distribution graphs. A is the TCGA training set sample, and B is the GSE50081 validation set sample. The top graph shows the RS distribution in the samples, the middle graph shows the survival prognosis information scatter plot, and the bottom graph shows the ROC curve based on the PS feature prognosis.
[0021] Figure 9 This is a prognostic KM curve based on the RS prediction model. A is the TCGA training set sample, and B is the GSE50081 validation set sample. The top figure is the survival rate and survival time relationship curve of the low-risk group and the high-risk group, i.e. the prognostic KM curve. The middle figure is the number of people at risk in the low-risk group and the high-risk group at different survival times. The bottom figure is the histogram of the number of samples reviewed in the low-risk group and the high-risk group at different survival times.
[0022] Figure 10 In the diagram, A represents the distribution of RS in different subtypes, and B represents the distribution of RS in different recurrence state groups.
[0023] Figure 11 This is a scatter plot showing the correlation between RS prognostic information and different types of immune cells in TIMER. Detailed Implementation Example
[0024] A prognostic marker for non-small cell lung cancer, said marker is m 6 A-related genes, including methylation genes, demethylation genes, and m 6 A binds to at least one of the effector proteins.
[0025] Furthermore, the methylation genes include METTL3, METTL14, METTL15, WTAP, VIRMA, RBM15, RBM15B, KIAA1429, and ZC3H13, and the demethylation genes include FTO, ALKBH5, and m 6 A-binding effector proteins include RBMX, YTHDC1, YTHDC2, IGF2BP1, IGF2BP2, IGF2BP3, YTHDF1, YTHDF2, YTHDF3, HNRNPA2B1, and HNRNPC.
[0026] A method for constructing a prognostic model for non-small cell lung cancer, using the aforementioned markers, is described below. Data Sources and Preprocessing: LUAD and LUSC gene expression level data (both standardized log(FPKM+1,2) expression values) were downloaded from the TCGA database, containing 585 and 550 samples respectively. The detection platform for both was Illumina HiSeq 2000 RNA Sequencing. Based on the clinical information of samples downloaded at the same time, LUAD and LUSC samples with survival prognostic information were retained, resulting in 994 NSCLC tumor samples and 107 normal control samples included in this analysis. TCGA samples were used as the training dataset. Since LUAD and LUSC came from different batches of gene expression level data, batch effect removal was performed on the LUAD and LUSC expression profile data using the sva package version 3.38.0 in R3.6.1, resulting in the merged expression profile data of 994 samples. Meanwhile, NSCLC expression profile data numbered GSE50081 was downloaded from the NCBI GEO (National Center for Biotechnology Information, Gene Expression Omnibus) database. This data contains 181 NSCLC tumor samples with clinical prognostic information. The detection platform was GPL570 Affymetrix Human Genome U133 Plus 2.0 Array. This data served as a validation dataset.
[0027] Merging LUAD and LUSC expression profiles and removing batch effects: Download the corresponding expression profile data. As described in the method, first, remove batch effects from the LUAD and LUSC gene expression profile data using the SVA algorithm, and merge them into a single dataset. The sample relationship before and after batch effect removal is shown in the figure below. Figure 1As shown in the figure, before batch effect removal, the LUAD and LUSC samples were separate and very different. After batch effect removal, the LUAD and LUSC samples were completely mixed together, indicating that the difference in detection level between the samples was smaller and the correlation was higher after removing the batch effect, thus achieving the purpose of batch effect removal. The expression levels of the datasets before and after batch effect removal are shown in Appendix Table 1.
[0028] m 6 Comparative analysis of gene A expression levels: m was extracted from NSCLC tumor samples in the compiled TCGA dataset. 6 Expression levels of A-related genes: methylated genes (METTL3, METTL14, METTL15, WTAP, VIRMA, RBM15, RBM15B, KIAA1429, ZC3H13); demethylated genes (FTO, ALKBH5); m6A binding and effector proteins (RBMX, YTHDC1, YTHDC2, IGF2BP1, IGF2BP2, IGF2BP3, YTHDF1, YTHDF2, YTHDF3, HNRNPA2B1, HNRNPC); then, the intergroup t-test in R language 3.6.1 was used to compare the expression levels of m... 6 The expression levels of gene A in TCGA NSCLC tumor samples and control samples were analyzed. A p-value less than 0.05 was selected as the screening threshold, and samples with significant differences were retained. 6 Gene A, based on its expression level, was subjected to hierarchical clustering using the centered Pearson correlation algorithm based on pheatmap Version 1.0.8 in R3.6.1; m was extracted from the merged LUAD and LUSC samples. 6 The expression levels of A-related genes are shown in Appendix Table 2. Then, the intergroup t-test in R language 3.6.1 was used to compare the expression levels of m-related genes. 6 The expression levels of gene A in TCGA NSCLC tumor samples and control samples were analyzed, and a total of 15 significantly differentially expressed m were obtained. 6 For a list of A-related genes and p-values, please refer to Appendix 2. The m-values for significantly differentially expressed genes are also shown. 6 The heatmap showing the expression levels of A-related genes in the samples is shown below. Figure 2 For the expression level distribution of A in tumor and control samples, please see [link to relevant documentation]. Figure 2 As shown in B.
[0029] Based on m 6 Subtype analysis of A-related genes: Based on the 15 m genes whose expression levels differed significantly between TCGA NSCLC tumor samples and control samples, retained from the previous step.6 A-related genes were used to perform subtype analysis on NSCLC tumor samples, and the results were as follows: Figure 3 As shown in Figure A, two different subtypes were obtained: Subtype 1 and Subtype 2, containing 696 and 298 tumor samples, respectively. The survival Kaplan-Meier curve method from the survival package Version 2.41-1 in R3.6.1 was then used to evaluate the survival prognostic correlation between the different disease subtype sample groups, as shown below. Figure 3 As shown in Table B, the clinical information of samples from different subtypes was then statistically analyzed and compared. The results showed that there were significant differences in survival prognosis information among different subtypes. Finally, the clinical information of samples included in different subtypes was statistically compared using the chi-square test, as shown in Table 1.
[0030] Table 1. Comparison of Statistical Information on Clinical Information of Samples in Different Subtype Groups CIBERSORT assesses and compares the proportion of immune cells in samples: The concept of the tumor immune microenvironment refers to the large number of immune cells that often accumulate inside and around the tumor. These immune cells are interconnected, and there are intricate relationships between them, as well as between tumor cells and immune cells. Immune cells are diverse, so the analysis of the immune microenvironment, or immune infiltration, essentially aims to clarify the compositional proportions of immune cells within tumor tissue. This invention uses CIBERSORT to calculate the proportions of various immune cells based on the expression levels of the NSCLC tumor samples included in the analysis. CIBERSORT is a tool that deconvolves the expression matrix of immune cell subtypes based on the principle of linear support vector regression. Then, the differences in the proportions of different immune cell subtypes obtained in the previous step are compared using the inter-group t-test function in R3.6.1.
[0031] Based on the expression profile data of TCGA NSCLC tumor samples, the CIBERSORT algorithm was used to calculate the immune cell type classification for each sample, yielding the proportions of 22 immune cell types. Figure 4 As shown, the differences in the proportions of various immune cells among different subtype groups were then analyzed and compared using the t-test between groups, and a total of 13 significantly different immune cells were obtained.
[0032] Estimate the subtype immune score: Malignant tumor tissue includes not only tumor cells, but also tumor-associated normal epithelial cells and stromal cells, immune cells, and vascular cells. Infiltrating stromal cells and immune cells are the main components of normal cells in tumor tissue, interfering with tumor signaling in molecular studies and playing an important role in tumor biology. This application uses the estimate package in R3.6.1 to calculate the Immune Score, and then uses the intergroup t-test function in R3.6.1 to compare the distribution differences of Immune Score among different subtypes. The subtype with a high Immune Score is defined as: High tumor immune microenvironment (High TIME), and the subtype with a low Immune Score is defined as: Low tumor immune microenvironment (Low TIME). Finally, based on the gene expression levels detected in all samples, GSEA (Gene Set Enrichment Analysis) is used to screen for KEGG signaling pathways significantly associated with different TIME environments. The GSEA analysis results include three key statistical values: Enrichment score (ES): This is the raw result of GSEA analysis. It reflects the degree of enrichment of a functional gene set before or after the sequence after all hybridization data are sorted. The basic principle of ES value calculation is to scan the sorted sequence. When a gene in a functional gene set appears, the ES value is increased; otherwise, the ES value is decreased. Therefore, ES is a dynamic value throughout the scanning process. The final ES value is determined by defining the position of the hybridization data sorted sequence as 0 and the ES value as the maximum deviation from the sorted sequence. When the ES value is positive, it indicates that a functional gene set is enriched before the sorted sequence. When the ES value is negative, it indicates that a functional gene set is enriched after the sorted sequence. Normalized enrichment score (NES): This is the main statistical measure of the results of functional gene set enrichment analysis. It is calculated based on whether the genes in the analyzed microarray dataset appear in a functional gene set. However, the number of genes contained in each functional gene set is different, and the correlation between different functional gene sets and microarray data is also different. Therefore, when comparing the enrichment degree of microarray datasets in different functional gene sets, it is necessary to standardize the ES. The nominal p-value describes the statistical significance of the enrichment score obtained for a specific subset of functional genes. A smaller p-value indicates better gene enrichment. Generally, the larger the absolute value of the NES, the smaller the significance p-value, indicating a higher enrichment level for the functional gene set and thus higher reliability of the analysis results. A corrected p-value (FDR) less than 0.05 was selected as the threshold for screening significantly associated KEGG signaling pathways. The Immune Score was calculated using the estimate package, and then the differences in the distribution of scores between different subtypes were compared using the t-test function in R3.6.1. Figure 5 As shown, the results indicate that Subtype 1 had higher scores across all categories. Following the definition in the methodology, Subtype 1 with higher scores was defined as the High TIME group, and Subtype 2 with lower Immune scores was defined as the Low TIME group. Finally, based on the gene expression levels detected in all samples, the GSEA algorithm was used to screen for KEGG signaling pathways associated with High and Low TIME environments. The results are shown in [Figure number missing]. Figure 6 A total of 10 KEGG pathways were identified, such as CELL_CYCLE.
[0033] Analysis of PD-L1 gene expression levels: First, the expression level of the PD-L1 gene in NSCLC tumor samples was obtained. Then, the intergroup t-test function in R3.6.1 was used to compare the differences in expression levels of this gene in NSCLC and control samples, and in different subtypes. Figure 7 As shown in Figure A; then, the `cor` function in R3.6.1 is used to calculate the m-values of the PD-L1 gene and the significantly differentially expressed genes obtained from screening in the TCGA and GSE50081 datasets, respectively. 6 The expression correlation between gene A expression levels is as follows: Figure 7 As shown in B.
[0034] Constructing a prognostic risk prediction model: based on the 15 m values obtained from previous screening. 6 For genes related to A, as described in the method, the optimal gene combination was screened using the LASSO algorithm (parameter diagram attached). Figure 2 As shown, when the optimal combination is obtained, 1se = 0.02935), and finally these 5 m values are obtained. 6 The genes related to A are HNRNPA2B1, HNRNPC, IGF2BP1, METTL3, and RBM15B. Based on the LASSO regression coefficients of each gene, the following RS calculation formula was constructed: RS = (0.003612586) *ExpHNRNPA2B1 + (0.052669305) *Exp HNRNPC + (0.006522287) *Exp IGF2BP1 + (-0.079231555) *Exp METTL3 + (0.001102796) *ExpRBM15B RS scores were calculated for samples in both the TCGA training set and the GSE50081 validation set. For the distribution of RS scores and clinical survival information, please refer to [link to relevant documentation]. Figure 8 As shown in A and B.
[0035] Performance Evaluation and Comparison of RS Survival Prognostic Risk Prediction Models: RS values were calculated for the TCGA training set and the GSE50081 validation dataset. Then, based on the median RS value, samples from both sets were divided into High (RS score equal to or higher than the median RS value) and Low (RS score lower than the median RS value) groups. The Kaplan-Meier curve method from the survival package in R3.6.1 was used to evaluate the correlation between the High and Low groupings and actual disease prognostic information. The KM curves are shown below. Figure 9 As shown, the ROC curve results based on characteristic genes are as follows: Figure 8 As shown in the figure. The results show that in the TCGA training set and GSE50081 validation set samples, there is a significant correlation between the different risk groups obtained after the samples are predicted based on the RS model and the actual prognosis; in the TCGA dataset samples, univariate and multivariate Cox regression analysis in the survival package Version 2.41-1 was used to examine whether the RS factor is an independent prognostic factor; finally, the differences in the distribution of RS values among different clinical factors were examined in the selected independent clinical prognostic factors and different subtypes. The clinical information of the NSCLC tumor samples included in the analysis was analyzed, and the results are shown in Table 2. The results show that based on 5 m 6 The prognostic characteristics of gene A are independent of other clinical factors. Furthermore, tumor recurrence is also an independent clinical prognostic factor. Analysis of variance was then used to calculate the differences in RS distributions among different tumor recurrence states and different subtype groups for each clinical factor. The results are as follows: Figure 10 As shown, the distribution of RS eigenvalues differed significantly between the two clinical factor groups, with significantly higher RS values in the Subtype 2 and tumor recurrence sample groups, which had poorer prognoses.
[0036] Table 2. Screening information related to prognosis of each clinical factor Correlation analysis between prognostic features and immunity based on m6A gene: First, based on the expression level of m6A gene in TCGA NSCLC tumor samples, the online tool TIMER (Tumor Immune Estimation Resource) was used to analyze relevant immune cells. This tool can estimate the composition ratio of six immune infiltrating cells (B cells, CD4+ T cells, CD8+ T cells, neutrophils, macrophages, and dendritic cells). Results are attached. Figure 10 As shown in the figure. Then, the correlation between each immune cell subtype and the RS values based on TCGA samples was calculated to observe the association between prognostic signals based on characteristic RNAs and the infiltration of various immune cell subtypes. The Pearson correlation coefficient between RS eigenvalues and various immune cells was calculated using the `cor` function in R language 3.6.1, and the results are shown in the figure. Figure 11 As shown.
[0037] This invention also provides an application of a prognostic marker for non-small cell lung cancer, which is m 6 At least one of the A-related genes is used as a molecular marker to assess the prognosis of non-small cell lung cancer, or is used to formulate a non-small cell lung cancer prognostic diagnostic reagent; it also includes the application of a non-small cell lung cancer prognostic model, which is used to assess the prognosis of non-small cell lung cancer patients and improve prognosis.
Claims
1. A prognostic marker for non-small cell lung cancer, characterized in that: With m 6 A combination of genes HNRNPA2B1, HNRNPC, IGF2BP1, METTL3, and RBM15B is used as a prognostic marker for non-small cell lung cancer.
2. The application of the prognostic marker for non-small cell lung cancer as described in claim 1, characterized in that: m 6 The application of a combination of A-related genes HNRNPA2B1, HNRNPC, IGF2BP1, METTL3, and RBM15B as prognostic markers for non-small cell lung cancer in the preparation of prognostic assessment products for non-small cell lung cancer.
Citation Information
Patent Citations
Lung-cancer prognostic markers, lung-cancer prognostic typing model and application of lung-cancer prognostic typing model
CN111206097A