Gene signature for assessing overall survival and characterizing tumor immune microenvironment and research method thereof

CN117059166BActive Publication Date: 2026-09-15AFFILIATED HUSN HOSPITAL OF FUDAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310915518.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-25
Publication Date
2026-09-15
Estimated Expiration
2043-07-25

AI Technical Summary

Benefits of technology

[0032] 1. This invention proposes an m-type method for assessing overall survival and characterizing the tumor immune microenvironment. 6 The A-related gene features specifically include 20 protein-coding genes: IGFBP3, HDAC2, HMGN1, TCP1, PABPC4, RDH16, HDDC2, CFHR5, GYS1, MAPRE1, GYS2, BLMH, YBX1, NAP1L1, INTS8, MARCKSL1, STX6, MASP2, UQCRH, and XPNPEP1. This invention identifies and verifies that these protein-coding gene features are reliable and robust for assessing the prognosis of HCC.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117059166B_ABST
    Figure CN117059166B_ABST
Patent Text Reader

Abstract

This invention discloses a gene signature for assessing overall survival and characterizing the tumor immune microenvironment, and its research method, belonging to the field of biomedical technology. The gene signature for assessing overall survival and characterizing the tumor immune microenvironment is m. 6 The A-related gene characteristics specifically include 20 protein-coding genes. This invention comprehensively analyzes differentially expressed genes in TCGA to identify m-related genes associated with HCC prognosis. 6 A characteristics were assessed; furthermore, GSEA, tumor mutational burden (TMB), immune infiltration, and treatment response were evaluated, and mass spectrometry proteomics and multiplex immunofluorescence assays were performed for validation, ultimately yielding m 6 Feature A has the ability to significantly assess overall survival and characterize the tumor immune microenvironment in HCC, and can serve as a useful method for risk stratification management, providing valuable clues for selecting appropriate treatment strategies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedical technology, and in particular to gene characterization and research methods for assessing overall survival and characterizing the tumor immune microenvironment. Background Technology

[0002] Hepatocellular carcinoma (HCC) is the sixth most common malignant tumor worldwide and the fourth leading cause of cancer-related death. Multiple treatment options, including targeted therapy, chemotherapy, and immunotherapy, have been approved for HCC. However, the overall prognosis and survival for HCC patients remain poor, with approximately 781,000 deaths annually. Therefore, it is necessary to develop effective methods to identify patients with poor prognoses. For high-risk HCC subgroups, better clinical management and rational treatment strategies are needed to improve overall outcomes.

[0003] Currently, with the advancement of new detection technologies, N in RNA molecules 6 -Methyladenine modification (m 6 A) This is receiving increasing attention because it has a significant impact on RNA stability, export, splicing, or translation. 6 The A modification is catalyzed by an assembled complex consisting of multiple methyltransferases acting as "writers" and m acting as "readers". 6 It consists of an RNA-binding protein and a demethylase that acts as an "eraser." Mature m 6 A "writer" includes METTL3, METTL14, WTAP, RBM15, and ZC3H13. These "writers" are responsible for adding methylation units to the target RNA. It is worth noting that m 6 A modification is a dynamic and reversible process that can be removed by "erasers"—RNA demethylases such as FTO and ALKBH5. 6 A-modified RNA can be recognized by various RNA-binding protein "readers" to determine RNA fate, including YTHDC1, YTHDF1 / 2 / 3, the insulin-like growth factor mRNA-binding protein family IGFBP1 / 2 / 3, and RBMX.

[0004] Increasing evidence suggests that m 6 A-modification plays a crucial role in the malignant evolution of cancer and has become a hot topic in cancer biology research. Numerous studies have described m... 6 How A modifications regulate the development of HCC. For example, the "writer" METTL14 can mediate the m of EGFR. 6 A methylation inhibits EGFR / PI3K / AKT activity and suppresses EMT and translocation in HCC. 6 A “reader” YTHDF1 can communicate with m6 A-modified ATG2A and ATG14 mRNAs bind to promote autophagy in HCC cells, thereby increasing their translation under hypoxic conditions. The "eraser" FTO is also involved in HCC progression; it can modulate cancer stem cell properties by demethylating SOX2, KLF4, and NANOG mRNAs. Given m... 6 The important regulatory function of A modification, m 6 A-related genes may be a promising molecular profiling tool for patient stratification. Recently, several m-related genes have been identified in various cancers, including HCC. 6 A related gene signature. However, most signatures lack validation in independent cohorts, and the reliability and robustness of these models remain elusive. Summary of the Invention

[0005] To address the aforementioned problems, this invention aims to provide a gene signature and research method for assessing overall survival and characterizing the tumor immune microenvironment, identifying and validating m 6 A-related protein encoding gene features can help indicate prognosis, characterize the tumor immune microenvironment, and predict the potential efficacy of treatment for HCC patients, thereby aiding in risk stratification management and selecting appropriate treatment strategies for high-risk HCC patients.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0007] Genetic features used to assess overall survival and characterize the tumor immune microenvironment, characterized in that: the genetic features are m 6 A-related gene characteristics.

[0008] Furthermore, the m 6 The A-related gene signature includes 20 protein-coding genes: IGFBP3, HDAC2, HMGN1, TCP1, PABPC4, RDH16, HDDC2, CFHR5, GYS1, MAPRE1, GYS2, BLMH, YBX1, NAP1L1, INTS8, MARCKSL1, STX6, MASP2, UQCRH, and XPNPEP1.

[0009] Furthermore, a research method for assessing overall survival and characterizing the genetic features of the tumor immune microenvironment is characterized by comprising the following steps:

[0010] S1: Obtain gene expression data from HCC tumor samples and normal samples;

[0011] S2: Construct m 6 A score system was used to evaluate all HCC tumor samples.

[0012] S3: Identify high-risk and low-risk HCC populations and perform GO and KEGG pathway enrichment analysis to determine gene signatures for assessing overall survival and characterizing the tumor immune microenvironment;

[0013] S4: Assess GSEA, tumor mutational burden, immune infiltration, and treatment response in HCC tumor samples;

[0014] S5: Validated by mass spectrometry proteomics and multiplex immunofluorescence assays.

[0015] Furthermore, the specific operation of step S2 includes the following steps:

[0016] S201: The expression of all genes in HCC tumor samples is analyzed using a consensus clustering algorithm to generate different clusters;

[0017] S202: Obtaining overlapping differentially expressed genes (DEGs) using Venn diagrams;

[0018] S203: Using Cox proportional hazards regression analysis to determine different m 6 The prognostic value of protein-coding genes with A modification patterns was used to determine the genes used in constructing m 6 A Score scoring system for protein-coding genes;

[0019] S204: Perform principal component analysis on the protein-coding genes identified in step S203 to establish m 6 A Score scoring system.

[0020] Furthermore, the specific operations of mass spectrometry proteomics in step S5 include the following steps:

[0021] S501a: Protein isolation from formalin-fixed paraffin-embedded tumor samples;

[0022] S502 a: Digest the isolated proteins;

[0023] S503 a: Peptide sequencing by mass spectrometry;

[0024] S504 a: Data analysis to identify proteins.

[0025] Furthermore, the specific procedures for multiplex immunofluorescence assay in step S5 include the following steps:

[0026] S501b: After dewaxing and dehydrating paraffin sections, glass slides are immersed in citrate antigen retrieval buffer and boiling water for antigen retrieval;

[0027] S502 b: Incubate the sample with 3% H2O2 at room temperature, block the solution with 5% FBS at 37°C for 30 minutes, and then incubate with the primary antibody at 4°C overnight;

[0028] S503 b: Wash the sample with phosphate-buffered saline (PBS), incubate with Cy5 / Sp Green-labeled secondary antibody at 37°C, and then wash again with PBS;

[0029] S504 b: Incubate the solution with Hoechst at room temperature for 15 minutes, remove the previous primary / secondary antibody by microwave treatment, and repeat the continuous staining process.

[0030] S505 b: Perform multiplex immunofluorescence assay on the sample.

[0031] The beneficial effects of this invention are:

[0032] 1. This invention proposes an m-type method for assessing overall survival and characterizing the tumor immune microenvironment. 6 The A-related gene features specifically include 20 protein-coding genes: IGFBP3, HDAC2, HMGN1, TCP1, PABPC4, RDH16, HDDC2, CFHR5, GYS1, MAPRE1, GYS2, BLMH, YBX1, NAP1L1, INTS8, MARCKSL1, STX6, MASP2, UQCRH, and XPNPEP1. This invention identifies and verifies that these protein-coding gene features are reliable and robust for assessing the prognosis of HCC.

[0033] 2. This invention uses a comprehensive analysis of differentially expressed genes in TCGA to identify m genes associated with HCC prognosis. 6 A characteristics were assessed; furthermore, GSEA, tumor mutational burden (TMB), immune infiltration, and treatment response were evaluated, and mass spectrometry proteomics and multiplex immunofluorescence assays were performed for validation, ultimately yielding m 6 Feature A has the ability to significantly assess overall survival and characterize the tumor immune microenvironment in HCC, and can serve as an effective method for risk stratification management, providing valuable clues for selecting appropriate treatment strategies. Attached Figure Description

[0034] Figure 1 These are the D15 and D17 mass spectra of some samples in this invention.

[0035] Figure 2 This represents the protein identification quantity results of tissue samples in this invention.

[0036] Figure 3 To establish m in HCC in this invention 6 A flowchart illustrating the characteristics of genes encoding A-related proteins.

[0037] Figure 4 For the present invention in HCC m6 When identifying A-related gene characteristics, 15m uniCox regression analysis was used. 6 A forest plot showing the relationship between related genes, OS, and gene expression, and a scree plot used to determine the optimal number of principal components.

[0038] Figure 5 In this invention, m in HCC 6 Bubble charts of GO and KEGG enrichment analyses of differentially expressed genes (DEG) in the three clusters during A-related gene feature identification.

[0039] Figure 6 In this invention, m in HCC 6 When identifying A-related gene characteristics, Venn diagrams of the number of DEGs in the three clusters were used, along with mass spectrometry to verify genes and prognostic-related genes.

[0040] Figure 7 In this invention, m in HCC 6 When identifying A-related gene characteristics, the expression levels and copy number variation (CNV) frequencies of 20 genes in HCC tumors and adjacent normal specimens were analyzed.

[0041] Figure 8 m is the SRAMP prediction in this invention. 6 The m of BLMH, HDAC2, INTS8, and NAP1L1 in the A-related gene sequences 6 A methylation site.

[0042] Figure 9 m is the SRAMP prediction in this invention. 6 The m of CFHR5, HDDC2, MAPRE1, and PABPC4 in the A-related gene sequences 6 A methylation site.

[0043] Figure 10 m is the SRAMP prediction in this invention. 6 The m of GYS1, HMGN1, MARCKSL1, and RDH16 in the A-related gene sequences 6 A methylation site.

[0044] Figure 11 m is the SRAMP prediction in this invention. 6 The m-sequence of GYS2, IGFBP3, MASP2, and STX6 in the A-related gene sequence 6 A methylation site.

[0045] Figure 12 m is the SRAMP prediction in this invention. 6 The m-sequence of related genes TCP1, UQCRH, XPNPEP1, and YBX1 6A methylation site.

[0046] Figure 13 For m in the TCGA-HCC queue of this invention 6 Clinical significance of related features A.

[0047] Figure 14 m in the independent HCC queue of this invention 6 Proteomics validation results of A-related gene characteristics.

[0048] Figure 15 m in this invention 6 Functional analysis of related features and prediction results of sensitivity to targeted / chemotherapy drugs.

[0049] Figure 16 m in this invention 6 A-related features predict tumor somatic mutations and immunotherapy sensitivity.

[0050] Figure 17 m in this invention 6 The results of a study on the correlation between A-related features and the infiltration of immune cells and related genes.

[0051] Figure 18 This is the verification result of immune cell infiltration in hepatocellular carcinoma specimens in this invention. Detailed Implementation

[0052] To enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0053] Example 1:

[0054] Example 1 provides a gene signature for assessing overall survival and characterizing the tumor immune microenvironment, wherein the gene signature is m 6 A-related gene characteristics.

[0055] Specifically, the m 6 The A-related gene signature includes 20 protein-coding genes: IGFBP3, HDAC2, HMGN1, TCP1, PABPC4, RDH16, HDDC2, CFHR5, GYS1, MAPRE1, GYS2, BLMH, YBX1, NAP1L1, INTS8, MARCKSL1, STX6, MASP2, UQCRH, and XPNPEP1.

[0056] Example 2:

[0057] Example 2 provides a research method for assessing overall survival and characterizing the gene features of the tumor immune microenvironment, including the following steps:

[0058] S1: Obtain gene expression data from HCC tumor samples and normal samples;

[0059] S2: Construct m 6 A score system was used to evaluate all HCC tumor samples.

[0060] S3: Identify high-risk and low-risk HCC populations and perform GO and KEGG pathway enrichment analysis to determine gene signatures for assessing overall survival and characterizing the tumor immune microenvironment;

[0061] S4: Assess GSEA, tumor mutational burden, immune infiltration, and treatment response in HCC tumor samples;

[0062] S5: Validated by mass spectrometry proteomics and multiplex immunofluorescence assays.

[0063] Specifically, the research process includes the following:

[0064] Materials and methods:

[0065] Gene expression data from HCC samples

[0066] Raw gene expression data and corresponding clinical information for HCC were extracted from the Cancer Genome Atlas (TCGA) database (https: / / portal.gdc.cancer.gov / ). Transcriptional data from 374 tumor specimens and 50 normal specimens were included. HTseq counts were normalized based on the TPM (transcripts per million) method.

[0067] Exclusion criteria were as follows: (i) histological diagnosis excluding HCC, (ii) extremely low gene expression values, and (iii) incomplete clinical data and follow-up time <30 days. ID conversion and RNA classification were performed by Perl (The Perl Programming Language version 5.30.1, https: / / www.perl.org / ).

[0068] m in HCC 6 Construction of related signatures

[0069] To generate m 6 A related feature and quantification of m for each patient 6 In this invention, a scoring system (referred to as m) is constructed, modified by A. 6 A score was used to evaluate all HCC patients.

[0070] First, the expression of all genes in HCC tumor samples was analyzed using a consensus clustering algorithm to generate different clusters; then, overlapping differentially expressed gene DEGs were obtained using Venn diagrams; next, Cox proportional hazards regression analysis was used to determine the different m... 6 The prognostic value of protein-coding genes with A modification patterns was used to determine the genes used in constructing m 6 A Score scoring system uses protein-coding genes (DEGs with significant prognostic value to construct m) 6 A score scoring system); given the prognostic protein encoding DEG in HCC, principal component analysis (PCA) was then performed to establish m 6 A Score. PCA is a dimensionality reduction method that has been widely used in gene expression analysis. In this invention, principal components (PC)1 and 2 are added as the final gene feature score.

[0071]

[0072] In the formula, i and j are m in HCC 6 The order and total number of prognostic proteins related to A were determined. Finally, m... 6 Further analysis was conducted on the Z-score of the A score.

[0073] For different clusters, Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways were identified and visualized using bubble charts.

[0074] HCC Participant Registration

[0075] Between March 2015 and June 2020, Huashan Hospital of Fudan University enrolled an independent cohort of 101 HCC patients. None of the patients had received radiotherapy or chemotherapy prior to surgery. Informed consent was provided by each patient. The study protocol was approved by the Human Research Ethics Committee of Huashan Hospital of Fudan University. Basic information about the samples is shown in Tables 1 and 2 below.

[0076] Table 1. Clinical information of the independent HCC cohort at Huashan Hospital

[0077]

[0078] Table 2. Expression of 20 m6A-related genes in an independent HCC cohort at Huashan Hospital

[0079]

[0080] Continued from Table 2(1)

[0081]

[0082] Continued from Table 2(2)

[0083]

[0084] General mutation information

[0085] Copy number variation (CNV) data were downloaded from the TCGA database. Gene mutation data were also downloaded from the TCGA database in this invention, and data was analyzed at high / low m... 6 Gene mutations were identified in the AScore group. Furthermore, the tumor mutational burden (TMB) for each sample was calculated using a Perl script (https: / / www.perl.org / ), where TMB is defined as the total number of mutations per megabase in the tumor tissue.

[0086] Gene set enrichment analysis (GSEA)

[0087] GSEA is a computational method that performs GO and KEGG analyses on a given list of genes. GSEA software version 4.2.3 (https: / / www.gsea-msigdb.org / ) is used to predict m... 6 The potential functionality of A-related signatures. Combined with m 6 High-risk and low-risk populations were identified by a score. GO and KEGG pathway enrichment analyses were performed to visualize various genes involved in different pathways, biological functions, and expression patterns. Data have been corrected for multiple tests (replication count = 1000).

[0088] Chemotherapy response prediction

[0089] The R package in pRRhetic is used to predict sensitivity to common chemotherapy drugs. Sensitivity represents the effectiveness of a substance in inhibiting a specific biological or biochemical function. Differences between groups were tested using the Wilcoxon signed-rank test.

[0090] Predicting response to immunotherapy

[0091] The Tumor Immune Dysfunction and Exclusion (TIDE) algorithm was used for genetic characterization of T-cell dysfunction and prediction of patient response to cancer immunotherapy. TIDE utilizes T-cell dysfunction and exclusion features to model immune escape from tumors with varying levels of cytotoxic T lymphocytes, and the TIDE score correlates with the characteristics of tumor immune escape. Higher tumor TIDE scores are associated with poorer immune checkpoint blockade responses.

[0092] Analysis of immune cell infiltration and immune function

[0093] Single-sample gene set enrichment analysis (ssGSEA) was performed to calculate the immune function score for each HCC patient, further quantifying the composition of tumor-infiltrating immune cells. Based on the immune score, the degree of immune cell infiltration in HCC tissue was quantified. Table 3 below lists the gene characteristics of the immune cells involved in this study.

[0094] Table 3. Genetic characteristics of immune cells involved in this study (partial)

[0095]

[0096] Mass spectrometry proteomics detection

[0097] Proteomics testing is performed in four steps: protein isolation from formalin-fixed paraffin-embedded tumor samples, protein digestion, peptide sequencing by mass spectrometry (MS), and data analysis to identify proteins.

[0098] The processing details of the raw MS files are as follows: MS raw file entries were searched in the NCBI Human Refseq database using the Mascot search engine (v2.3, Matrix Science Inc.; version 04-07-2013, total 32015). The precursor and daughter ion mass biases were set to 20 and 50 ppm (QExactive HF) or 0.5 Da, respectively. The theoretical protein cleavage sites were arginine (R) and lysine (K), with a maximum of two missing cleavage sites allowed. Aminomethylation (C) was used for immobilization, and acetyl groups (N-terminus of the protein) and oxidation (M) were used for dynamic modification. All identified peptides were derived from the area calculated under the MS1 peak, with a false discovery rate controlled to 1%. At the protein characterization level, only proteins containing at least one unique peptide and two high-quality peptides (strict peptides, i.e., Mascot Ion Score > 20) were retained. For protein quantification, an intensity-based absolute quantization algorithm (i.e., the iBAQ algorithm) was used, and each iBAQ value was subsequently normalized to a FOT-iBAQ value. This is done by dividing the iBAQ value by the sum of the iBAQ values ​​of all detected proteins in the corresponding sample. Then, to make this FOT size easier to read and write, the FOT value is multiplied by 105 to obtain the iFOT value.

[0099] Proteomic sequencing of 101 HCC specimens obtained from the Department of General Surgery, Huashan Hospital, Fudan University, was performed by the National Institute of Metrology (China) as described above. Sample quality control and proteomic expression data are attached. Figure 1 and attached Figure 2 As shown, where Figure 1The following are the D15 and D17 mass spectra of some samples in this invention. The protein identification results for each sample are attached. Figure 2 As shown, in Figure 2 In this study, the number of proteins identified in each sample exceeded 4,000, with most samples having between 4,500 and 5,000.

[0100] Multiplex immunofluorescence (mIF) assay

[0101] After dewaxing and dehydration of paraffin sections, slides were immersed in citrate antigen retrieval buffer and boiling water for antigen retrieval. Before rinsing with phosphate-buffered saline (PBS), the samples were incubated with 3% H2O2 at room temperature, the solution was blocked with 5% FBS at 37°C for 30 minutes, and then incubated with primary antibody overnight at 4°C. The samples were rinsed with PBS, incubated with Cy5 / Sp Green-labeled secondary antibody at 37°C, and then rinsed with PBS again. The solution was then incubated with Hoechst at room temperature for 15 minutes, and the previous primary / secondary antibody was removed by microwave treatment. Continuous staining was repeated.

[0102] Statistical analysis

[0103] Kaplan-Meier analysis and log-rank test were performed to compare survival differences between the high / low-risk HCC groups. Test data were normally distributed; if so, a t-test was used, or a nonparametric Mann-Whitney U test was applied, unless otherwise specified. Correlation was assessed using the Pearson test for parametric data and the Spearman test for nonparametric data. All statistical analyses were performed using Prism (version 8, https: / / www.graphpad-prism.cn / ) and R statistical software (version 4.1.3, https: / / www.r-project.org / ). A p-value < 0.05 indicated statistical significance (*P < 0.05; **P < 0.01 and ***P < 0.001).

[0104] Experimental results

[0105] m in HCC 6 Identification of A-related gene characteristics

[0106] The TCGA database, containing detailed clinical information on 33 types of cancer, is considered a milestone in the cancer genomics initiative. It is used to identify m... 6Based on relevant characteristics, all raw gene expression data of 374 HCC patients were downloaded from the TCGA database. After carefully verifying the clinical information of each sample, 31 patients were excluded due to incomplete clinical data or follow-up time of less than 30 days. Ultimately, 343 patients with reliable transcriptional data and detailed clinical information were included in the study. Further analysis of 15 mature m... 6 The A-regulated genes (METTL3, METTL14, WTAP, ZC3H13, RBM15, YTHDC1, YTHDF1, YTHDF2, YTHDF3, IGFBP1, IGFBP2, IGFBP3, RBMX, FTO, and ALKBH5) were studied and analyzed. The entire analysis process is summarized in the appendix. Figure 3 The flowchart in the document.

[0107] Using univariate Cox regression analysis of gene expression levels, 15m 6 Eleven of the A-modifiers were significantly associated with HCC prognosis, while METTL14, IGFBP1 / 2, and ALKBH5 did not reach statistical significance, as shown in the attached table. Figure 4 As shown in A. According to 11m 6 The consensus algorithm of regulator A shows that when the x-axis is k=3, the slope of the scree plot drops sharply, indicating that patients can be divided into three categories to obtain better clustering representation, as shown in the attached figure. Figure 4 As shown in Figure B. GO and KEGG analyses show that each cluster exhibits distinct characteristics, as shown in the appendix. Figure 5 As shown. Cluster A is closely related to RNA processing and splicing. Figure 5 (Top left image). Cluster B is involved in histone modification and ribosome assembly. Figure 5 (Top right image). Cluster C is mainly rich in extracellular structures, collagen matrix, and integrin binding ( Figure 5 (Lower left figure). There was 85 overlapping differentially expressed genes across all three clusters, as shown in Table 4 below. Of these, 68 differentially expressed genes were significantly associated with HCC prognosis, as shown in the appendix. Figure 6 The left-middle figure and Table 5 below are shown. Given that protein-coding genes are the main executors of biological functions, this invention performed proteomics analysis using mass spectrometry (MS) on 101 HCC samples. A total of 7059 protein expression levels were successfully detected by MS. Subsequently, 68 prognostic-related genes were interpolated with the genes detected by MS, and finally 20 protein-coding genes were defined as m... 6 A. Relevant features are provided for further analysis, as shown in the attached document. Figure 6 As shown in the middle right figure, their coefficients are shown in Table 6 below.

[0108] Table 4. m in hepatocellular carcinoma 6 A modifies the DEGs of clusters A, B, and C.

[0109] FAM117B HDAC2 DZIP1L FANCE T×NDC12 BICD1 INTS8 UQCRH KDM1A CRLS1 SLC44A3 PABPC4 GYS1 DCUN1D5 MARCKSL1 MACROD2 AC104066.2 RAB11A CFHR3 GAS2L3 MAPRE1 CEP126 GLYATL1 AC022211.1 CIDEC CFAP298 PIGS RAVER2 GYS2 NAP1L1 KLRB1 XPNPEP1 H1-3 BMT2 PDE7A LINC01121 SFT2D1 AP000892.4 STX6 FAM86DP SNRNP40 CCN5 CDK19 RIBC2 BLMH AP002990.1 SOCS3 UBA2 C3P1 CDCA7 SNHG3 TMEM234 Y8X1 RALGPS2-AS1 SCLT1 SNORA15B-1 PAQR8 MSI2 LBX2-AS1 RDH16 IYD ARHGEF2-AS2 PROM1 LHFPL3-AS2 HMGN1 B9D2 HDDC2 AC114760.2 PTMAP4 MASP2 STXBP4 LINC01018 UTP11 CYP26B1 HAO2 AC084824.1 ACTG1P25

[0110] Table 5 m 6 Prognostic analysis of A-modified clusters A, B, and C DEG

[0111] IGFBP3 1.202152 1.066211 1.355424 0.002638 HDDC2 1.468087 1.164597 1.850666 0.001156 FAM117B 1.314641 1.090386 1.585017 0.004147 CYP2681 1.399713 1.220531 1.6052 1.5E-06 KDM1A 1.956574 1.518727 2.520651 2.07E-07 CFHR5 0.916698 0.858619 0.978706 0.009201 AC104066 2.797414 1.85927 4.208924 7.99E-07 TXNDC12 2.345491 1.650229 3.333675 2.01E-06 CIDEC 1.119483 1.01631 1.23313 0.022142 GYS1 1.40372 1.126197 1.749631 0.002549 SNRNP40 1.826473 1.397778 2.386646 1.02E-05 MAPRE1 1.443094 1.180244 1.764484 0.00035 C3P1 0.890395 0.82006 0.966762 0.005691 GYS2 0.884051 0.816449 0.957251 0.002394 PAQR8 1.172628 1.000384 1.374529 0.049449 SFT2D1 1.675212 1.327143 2.11457 1.41E-05 LHFPL3-A 1.203727 1.046051 1.38517 0.00964 BLMH 1.20832 1.042625 1.400346 0.011915 STXBP4 1.490245 1.126437 1.971554 0.005211 YBX1 2.334344 1.799421 3.028287 1.73E-10 HDAC2 1.93874 1.515903 2.479519 1.33E-07 IYD 0.793324 0.683902 0.920253 0.002232 RAB11A 1.363697 1.00219 1.855607 0.048395 BICD1 1.951282 1.470895 2.588561 3.55E-06 CFAP298 1.517849 1.106749 2.081652 0.009617 DCUN1D5 1.888181 1.473202 2.420053 5.17E-07 BMT2 1.649991 1.270008 2.143665 0.000177 CEP126 2.008702 1.379218 2.925487 0.000277 CDCA7 1.267975 1.119948 1.435566 0.000178 NAP1L1 1.541312 1.265808 1.87678 1.66E-05 MSI2 1.263862 1.017455 1.569945 0.034315 AP002990 1.650051 1.138396 2.391671 0.008184 HMGN1 1.298436 1.031847 1.633902 0.025924 ARHGEF2 1.978118 1.312059 2.982296 0.001128 LINC01018 0.901427 0.836688 0.971175 0.006349 PTMAP4 2.271206 1.706525 3.022738 1.86E-08 DZIP1L 1.523597 1.237509 1.875823 7.24E-05 AC084824 1.582702 1.122336 2.231903 0.008843 CFHR3 0.858593 0.79341 0.929131 0.000154 RPS4XP1 1.839692 1.02241 3.310282 0.041961 PIGS 1.484926 1.230384 1.792127 3.77E-05 INTS8 1.743036 1.353881 2.244049 1.63E-05 PDE7A 1.413364 1.158933 1.723652 0.000634 MARCKSL 1.37243 1.195007 1.576195 7.38E-06 CDK19 1.404927 1.147375 1.720292 0.001 GLYATL1 0.844672 0.7687 0.928153 0.000447 SNHG3 1.459384 1.24702 1.707913 2.46E-06 KLRB1 0.795015 0.672439 0.939936 0.007253 UTP11 3.015607 2.176226 4.178741 3.31E-11 STX6 1.667861 1.306571 2.129055 4.01E-05 TCP1 1.670928 1.28745 2.168628 0.000114 SCLT1 1.840914 1.287927 2.631333 0.000813 FANCE 1.583806 1.307757 1.918124 2.53E-06 MASP2 0.889218 0.822008 0.961923 0.00341 PABPC4 1.770198 1.355257 2.312183 2.78E-05 KDELR3 1.228993 1.086776 1.389821 0.001015 GAS2L3 1.688699 1.410232 2.022152 1.21E-08 UQCRH 1.92623 1.493764 2.4839 4.34E-07 RAVER2 1.215942 1.030797 1.434342 0.020348 MACROD2 1.282311 1.008208 1.630934 0.042705 LINC0112 2.441549 1.285963 4.635562 0.006356 AC022211 3.104223 1.505991 6.398578 0.002144 RIBC2 1.437506 1.225608 1.686039 8.19E-06 XPNPEP1 1.743086 1.268546 2.395143 0.00061 TMEM234 1.902631 1.454 2.489687 2.76E-06 FAM86DP 2.005571 1.442969 2.787526 3.43E-05 RDH16 0.879342 0.816205 0.947363 0.000719 UBA2 1.477286 1.173808 1.859225 0.000881

[0112] Table 6 20 m 6 Coefficient of A-related genes

[0113] IGFBP3 0.505342 0.016672 0.522013718 HDAC2 0.867325 0.119152 0.986476171 HMGN1 0.73316 0.056543 0.789703725 TCP1 0.676837 0.092023 0.768860137 PABPC4 0.728055 -0.01474 0.71331339 RDH16 -0.58298 0.522821 -0.060163531 HDDC2 0.736668 -0.10174 0.634928786 CFHR5 -0.41475 0.519119 0.104369803 GYS1 0.73258 0.032457 0.765037486 MAPRE1 0.872042 0.062224 0.934266148 GYS2 -0.61596 0.641265 0.025300886 BLMH 0.685407 -0.10318 0.58222723 YBX1 0.692319 0.071267 0.763585884 NAP1L1 0.845052 -0.10595 0.739100649 INTS8 0.770094 0.042677 0.812771311 MARCKSL1 0.746138 -0.26199 0.484149414 STX6 0.830609 0.091894 0.922502372 MASP2 -0.67542 0.218138 -0.457282071 UQCRH 0.58297 -0.1771 0.405866493 XPNPEP1 0.729938 0.135524 0.865462231

[0114] Comparison of HCC tissue and adjacent normal tissue over 20m 6 Expression levels of A-related genes. Most were upregulated, while five were downregulated, as shown in the attached table. Figure 7 As shown in Figure A, red boxes represent tumors, and blue boxes represent normal tissue. Loss and addition copy number variations (CNVs) are shown in the appendix. Figure 7 As shown in Figure B, CNV gains and deletions are represented in red and green, respectively. Furthermore, online SRAMP analysis revealed a large number of highly dense m-elements within these 20 genes. 6 Cluster A, as shown in the appendix Figure 8-12 As shown, this indicates that these genes may be m 6 Potential targets modified by A.

[0115] In summary, this invention establishes a m-system composed of 20 protein-coding genes (IGFBP3, HDAC2, HMGN1, TCP1, PABPC4, RDH16, HDDC2, CFHR5, GYS1, MAPRE1, GYS2, BLMH, YBX1, NAP1L1, INTS8, MARCKSL1, STX6, MASP2, UQCRH, and XPNPEP1). 6 A related feature is significantly associated with the prognosis of HCC.

[0116] m 6 Clinical significance of A-related features in HCC

[0117] m 6 The results of a study on the clinical significance of A-related features in HCC are attached. Figure 13 As shown, where A is based on m from the TCGA HCC sample. 6 m of A related features 6 AScore distribution; red represents high m 6 AScore (high risk) group, blue represents low m 6 AScore (low-risk) group. B is m 6The association between AScore and survival status; red represents deceased patients, and blue represents surviving patients. The box plot in C shows the m... 6 The relationship between AScore and clinicopathological stage; blue represents stage I-II patients, and red represents stage III-IV patients. D indicates low m 6 AScore group (blue image) and height m 6 Kaplan-Meier survival curves for the AScore group (red plot). E represents m for 1 (red plot), 3 (blue plot), and 5 (green plot). 6 Receiver operating characteristic (ROC) curve and area under Ascore ROC. F is m 6 AScore calibration curves for 1 (red), 3 (blue), and 5 (green) years.

[0118] According to m 6 A related features, according to the method (m in HCC) 6 The description in the construction of the relevant signature is to calculate m for each patient. 6 AScore. Based on the optimal cut-off point of the survival curve, 343 HCC patients in TCGA were divided into low-risk and high-risk groups. (See attached...) Figure 13 As shown in A and B. With m 6 An increase in AScore is associated with a gradually increasing risk of death. Meanwhile, high m... 6 AScore is significantly correlated with the stage of late-stage HCC, as shown in the attached figure. Figure 13 As shown in C. Therefore, it has a height of m 6 Patients with AScore were stratified into the high-risk HCC group.

[0119] Then, Kaplan-Meier analysis of overall survival (OS) showed that high-risk individuals typically had shorter survival times, as shown in the attached figure. Figure 13 As shown in Figure D. Compared with low-risk patients, the 1-year, 3-year, and 5-year survival rates of high-risk patients were 62.2% vs. 91.7%, 40.1% vs. 70.7%, and 21.9% vs. 57.0%, respectively. Receiver operating characteristic (ROC) curve analysis showed that m 6 AScore is a useful prognostic indicator for predicting overall survival (OS) in HCC patients because the area under the ROC curve (AUC) reaches 0.746, 0.669, and 0.628 at 1, 3, and 5 years, respectively, as shown in the attached figure. Figure 13 As shown in E. Furthermore, m 6 The AScore model showed good calibration; the calibration curves for 1-year, 3-year, and 5-year survival were close to the ideal curve, as shown in the attached figure. Figure 13 As shown in F.

[0120] In summary, these findings indicate that m 6A-related coding gene features have a strong ability to predict the overall survival of HCC patients.

[0121] m in independent HCC queue 6 Proteomic validation of A-related gene features

[0122] In order to evaluate m 6 The reliability of A-related protein-encoding gene markers in prognostic assessment was further validated using an independent cohort of 101 HCC specimens. The results are attached. Figure 14 As shown, where A is the value of m based on independent HCC samples. 6 m of A related features 6 AScore distribution; red represents high m 6 AScore (high risk) group, blue represents low m 6 AScore (low-risk) group. B is m 6 The association between AScore and survival status; red represents deceased patients, and blue represents surviving patients. The box plot in C shows the m... 6 The relationship between AScore and clinicopathological stage; blue represents stage I-II patients, and red represents stage III-IV patients. D indicates low m 6 AScore group (blue image) and height m 6 Kaplan-Meier survival curves for the AScore group (red plot). E represents m for 1 (red plot), 3 (blue plot), and 5 (green plot). 6 The area under the receiver operating characteristic (ROC) curve in AScore ROC. F represents the m for 1 (red graph), 3 (blue graph), and 5 (green graph). 6 AScore calibration curve.

[0123] The basic information of the HCC queue verified in this invention is shown in Table 1 above. 6 Protein levels of the A-related characteristic were determined by mass spectrometry (MS), and m was calculated. 6 AScore categorizes HCCs into high-risk and low-risk groups as described above.

[0124] With appendix Figure 13 The results are consistent with those in the previous section, with high m 6 The AScore group had a higher risk of death (see attached). Figure 14 (as shown in A and B) and later clinical stages (as shown in the appendix) Figure 14 (As shown in C). Furthermore, the survival time of the high-risk group was significantly shorter (as shown in Appendix C). Figure 14 (As shown in Figure D). Compared with low-risk patients, the 1-year and 3-year survival rates of high-risk patients were 87.5% vs. 92.2% and 44.1% vs. 71.8%, respectively. ROC curve analysis further confirmed m6 The accuracy and sensitivity of the relevant features are improved, as the AUC reaches 0.722 and 0.604 at 1 year and 3 years, respectively (see attached). Figure 14 (As shown in Figure E). The 1-year and 3-year calibration curves are very close to the ideal curve, showing satisfactory calibration results (as shown in the attached figure). Figure 14 (As shown in F).

[0125] Therefore, these verified results confirm m 6 The reliability and robustness of the A-related characteristics make it easy to stratify high-risk HCC patients and assess overall survival. 6 A related characteristic may help in providing better clinical management for high-risk HCC at an earlier stage.

[0126] m 6 Functional characterization of A-related gene features and potential treatment strategies for high-risk HCC populations

[0127] In order to utilize m 6 To explore the potential mechanisms underlying the A-related features, enrichment analysis was performed on GO and KEGG in this invention to investigate possible biological functions. The results are attached. Figure 15 As shown, A represents the gene set enrichment analysis from the Gene Ontology (GO) model; the figure below shows the low m 6 The gene set enrichment of the AScore group is shown in the figure above, which illustrates the high m 6 Gene set enrichment in the AScore group. B represents gene set enrichment analysis of the KEGG pathway; the figure below shows low m 6 The gene set enrichment of the AScore group is shown in the figure above, which illustrates the high m 6 The AScore group showed gene set enrichment. CK was high or low m 6 Estimated logIC values ​​of several chemotherapy drugs in the AScore group 50 Box plot.

[0128] As attached Figure 15 As shown in Figures A and B, in the high-risk HCC group, abnormal DNA replication and recombination, epigenetic dysregulation, cell cycle disorders, and activation of oncogenic signaling pathways such as MAPK, mTOR, and VEGF were significantly increased. Meanwhile, amino acid and fatty acid metabolism disorders and cytochrome p450 pathway abnormalities were significantly increased in the low-risk HCC group.

[0129] Based on the enrichment analysis results, this invention aims to explore potential effective strategies for the treatment of high-risk HCC. Several FDA-approved inhibitors and chemotherapeutic agents that target the cell cycle (CGP.082996), inhibit DNA replication and synthesis (etoposide and mitomycin C), and inactivate the MAPK signaling pathway (PLK1, AKT, and p38 inhibitors) were selected for further analysis. The half-maximal inhibitory concentration (IC50) of these inhibitors was then evaluated using the pRRophetic platform in low-risk and high-risk subgroups. 50 pRRheetic is an independent, licensed public resource that predicts inhibitor sensitivity based on gene expression microarray data. (See attached...) Figure 15 As shown in the Chinese study, high-risk HCC patients are more sensitive to the above-mentioned inhibitors because low drug concentrations can effectively inhibit cell proliferation.

[0130] These findings suggest that low-risk and high-risk individuals may have different carcinogenic mechanisms that promote the development of HCC. High-risk HCC patients may respond better to CDK and MAPK inhibitors, as well as chemotherapy drugs such as etoposide and mitomycin C. 6 A related characteristic helps in selecting an appropriate treatment strategy.

[0131] m 6 A related characteristic is associated with tumor mutational burden and immune checkpoint blockade therapy.

[0132] Tumor mutation burden (TMB) is an important indicator for assessing cancer progression and the efficacy of immunotherapy. 6 The prediction results of tumor somatic mutations and immunotherapy sensitivity related to the A-related features are shown in the appendix. Figure 16 As shown, where A is the low m 6 Gene mutation frequency in the AScore group. B represents high m. 6 Gene mutation frequency in the AScore group. C represents the Kaplan-Meier survival curves of the four groups divided by tumor mutation burden (TMB) and m. 6 AScore. L-TMB is short for Low TMB, H-TMB is short for High TMB; L-m6AScore is lowm 6 AScore is short for Hm 6 AScore is high-m 6 AScore is an abbreviation for TIDE score. D represents the analysis of predicting immunotherapy response in four groups using the TIDE score; a lower TIDE score indicates greater sensitivity to immunotherapy. The box plot in E shows high m... 6 AScore group and low m 6 Differential expression of immune checkpoint genes (PD-1, PD-L1, CTLA4, TIM3, LAG3) among AScore groups.

[0133] As attached Figure 16 As shown in the waterfall plots A and B, the gene mutation profiles of the low-risk and high-risk groups are significantly different. In low-risk patients, CTNNB1 mutation is the most common event (see attached figure). Figure 16 (As shown in Figure A); while in the high-risk group, the frequency of TP53 mutations was present in more than half of the patients (as shown in the attached figure). Figure 16 (As shown in B). Then, based on high- / low-TMB and m 6 The AScore combination divides HCC patients into four subgroups, as shown in the appendix. Figure 16 As shown in C, L-TMB+Lm 6 AScore, L-TMB+Hm 6 AScore, H-TMB+Lm 6 AScore and H-TMB+Hm 6 AScore. Ultimately, L-TMB+Hm was discovered. 6 The AScore group had the worst prognosis, while the L-TMB+Lm group had the worst prognosis. 6 The AScore group had the best prognosis.

[0134] Immunotherapy using immune checkpoint inhibitors has achieved encouraging clinical results in various cancers, including HCC. TIDE is an algorithm used to predict response to immunotherapy; a high TIDE score generally indicates a poor response. This invention also found that, regardless of TMB, a high m 6 HCC patients with AScore scores tend to have lower TIDE scores (see attached). Figure 16 As shown in Figure D), this suggests that high-risk individuals may benefit from immunotherapy. Furthermore, we analyzed each TMB combination with m 6 The AScore subgroup contained documented expression levels of immune checkpoint molecules, including CTLA4, PD-1, PD-L1, LAG3, and TIM3. (See attached image.) Figure 16 As shown in Figure E, most of these immune checkpoint molecules are in high m 6 Higher expression levels were observed in the AScore subgroup, suggesting that high-risk HCC may suppress T cell activation and evade anti-tumor immunity. On the other hand, this implies that high-risk HCC patients may be sensitive to immune checkpoint therapies, such as antibodies targeting PD-1, CTLA-4, TIM3, and LAG3.

[0135] In summary, m 6 AScore combined with TMB can stratify the HCC subgroup with the worst prognosis. High-risk HCC patients are prone to immune escape through upregulation of immune checkpoint molecule expression and may benefit from immune checkpoint therapy.

[0136] m 6 Feature A can characterize the tumor immune microenvironment.

[0137] To systematically describe the state of the tumor immune microenvironment, this invention uses ssGSEA analysis to compare the infiltration percentages of various immune cells, and the results are attached. Figure 17 As shown, the box plot in section A displays the m in the HCC. 6 AScore and ssGSEA analysis of immune infiltration level; red indicates high m 6 AScore group, blue represents low m 6 AScore group. The relevant heatmap in section B shows m. 6 The relationship between AScore, expression of 20 genes, and immune cell infiltration. CF box plots show the expression of Treg (regulatory T cells) and M2 macrophage-related genes in high m 6 AScore group and low m 6 Expressions between AScore groups.

[0138] Notably, immunosuppressive cells such as M2 polarized macrophages and T regulatory cells (Tregs) significantly accumulated in the high-risk group, as shown in the appendix. Figure 17 As shown in Figure A. m 6 Each of the 20 genes in trait A contributes to immune cell infiltration in its own way (see attached figure). Figure 17 (See Figure B). Foxp3 is a well-known important transcription factor for Treg cells; CD163, IRF4, and VSIG4 are markers of M2 macrophages. Therefore, the expression levels of these genes were analyzed in the TCGA dataset. The high-risk HCC group showed significant upregulation of Foxp3, CD163, IRF4, and VSIG4 (see Appendix B). Figure 17 (as shown in CF).

[0139] To confirm and validate these findings, this invention performed multiplex immunofluorescence staining to visualize the tumor microenvironment characteristics in a subset of HCC cohorts, the results of which are attached. Figure 18 As shown, where A is the value at low m 6 In the Ascore group, multiplex immunofluorescence staining results for DAPI, CD163, and Foxp3 were shown; each stain was displayed separately and then combined (scale bars are shown in the figure). B represents the staining at a height of m. 6 In the Ascore group, multiplex immunofluorescence staining results for DAPI, CD163, and Foxp3 are shown; each stain is displayed separately and then merged (scale bars are shown in the figure). (See attached...) Figure 18 As can be seen, in low-risk HCC samples, only a small number of M2-macrophages (green) and Tregs (purple) were observed (see attached image). Figure 18(As shown in Figure A). Consistently, in high-risk HCC samples, there was significant infiltration of M2 macrophages and Tregs in the tumor microenvironment (as shown in the attached figure). Figure 18 (As shown in B).

[0140] Therefore, these results indicate that m 6 Feature A can characterize the tumor immune microenvironment. High-risk HCC individuals may develop an immunosuppressive microenvironment with a large number of M2 macrophages and Treg infiltration.

[0141] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. m used to assess overall survival in HCC and characterize the tumor immune microenvironment in HCC 6 A-related gene combination, characterized by: The m 6 A related gene combination consists of 20 protein-coding genes; the 20 protein-coding genes include IGFBP3, HDAC2, HMGN1, TCP1, PABPC4, RDH16, HDDC2, CFHR5, GYS1, MAPRE1, GYS2, BLMH, YBX1, NAP1L1, INTS8, MARCKSL1, STX6, MASP2, UQCRH, and XPNPEP1.

Citation Information

Patent Citations

  • Immune gene prognosis model for predicting hepatocellular carcinoma tumor immune infiltration and postoperative survival time

    CN112011616A

  • Liver cancer prognosis risk model based on m6A characteristic gene and application thereof

    CN114292917A