Method for typing polymorphic glioblastoma
Through single-cell omics analysis and genomic pathway enrichment, GBM was classified into proliferative, neurosynaptic, metabolic, and immunomodulatory types, which solved the problem of inaccurate classification of glioblastoma multiforme and provided GBM treatment strategies for different populations.
Patent Information
- Application Number
- CN202511140359.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-12-09
AI Technical Summary
Current classification methods for glioblastoma multiforme are inaccurate, lack population-specificity, and cannot effectively guide clinical treatment.
By acquiring single-cell omics analysis data from multiple tumor cell samples, performing whole-exome sequencing and RNA sequencing, and utilizing clustering computation and genomic pathway enrichment analysis, glioblastoma multiforme was classified into proliferative, neurosynaptic, metabolic, and immunomodulatory types, and identified through gene and protein biomarkers.
It provides a more accurate classification method for glioblastoma multiforme, reveals the genetic drivers and clinical characteristics of different human ancestral tumors, and enhances the understanding of GBM occurrence and treatment strategies in different populations.
Smart Images

Figure CN121096451A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of bioinformatics technology, and more specifically, to a method for classifying glioblastoma multiforme. Background Technology
[0002] Glioblastoma multiforme (GBM) is a common primary malignant tumor of the central nervous system and a leading cause of death related to primary brain cancer. Surgery, chemotherapy, and radiotherapy remain the standard treatments. Although research on immunotherapy for GBM is increasing, no major breakthroughs have been achieved to date. Most characterizations of GBM to date have focused on the predominantly single-ancestor population (EUR-GBM), and whether these findings are applicable to GBM patients with other ancestral origins remains to be explored.
[0003] The transcriptome-based classification of GBM was initially proposed but later revised into three molecular subtypes. However, these groupings are not routinely used to guide clinical treatment management, and their relevance to other ancestral GBM remains unclear. Summary of the Invention
[0004] This application addresses the shortcomings of existing methods by proposing a classification method for glioblastoma multiforme, which solves the technical problems of inaccurate classification and lack of population diversity in related technologies.
[0005] This application provides a method for classifying glioblastoma multiforme, which includes the following steps: Acquire single-cell omics analysis data from multiple tumor cell samples; Clustering calculations were performed on the whole exon sequencing data and / or RNA sequencing data based on the single-cell omics analysis data; Genomic pathway enrichment analysis was performed on each cluster based on the clustering calculation results; Based on the results of the genomic pathway enrichment analysis, glioblastoma multiforme was classified into: proliferative, neurosynaptic, metabolic, and immunomodulatory types.
[0006] Alternatively, after acquiring single-cell omics analysis data from multiple tumor cell samples: Different human ancestral origins can be distinguished based on the single-cell omics analysis data; Compare somatic copy number variations and / or somatic single nucleotide variations among different human ancestral origins; The population-specific differences in glioblastoma multiforme were identified based on the differential data.
[0007] Furthermore, the comparison of somatic copy number variation and / or somatic single nucleotide variation among different human ancestral sources includes: The typing results based on the genomic pathway enrichment analysis are mapped to other typing criteria.
[0008] Alternatively, the other classification criteria may include at least one of the following: traditional pathological classification based on histological features, molecular classification based on key molecular markers, and TCGA molecular classification based on gene expression profiles.
[0009] Furthermore, the clustering calculation includes: Perform NMF clustering or k-means clustering on whole exome sequencing data and / or RNA sequencing data; The number of subtypes is determined based on the number of clusters after clustering.
[0010] Optionally, the method may also include screening for genotyping markers based on the results of the genomic pathway enrichment analysis, the genotyping markers including gene markers and / or protein markers.
[0011] Further, alternatively, the genotyping markers include: Gene markers used to identify the proliferative type MYB, MYCN, BRIP1, CCND2, PDGFRA, PIK3CA, MAPK1, AKT3 At least one of the following, protein markers DLL3 and / or Ki67; Genetic markers used to identify the neural synaptic type WIF1, ATP2B3, GRM3, GRIN2A, KIAA1598, FLT3SIX1, FGFR2 At least one of the following, protein markers: GRMM3 and / or FGFR2; Gene markers used to identify the metabolites ASPSCR1, NTHL1, MADIL1, TERT, GADD45G, ACSF2, ACADVL, ACSF3 At least one of them, protein marker P16; Gene markers used to identify the immunomodulatory type REL, TGFBR2, COL3A1, COL1A1, SELE, CD274, PDCD1LG2, CCR4L, IL21R, IL7R, FAS At least one of the following, and at least one of the following protein markers: CD44, C1R, and Ki67.
[0012] Alternatively, the single-cell omics analysis data of the plurality of tumor cell samples exclude data from IDH mutant samples.
[0013] Furthermore, the single-cell omics analysis data includes genomic, transcriptomic, epigenomic, and / or proteomic data of individual tumor cells.
[0014] The beneficial technical effects of the technical solutions provided in this application include: Tumor samples from 367 EAS-GBM patients were analyzed using whole-exome sequencing (WES, n = 183) and RNA sequencing (RNA-Seq, n = 258). These data were used to define the molecular categories of EAS-GBM and compare them and their genetic drivers with known EUR-GBM data. The EAS-GBM subgroup includes an immune-infiltrating subgroup (immunomodulatory type) with potential clinical applications. This application provides a comprehensive profile of EAS-GBM, enhancing our understanding of tumor development, clinical stratification, and treatment strategies across different ancestral origins.
[0015] Additional aspects and advantages of this application will be set forth in part in the description which follows, and will become apparent from the description or may be learned by practice of this application. Attached Figure Description
[0016] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 Differences between EAS-GBM and EUR-GBM used in this application in terms of age, sex, and treatment status (including adjuvant therapy); bar color indicates the frequency of each change in each cohort; gray bars indicate unknown change status; Figure 2 Correlation analysis of PRS status and molecular subtype in the sample data used in this application; Figure 3 The t-SNE plot of the EAS-GBM samples is colored according to the groups defined by NMF; Figure 4 To display the distances between gene expression combinations between EAS-GBM subtypes in a phylogenetic tree; Figure 5 The EAS-GBM subtype tags defined in this application are (PL = proliferative, n=50, NS = synaptic, n=37, MB = metabolitic, n=55, IM = immunomodulatory, n=116). Figure 6 The significance of ssGSEA scores for signaling pathways (I) and GO items (J) among EAS-GBM subtypes; colors indicate the significance of enrichment (red) or deletion (blue); Figure 7The graph shows the enrichment or deletion of events at the chromosome arm level between EAS-GBM and EUR-GBM; the X-axis represents the log2 (plurality ratio) showing the magnitude of the difference in prevalence between the EAS-GBM (n=263) and EUR-GBM cohorts (n=330); the Y-axis represents the significance of enrichment as shown by -log10 (q value). Figure 8 This shows the difference in SCNA between EAS (n=41) and EUR (n=109) GBM in the GLASS cohort; the X-axis represents log2 (variance ratio) showing the magnitude of the difference in prevalence between EAS (n=41) and EUR (n=109); the Y-axis represents the enrichment significance shown by -log10 (q value); Figure 9 Enrichment and expression network analysis of PL-GBM subtype pathways; Figure 10 Enrichment and expression network analysis of NS-GBM subtype pathways; Figure 11 Enrichment and expression network analysis of MB-GBM subtype pathways; Figure 12 For IM-GBM subtype pathway enrichment and expression network analysis; Figure 13 Significance of overlap between existing GBM subtype labels and EAS-GBM subtype; positive correlations are shown in red, and significant negative correlations are shown in blue; color intensity indicates significance level, and values q>0.05 are kept in white; Figure 14 The overlap between the EAS-GBM subtype label in this application and existing classifications; Figure 15 The overlap of marker genes between the EAS-GBM subtype tag in this application and existing classifications; Figure 16 The expression differences of marker genes between EAS-GBM subtypes are shown; colors represent relative normalized expression, with red indicating overexpression and blue indicating low expression. The size of the squares indicates the significance of the expression differences calculated by the Wilcoxon rank-sum test; solid squares represent q values q < 0.01, while diamonds represent q > 0.1. Figure 17 For the analysis of the significance of expression differences of previously classified marker genes between EAS-GBM subtypes; colors represent relative normalized expression, red indicates overexpression, and blue indicates low expression; the size of the squares represents the significance of expression differences calculated by the Wilcoxon rank-sum test; the size of solid squares represents q values q < 0.01, while diamonds represent q > 0.1; Figure 18 Summarize the marker genes for each molecular subtype in EAS-GBM; Figure 19 SCNA enrichment / deletion analysis at the chromosome arm level for each EAS-GBM subtype based on Fisher's exact test: The X-axis represents log2 (fold change ratio) of the difference between the subtype sample and the other EAS-GBM samples; the Y-axis represents -log10 (q value) of the enrichment significance. Figure 20 Genes with significantly recurring single nucleotide variants (SNVs) in EAS-GBM (q<0.1) are shown; the top row shows the mutation burden for each patient; the bottom row shows the mutation type for each gene in all patients, with the right column showing the significance of each mutation; SNV frequency is shown in the left column; the bottom row shows the ratio of tumor reads to total reads in each sample. Figure 21 IHC staining of immune cell markers CD4+ (helper T cells), CD8+ (cytotoxic T cells), CD68+ (microglia), CD163+ (M2 microglia) and EAS subtype markers DLL3 (PL), GRM3 (NS), FGFR2 (NS), p16 (MB), CD44 (IM), and C1R (IM) in four EAS-GBM subtype samples; Figure 22 FGFR2 and GRM are significantly overexpressed in NS tumors; Figure 23 The levels of marker proteins between PL-GBM and IM-GBM were displayed; Figure 24 The levels of marker proteins between MB-GBM and NS-GBM were displayed; Figure 25 For pathologists to perform quantitative analysis and statistics on the staining intensity of tumors; Figure 26 IHC staining of B cell markers in each EAS-GBM subtype is shown. Detailed Implementation
[0017] The embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions of the embodiments of this application.
[0018] Those skilled in the art will understand that, unless specifically stated otherwise, the terms "described" and "the" as used herein may also include plural forms. It should be further understood that the term "comprising" as used in the specification of this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude other features, information, data, steps, operations, elements, components, and / or combinations thereof supported by the art. The term "and / or" as used herein refers to at least one of the items defined by the term; for example, "A and / or B" can be implemented as "A," or as "B," or as "A and B."
[0019] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0020] First, let's introduce and explain several terms used in this application: EUR-GBM patients refer to patient data used to construct known GBM subtypes, primarily derived from first human ancestry.
[0021] EAS-GBM patients refer to the patient data used to construct the GBM subtyping in this application, which are primarily derived from second human ancestry.
[0022] Human ancestry: Genetic research distinguishes different human ancestry by analyzing specific genetic markers in DNA. It mainly relies on features such as the Y chromosome, mitochondrial DNA, and single nucleotide polymorphisms (SNPs), combined with genetic geography and population genetics methods, to reveal human migration and evolutionary pathways.
[0023] Single-cell omics refers to research methods that use high-throughput techniques to analyze information such as the genome, transcriptome, epigenome, and proteome of a single cell. It is often used to reveal cellular heterogeneity, developmental trajectories, and disease mechanisms.
[0024] Non-negative Matrix Factorization (NMF) is an unsupervised learning algorithm focused on dimensionality reduction and feature extraction. Its core idea is to decompose a non-negative matrix into the product of two low-dimensional non-negative matrices, thereby preserving the main features of the data and removing redundant information.
[0025] K-means is a partition-based clustering algorithm that aims to divide data into K clusters such that the sum of the squared distances from samples within a cluster to the cluster center is minimized.
[0026] Principal Component Analysis (PCA) is a classic unsupervised dimensionality reduction and feature extraction method that projects high-dimensional data into a low-dimensional space through linear transformation while preserving the main information (variance) of the data.
[0027] The classification method for glioblastoma multiforme in this application specifically includes the following steps: S1, acquire single-cell omics analysis data from multiple tumor cell samples.
[0028] Patient characteristics: The study included 367 tumors from adult GBM patients in the second population. This included 109 tumors from a multicenter study in country A that underwent de novo whole-exome sequencing, as well as previously collected tumors. RNA sequencing was performed on GBM samples from 258 patients and 20 normal brain tissue control samples, with an average sequencing depth of 46.5 million reads (range 28.1 million–92.3 million reads). A subset of 183 tumors from patients aged 23 years and older was analyzed using whole-exome sequencing (WES), with an average coverage depth of 149x (range 91–195x). Forty-six patients had both WES and RNA-seq data.
[0029] Clinical and molecular data included primary / relapsed / secondary (PRS) type, sex, radiotherapy status, chemotherapy status, IDH mutation status, 1p19q co-deletion status, MGMT promoter methylation status, and age of onset for all recruited individuals, such as Figure 1 As shown, the age of EAS-GBM patients was significantly younger than that of EUR-GBM patients in the TCGA cohort (53 years vs. 61 years, bilateral Wilcoxon test). Among patients with a known treatment history, EUR-GBM patients received more radiotherapy (chi-square test). No significant difference was detected in the proportion of patients receiving chemotherapy between the two cohorts (chi-square test, p = 0.28). The age difference was consistent with previously published studies on EAS-GBM. No significant difference in sex ratio was detected between the two ancestral origins. The EAS-GBM sample was collected sequentially from multiple centers in country A. Using this large multi-omics dataset, we tested the applicability of previous findings, primarily based on EUR-GBM patients, to EAS-GBM patients.
[0030] Furthermore, additional single-cell omics analysis data of these tumor cell samples can be obtained, including but not limited to genomic, transcriptomic, epigenomic, and / or proteomic data of individual tumor cells in the samples, to facilitate more in-depth data mining.
[0031] S2, perform clustering calculations on whole exon sequencing data and / or RNA sequencing data based on the single-cell omics analysis data.
[0032] Transcriptome clustering using nonnegative matrix factorization: Raw RNA sequencing reads from 388 tumors and 20 normal samples from the EAS-GBM patient cohort, and 160 tumor samples from the EUR-GBM patient cohort, were aligned to a reference genome (UCSChg19, annotated with GENCODE v.37) using STAR (v.2.7.3a). Thirty RNA sequencing samples were removed from the cohort because PCA results indicated potential RNA degradation. Gene expression profiles were then quantified using RSEM (v.1.3.3). Expression counts were estimated using RSEM. TPMs were then calculated by dividing the read count by the kilobase length of each gene. For tumor sample clustering, the top 3000 most variable coding genes (based on median absolute bias) were selected from the EUR-GBM and EAS-GBM tumor transcriptome data, respectively. 2,771 overlapping genes were used for subsequent NMF clustering. To obtain robust clustering consensus, an unsupervised clustering method called Nonnegative Matrix Factorization (NMF) (NMF R-package v. 0.23.0) was used. First, 50 runs with random starting points were performed to determine the optimal rank parameter (number of clusters), followed by 300 runs using the optimal rank to obtain the final NMF clustering scheme. The rank was determined by evaluating the cophenotypic correlation coefficient (a measure of cluster stability) and the coefficient of variation (a measure of cluster repeatability obtained from multiple NMF runs). Following standard practice, the number of clusters with the largest slope among the two parameters was selected. Furthermore, profilometry was performed to evaluate the consistency of the proposed clustering scheme, evar (the explained variance of the NMF estimate of the target matrix), and residuals, i.e., the sum of squared residuals (RSS) between the target matrix and its NMF estimate, were calculated. The optimal number of clusters was selected at the first inflection point of the RSS, residual, and evar curves. Cluster significance was also evaluated using SigClust, and all class boundaries were proven to be statistically significant. In addition to enrichment of primary / untreated GBM samples in metabolic subtypes ( Figure 2 No significant difference was detected between the proportions of untreated / primary GBM samples and treated GBM samples across the clusters. Figure 1 As shown.
[0033] Depending on the specific circumstances, NMF clustering can be replaced by k-means clustering.
[0034] Principal component analysis and clustering: First, principal component analysis (PCA) was used to examine RNA degradation in these samples. Next, PCA was performed again, based on the top 3000 variant genes used in NMF clustering, to examine the correlation between the proposed molecular isotypes and batches. PCA was implemented using the built-in R function `prcomp`. Sample points were colored based on their NMF clusters and batches, while shapes were labeled according to PRS type. Furthermore, t-SNE clustering was performed using the `t-SNE` R package to examine grouping efficiency. The phylogenetic tree of the EAS cohort was analyzed using the `monocle` R package (v.0.2.3), and distance plots of expressed gene combinations between different molecular isotypes were drawn, such as... Figures 3-4 As shown. We deconvolve the expression of a large number of RNAs in EAS-GBM using single-cell datasets to infer the relative contribution of different published single-cell subtypes to the overall expression profile of our EAS-GBM tumor samples.
[0035] Detection of somatic copy number variations (SCNAs) in EAS and EUR-GBMs: Sequenza (v.3.0.0) was used to estimate copy number profiles (including allele-specific copy numbers), tumor purity, and ploidy for each patient. Default settings were used according to the Sequenza documentation. Significantly amplified and deleted regions were identified using GISTIC (v.2.0.23). The output segmentation file generated by Sequenza was used as input for GISTIC. GISTIC parameters were set to default values according to the manual. Chromosome arms were marked as “altered” in each cohort if GISTIC q < 0.1. GII was calculated as the percentage of the tumor genome showing a copy number deviation from the genomic median (all SCNAs inferred by Sequenza are integers). GII deletion was defined as the percentage of tumors with a significantly reduced copy number. Microsatellite instability (MSI) scores were then assessed using MSISensor by inputting aligned bam files.
[0036] SCNAs for EAS-GBM samples were computed from RNA-seq data using CaSpER. Using WES SCNA calls as the gold standard, we observed an accuracy of 89% for arm-level SCNA calls from RNAseq, for a subset of samples (n=46) that underwent both RNAseq and WES. Fisher's exact test was then performed to detect SCNA enrichment for GBM molecular subtypes. Comparisons were made with the TCGA cohort using WES SCNA calls, with a cutoff of + / - 0.5 for arm-level events in GISTIC, or with RNA SCNA calls in the absence of WES. We also used Fisher's exact test to assess the significance of differences between EAS GBM and the TCGA cohort (after subsetting the TCGA GBM cohort into 450 first human ancestry samples). TCGA ancestry information was derived from publicly available research data.
[0037] For the SCNA analysis of the GLASS dataset, we used published Titan results. Amplification was defined as a log ratio > 2.7, and missing data was defined as < 1.3. To validate the unique patterns of variation in EAS-GBM, we used a simplified pedigree marker approach on the GLASS dataset.
[0038] S3, based on the clustering calculation results, performs genomic pathway enrichment analysis on each cluster.
[0039] Marker gene selection and pathway enrichment analysis: To identify marker genes for each glioblastoma subtype, a comparative marker gene selection algorithm was performed on the GenePatter online platform using default settings and a threshold score >0.5. The top 500 differentially expressed genes (DEGs) in each subtype were selected as marker genes. To further compare the similarity between the EAS-GBM and EUR-GBM subtypes, Venn diagrams were constructed using the Venn diagram tool to examine overlapping marker genes, such as... Figure 5 As shown.
[0040] To reveal the gene ontology (GO) and pathway (Kyoto Encyclopedia of Genes and Genomes, KEGG) enrichment of each glioblastoma subtype, marker genes for each subtype were uploaded to the online tool Metascape. Gene set enrichment analysis (GSEA) was performed using the feela R package (v.3.12) based on the marker gene set in MSigDB (v.74). Significantly enriched gene sets were selected based on a threshold of q < 0.01. Figure 6 As shown.
[0041] Annotation of somatic mutation pathways: Furthermore, somatic mutation pathway annotation can be performed for chromosomal mutations. Somatic mutation pathway enrichment analysis was performed using Maftools, followed by visualization using the online tool Pathway Mapper. Homologous recombination deficiency (HRD) refers to abnormal chromosome alignment during meiosis and is closely related to cancer development. We selected a group of genes, including BRCA1 / 2, CDK12, ATM, MLHI / 3, MSH2 / 3 / 6, PMS1 / 2, and PALB2 mutations, to assess patients' HRD status. The DNA mismatch repair (MMR) pathway plays a crucial role in correcting errors during DNA replication. Genes associated with MMR, including MLHI, MLH3, MSH2, MSH6, PMS1, and PMS2, were investigated to determine MMR status, such as... Figures 7-8 As shown.
[0042] S4. Based on the results of the genomic pathway enrichment analysis, glioblastoma multiforme is classified into: proliferative (PL), neuro-synaptic (NS), metabolic (MB), and immunomodulatory (IM).
[0043] EAS-GBM molecular classification based on differential gene expression: Using RNA sequencing data from EAS-GBM, we first extended the existing classification to EAS-GBM. We then analyzed all IDHs... WT Molecular subtype classification was performed on 258 IDH wild-type EAS-GBM samples. WT Unsupervised nonnegative matrix factorization (NMF) consensus clustering of EAS-GBM RNA sequencing samples revealed four distinct clusters. These four clusters remained consistent after excluding recurrent tumors and could also be identified by t-distributed random neighborhood embedding (t-SNE) (Figure 3). Reference Figures 9-12The largest subtype, accounting for 34% of the sample, was highly enriched in the expression of immune-related genes, including immune infiltration markers, cytokines and cytokine receptors, as well as many genes related to immune function. We named this subtype the immunomodulatory (IM) subtype. The second most common subtype was named the proliferative (PL) subtype, characterized by differential expression of genes involved in promoting proliferation, including cell cycle progression and DNA replication, as well as DNA damage repair. The third subtype was named the metabolotropic (MB) subtype, showing differential expression of many cytoplasmic and mitochondrial ribosomal subunits, translation-related genes, and genes involved in oxidative phosphorylation and metabolism. Finally, the fourth subtype was named the synaptic (NS) subtype, enriched in differential expression of genes involved in neuronal activity, including neurotransmitters and neurotransmitter receptors, ion channels and pumps involved in action potential propagation, and many genes involved in synapses. Metascape pathway analysis of the four subtypes confirmed the enrichment of pathways related to immune function, cell cycle and replication, translation and metabolism, and synaptic function, respectively. It is noteworthy that PL-GBM and IM-GBM exhibited the largest distance in the average expression profile within EAS-GBM ( Figure 4 ).
[0044] We therefore explored a comparison of this classification with the three established subtype classifications and the more recent classifications based on two neurodevelopmental clusters and two metabolic clusters. These three classification criteria are: traditional pathological classification based on histological features, molecular classification based on key molecular markers, and TCGA molecular classification based on gene expression profiles.
[0045] Molecular subgroup mapping between EAS and EUR cohorts: To investigate the correlations between molecular subtypes identified from the EAS and EUR cohorts, we performed TCGA and EAS subtyping on each sample and examined the frequency of subtype overlap between the two groups. We reconstructed the contribution of each cluster feature to each sample based on the gene × cluster matrix of Wang et al. and our own using the fast combined nonnegative least squares function from the R nmf package. Samples were then assigned to the cluster with the highest contribution. We used Fisher's exact test to calculate the significance of overlap. Analyzing EAS-GBM and EUR-GBM together using three classification labels, we observed a significant overlap (53% overlap) between the immunomodulatory (IM-GBM) and mesenchymal (Mesenchymal) labels defined by Wang et al. The synaptic type has significant overlap with the preneuronal type label defined by Wang et al. (41% overlap). The proliferative type has significant overlap with the preneuronal type label defined by Wang et al. (30% overlap). The immunomodulatory subtype showed a small degree of overlap with the classic subtype defined by Wang et al. (25% overlap, q = 0.006). The metabolic subtype showed no significant correlation with any of the labels defined by Wang et al. Figure 13 ).
[0046] refer to Figure 14 The differences in classification became more pronounced when we analyzed EUR-GBM and EAS-GBM separately. In the EAS-GBM cohort, the mitochondrial (MTC) category, determined according to the method described by Garofano et al., significantly overlapped with our metaboloid GBM (MB-GBM) expression group. EAS-GBMs of the Garofano glycolytic / multi-metabolic subtype (GPM) were all classified as immunomodulatory (IM-GBM) in our classification. The two Garofano subtypes, based on neuronal (NEU) and proliferation markers (proliferating / progenitor cells; PPR), significantly overlapped with our synaptic (NS-GBM 100% classified as NEU) and proliferative (PL-GBM 89% classified as PPR) groups. However, among the EAS-GBMs classified as PPR by Garofano et al., only 62% mapped to PL-GBM in our classification, while 27% mapped to IM-GBM. The results from the EUR-GBM cohort highlight similar problems in existing classifications due to a lack of microenvironmental consideration. Almost all EUR-GBMs classified as GPM by Garofano et al. were further subclassified as immunomodulatory based on their expression characteristics. Of the other three subtypes of EUR-GBM classified by Garofano et al., a significant proportion were classified as IM-GBM by our method (68% of all MTCs, 17% of NEUs, and 60% of PPRs). These data suggest the existence of a subset of GBMs with strong immune system activity across all Garofano et al. subtypes. Therefore, a classification system is needed to identify samples with high immune system activity that may respond differently to treatment.
[0047] Analysis of EUR-GBM using our classification method revealed that MB-GBM, similar to that in EAS-GBM, was not detected in EUR-GBM. Figure 14 However, different metabolic states of tumor cells may exist within EUR-GBM. These findings suggest that our classification is specific to EAS-GBM. Differences in the distribution frequency of genetic drivers between EAS-GBM and EUR-GBM may lead to this expression difference.
[0048] Further extending this concept, by identifying differential characteristics among populations with different human ancestry or genetic backgrounds, the origin of samples can be determined, thus allowing for the adaptation of specific GBM genotyping systems. Furthermore, suitable biomarkers can be screened from these differential characteristics for clinical applications. The following section further describes the screening and use of biomarkers.
[0049] EAS-GBM expression classification marker genes: There is a large amount of overlapping marker genes between the EUR proneural type and the EAS proliferative type (PL, n = 72) and the EAS synaptic type (NS, n = 173). Figures 15-17 EUR preneuronal GBM marker genes ASCL1, DLL3, NKX2-2 and OLIG2 High expression in both the EAS proliferative and EAS synaptic types indicates that both are similar to EUR preneuronal GBM. The clear separation of EAS proliferative and EAS synaptic types in terms of gene ontology and enriched signaling pathways suggests that preneuronal GBM may be a mixed population, not segregated in EUR-GBM due to limited sample size. (Note: The last sentence appears to be incomplete and possibly refers to a separate point about preneuronal marker genes.) ASCL1, DLL3, NKX2- 2 and OLIG2 The expression level was higher in the EAS proliferative subtype than in the EAS synaptic subtype, indicating an intrinsic difference between the two subtypes.
[0050] Next, we identified the characteristic marker genes of proliferating GBM. (DLL3, MYB, MYCN, BRIP1, CCND2, ASCL1 Characteristic marker genes of synaptic GBM GRM3, FGFR2, SIX1, GRIN2A, WIF1, ATP2B3, KIAA1598, FLT3, and TERT The lowest expression levels were observed. Cell cycle, PI3K-AKT, and FoxO signaling pathways were significantly enriched in the EAS proliferative type, while synaptic, Rap1, and Ras signaling pathways were enriched in the EAS synaptic type. Only two marker genes overlapped between the EAS metabolotype and the EUR classic type. EUR classic type GBM marker genes NOTCH3、 TP53, GLI2, and SMO Expression in the EAS metabolite is lower than in other EAS-GBM subtypes. This also indicates that there is no corresponding group for the EAS metabolite GBM in the classification by Wang et al.
[0051] The marker genes for EAS metabolotype GBM include ASPSCR1, NTHL1, MAD1L1 and GADD45GThese genes are primarily enriched in fatty acid metabolism and transcriptional dysregulation signaling pathways. The EAS immunomodulatory GBM and EUR mesenchymal types showed a large number of overlapping marker genes (n = 304), including mesenchymal characteristic genes. TRADD, CD44, MET and immune-related genes C1R, CTLA4, PDCD1, CD86, CD274 and RELB Our results also showed that immune-related pathways such as the PD-1 checkpoint and TNF signaling pathways were activated in EAS-immunomodulatory GBM. A summary of marker genes for each EAS-GBM subtype is available in [link to relevant documentation]. Figure 18 .
[0052] Somatic copy number variations (SCNAs) among the four subtypes were highly consistent with their characteristic signaling pathways, and significant differences were found in GBM between the second population (EAS) and the European population (EUR): refer to Figure 19Significantly recurring SCNAs in each group were consistent with characteristic pathways in each group. Overall, proliferative GBM had the highest number of significantly enriched arm-level copy number variations. In contrast, the genome of synaptic GBM was significantly “quieter,” with lower rates of chr7q amplification and chr10, chr14q, and chr22q deletions (very common in other types). Chromosome 10 deletions were present in almost all metabolic and immunomodulatory GBM, leading to the loss of PTEN, FAS, MGMT, and other tumor suppressor genes. Similarly, chromosome 19 amplification, associated with better prognosis in EUR-GBM, was significantly enriched and present in almost all metabolic GBM. These data show a gradient of genomic instability among EAS GBM subgroups, from highest in proliferative GBM to most stable in synaptic GBM. Specific focal copy number variations in each EAS-GBM subtype were significantly correlated with the expression of their marker genes. Metabolic GBM showed significant enrichment of FGFR3 and CCND1 amplification (but no changes in RNA expression), and RB1 deletion was almost universally present. RB1 deletion also led to a significant decrease in RB1 expression in the metabolic GBM group (q = 0.01). The immunomodulatory group showed enrichment of MET amplification, which has recently been found to regulate immune activity in GBM. MET expression in this cluster was on average 30% higher than in other groups, but this was not statistically significant (q = 0.08). Proliferative GBM showed enrichment of CCNE1 (q = 0.07) amplification, which was associated with significant overexpression in this group (q = 0.01). TERT amplification was also significantly enriched in this type, leading to a 40% increase in expression (q = 0.06). These data suggest that differences in SCNAs may be a contributing factor to differential pathway activation among EAS-GBM subgroups. Reference Figure 7 TCGA IDH, the first human ancestor WT Compared to GBM queues, including EGFR The acquisition of chr7p and chr7q in EAS IDH WT The incidence rates in GBM were all low (respectively) and When comparing untreated subsets, the lower incidence of chr7 obtained by EAS compared to EUR remained statistically significant (chr7p: and chr7q: This low incidence of chr7 was also present in nine TCGA GBM samples from second human ancestry, although the q values were lower due to the much smaller sample size (chr7p obtained q = 0.03, chr7q obtained q = 0.02). This is consistent with IDH samples from the same study but primarily from the B ancestral region. WT Compared to GBM samples, our IDH in the GLASS queue WT A significantly low incidence of chr7 was also detected in EAS-GBM (from regions C, D, and E). ). chr19 was obtained in EAS IDH WT Enrichment was also observed in GBM (chr19p: q = 0.0004, chr19q: q = 0.09), but it did not reach significance in the GLASS cohort. In the treated and untreated GBM subgroups of the GLASS dataset, deletions were obtained on the short arm of chromosome 7 in EAS-GBM (including...). EGFR All genes remained statistically significant (chr7p in the untreated group). (Treatment group q = 0.01). These data indicate that classic GBM driver genes... EGFR The amplification rate in EAS-GBM was significantly lower than that in EUR-GBM, which may be an important reason for the lack of classical expression subpopulations in EAS-GBM.
[0053] EAS synaptic GBMs exhibit low levels of genomic instability. They are mostly euploid with fewer copy number aberrations. They also show lower tumor mutational burden (TMB) and very low microsatellite instability rates. These data suggest that, compared to other subtypes, synaptic GBMs not only exhibit differences in gene expression but also demonstrate different mechanisms of variation generation.
[0054] Somatic single nucleotide variants (SNVs) in EAS GBM reflect SNVs in EUR-GBM and copy number variants in EAS-GBM: refer to Figure 20 Focusing on significantly recurring driver genes in EAS-GBM (defined by MutSigCV), we identified many genes that have been identified as driver factors in EUR-GBM, including the most frequently mutated driver genes: TP53, ATRX, PTEN and EGFR (.IDH) WT EAS-GBM display NF1 The incidence of SNVs was significantly higher (q = 0.01, 23% in EAS compared to 12% in EUR). ATRX SNVs occurred more frequently (q = 0.01, 13% in EAS compared to 6% in EUR), and were rare. H3F3A SNVs. EGFR SNVs were less frequent in EAS GBM (q = 0.075, 15% in EAS vs. 25% in EUR). These data suggest that in EAS-GBM, activation occurs via SCNAs and SNVs. EGFR The incidence rates were all lower than those of EUR-GBM.
[0055] We also performed mutation SNV characterization analysis. (Compared with EUR IDH) WT In comparison, EAS IDH WT GBM showed a lower TMB (median: 0.889 / Mbp vs. 0.920 / Mbp, Kruskal-Wallis test: p = 0.014). After controlling treatment status, we found no significant difference in SNV characteristics between EAS and EUR GBM.
[0056] Immunohistochemistry (LHC) staining: For histological examination, tumor samples from patients were fixed overnight at 4°C with 4% paraformaldehyde. The samples were then embedded in paraffin. Tissue sections were first stained with hematoxylin and eosin (H&E). For immunohistochemistry, tissues were blocked with bovine serum albumin (BSA) at room temperature. Subsequently, tissues were incubated overnight at 4°C with primary antibodies (anti-DLL3, anti-CRMPI1, anti-GRM3, anti-FGFR2, anti-p16, anti-CD44, anti-C1R, anti-Ki67, anti-CD4, anti-CD8, anti-CD20, anti-CD68, anti-CD163). The next day, secondary antibodies were added and incubated for 30 minutes at 37°C. Finally, sections were photographed using an optical microscope at 200x or 400x magnification. Protein expression levels were independently assessed and semi-quantitatively scored by two neuropathologists with over 10 years of experience using the widely used Immunoreactivity Rating System (IRS). In case of discrepancies, the final IRS score will be confirmed by a third senior neuropathologist. The IRS score is calculated based on the percentage of positive cells (4, >80%; 3, 51-80%; 2, 10-50%; 1, <10%; 0, 0%) and staining intensity (3, strong; 2, medium; 1, weak; 0, no staining), with the final IRS score ranging from 0 (no staining) to 12 (maximum staining). We averaged five high-power fields (technical replicates) for each patient, with each EAS subtype including five patients with IDH wild-type glioblastoma.
[0057] Multiple immunofluorescence (MLF) staining treatment: Multicolor staining and multispectral imaging were used to identify the expression of GFAP, Ki67, CD3, FOXP3, CD1IB, CD68, CD206, CD90, S100A4, and CD31 in the tumor microenvironment. Multicolor immunofluorescence staining was performed using a kit. Different primary antibodies were applied sequentially, followed by incubation with horseradish peroxidase-labeled secondary antibodies and tyramide signal amplification (TSA). After each TSA operation, the sections were microwave-treated. After all human antigens were labeled, the cell nuclei were stained with 4'-6'-diamino-2-phenylpyridine (DAPI). To obtain multispectral images, the stained sections were scanned using a Mantra system that captures fluorescence spectra from 420 nm to 720 nm at 20 nm wavelength intervals with identical exposure times; the scans were merged to construct a single stacked image.
[0058] Images of unstained and single-stained sections were used to extract the autofluorescence spectra of the tissue and each fluorescein, respectively. The extracted images were further used to construct a spectral library for multispectral dealiasing using inForm image analysis software. We then used this spectral library to obtain reconstructed images of the sections after removing autofluorescence.
[0059] The results are as follows Figure 21 and Figure 22 The image shows IHC staining of immune cell markers CD4+ (helper T cells), CD8+ (cytotoxic T cells), CD68+ (microglia), CD163+ (M2 microglia), and EAS subtype markers DLL3 (PL), GRM3 (NS), FGFR2 (NS), p16 (MB), CD44 (IM), and C1R (IM) in four EAS-GBM subtype samples. One representative image with magnified sections was selected from five samples of each subtype. The scale bar is 40µm for the full image and 20µm for the inset. n = 5 GBM individuals for each EAS subtype.
[0060] Figure 23 These are multiplex immunofluorescence images from two EAS-GBM patients. The left panel shows two sites from a proliferative GBM. The right panel shows two sites from an immunomodulatory GBM. In both cases, the top panel shows the protein levels of CD3, CD11B, FOXP3, GFAP, and Ki-67; the bottom panel shows the protein levels of CD31, CD68, CD90, CD206, and S100A. The scale bar is 40 µm.
[0061] We validated our single-cell findings using immunohistochemistry (IHC) and multiplex immunofluorescence (mIF) assays. The IHC and mIF results were consistent with the subtype markers obtained from bulk RNA and scRNA sequencing. DLL3 showed significantly higher expression in proliferative GBM compared to other subtypes (Mann-Whitney U test, after correction). Compared with NS and MB, Ki67 expression was increased in PL and IM (Mann-Whitney U test, after adjustment). GRM3 and FGFR2 were found only in NS GBM (Mann-Whitney U test, after correction). and p16 staining was strongest in metabolic GBM and also prominent in immunomodulatory GBM. Preneuronal and synaptic GBM were predominantly p16 negative. C1R, a proteolytic subunit of the complement system C1 complex, was highly and diffusely expressed in IM GBM tumor cells (Mann-Whitney U test, corrected). CD44, an adhesion molecule, was also found to be overexpressed in IM GBM tumor cells (Mann-Whitney U test, after correction). To rule out the possibility that the NS subtype was an artifact caused by non-tumor brain tissue contamination, we further examined the expression of NS markers between the tumor region and the peritumoral region. We found that FGFR2 and GRM3 were significantly more expressed in the tumor than in the peritumoral region. Figure 20 , These cells are selectively expressed in tumor cells (identified by enlarged, deeply stained nuclei, and atypical, mitotic nuclei). Based on morphology alone, our senior pathologists were unable to distinguish these subtypes. No significant increase in non-tumor tissue was found in NS tumors during pathological re-examination.
[0062] Compared to proliferative and synaptic GBM, immunomodulatory and metabolic GBM showed significantly greater infiltration of CD4+, CD8+ T cells and Treg cells. Figure 23 and 24 Mann-Whitney U test, after all corrections Despite the total count of CD4+ and CD8+ T cells (approximately 40-80 / mm), 2 The number of CD68+ microglia and CD163+ / CD206+ M2 microglia is significantly lower than in other cancers. CD68+ microglia and CD163+ / CD206+ M2 microglia are the most prevalent immune cells in the GBM microenvironment, and also show higher numbers in immunomodulatory and metabolic GBM (Mann-Whitney U test, after all adjustments). We also observed a greater number of S100A+CD90+- fibroblasts in the immunomodulatory EAS-GBM. Figure 23 and 24 Mann-Whitney U test, after correction For all cell types assessed, intratumoral cell count heterogeneity was less than intertumor heterogeneity, indicating that the differences between tumors were not caused by sampling artifacts. Figure 25 B cell infiltration (labeled with CD20) did not show a significant difference between the EAS-GBM subtypes. Figure 26These data suggest that the higher expression profile of immune RNA in the immunomodulatory GBM subtype is associated with higher T cell infiltration in this subtype.
[0063] In summary, biomarkers that can be used clinically to achieve GBM subtyping can be selected by combining at least one from each of the following categories: Gene markers used to identify the proliferative type MYB, MYCN, BRIP1, CCND2, PDGFRA, PIK3CA, MAPK1, AKT3 At least one of the following, protein markers DLL3 and / or Ki67; Genetic markers used to identify the neural synaptic type WIF1, ATP2B3, GRM3, GRIN2A, KIAA1598, FLT3SIX1, FGFR2 At least one of the following, protein markers: GRMM3 and / or FGFR2; Gene markers used to identify the metabolites ASPSCR1, NTHL1, MADIL1, TERT, GADD45G, ACSF2, ACADVL, ACSF3 At least one of them, protein marker P16; Gene markers used to identify the immunomodulatory type REL, TGFBR2, COL3A1, COL1A1, SELE, CD274, PDCD1LG2, CCR4L, IL21R, IL7R, FAS At least one of the following, and at least one of the following protein markers: CD44, C1R, and Ki67.
[0064] Using the largest second ancestor (EAS) multimodal GBM cohort to date, containing 367 samples, we demonstrated that EAS-GBM differs from EUR-GBM in the prevalence of GBM expression isotypes and common driving factors. EAS-GBM showed the absence of classic expression isotypes and EGFR Low prevalence of activation. We introduced four subtypes (proliferative, synaptic, metabolic, and immunomodulatory) for EAS-GBM, encompassing primary and recurrent GBM, IDH wild-type, and mutant GBM. These four subgroups exhibited significant differences in genetic, immunological, and clinical characteristics. EAS-GBM also includes a subgroup with high immune infiltration. The proliferative, synaptic, metabolic, and immunomodulatory subtypes in our study are in good agreement with single-cell data from previous studies. However, we expanded these findings by providing additional information about the EAS-GBM tumor microenvironment. We also provided IHC markers for each subtype for rapid classification and assessment of intratumoral heterogeneity in routine pathology workflows.
[0065] Those skilled in the art will understand that the steps, measures, and solutions in the various operations, methods, and processes discussed in this application can be alternated, modified, combined, or deleted. Furthermore, other steps, measures, and solutions in the various operations, methods, and processes discussed in this application can also be alternated, modified, rearranged, decomposed, combined, or deleted. Furthermore, steps, measures, and solutions in related technologies that are similar to those disclosed in this application can also be alternated, modified, rearranged, decomposed, combined, or deleted.
[0066] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0067] In the description of this specification, specific features, structures, materials, or characteristics may be combined in any suitable manner in one or more embodiments or examples.
[0068] The above description is only a partial implementation of this application. It should be noted that for those skilled in the art, other similar implementation methods based on the technical concept of this application, without departing from the technical concept of this application, also fall within the protection scope of the embodiments of this application.
Claims
1. A method for classifying glioblastoma multiforme, characterized in that, It includes the following steps: Acquire single-cell omics analysis data from multiple tumor cell samples; Clustering calculations were performed on the whole exon sequencing data and / or RNA sequencing data based on the single-cell omics analysis data; Genomic pathway enrichment analysis was performed on each cluster based on the clustering calculation results; Based on the results of the genomic pathway enrichment analysis, glioblastoma multiforme was classified into: proliferative, neurosynaptic, metabolic, and immunomodulatory types.
2. The method as described in claim 1, characterized in that, After obtaining single-cell omics analysis data from multiple tumor cell samples: Different human ancestral origins can be distinguished based on the single-cell omics analysis data; Compare somatic copy number variations and / or somatic single nucleotide variations among different human ancestral origins; The population-specific differences in glioblastoma multiforme were identified based on the differential data.
3. The method as described in claim 2, characterized in that, The comparison of somatic copy number variation and / or somatic single nucleotide variation among different human ancestral origins includes: The typing results based on the genomic pathway enrichment analysis are mapped to other typing criteria.
4. The method as described in claim 3, characterized in that, The other classification criteria include at least one of the following: traditional pathological classification based on histological characteristics, molecular classification based on key molecular markers, and TCGA molecular classification based on gene expression profiles.
5. The method as described in claim 1, characterized in that, The clustering calculation includes: Perform NMF clustering or k-means clustering on whole exome sequencing data and / or RNA sequencing data; The number of subtypes is determined based on the number of clusters after clustering.
6. The method as described in claim 1, characterized in that, It also includes screening for genotyping markers based on the results of the genomic pathway enrichment analysis, wherein the genotyping markers include gene markers and / or protein markers.
7. The method as described in claim 6, characterized in that, The typing markers include: Gene markers used to identify the proliferative type MYB, MYCN, BRIP1, CCND2, PDGFRA, PIK3CA, MAPK1, AKT3 At least one of the following, protein markers DLL3 and / or Ki67; Genetic markers used to identify the neural synaptic type WIF1, ATP2B3, GRM3, GRIN2A, KIAA1598, FLT3SIX1, FGFR2 At least one of the following, protein markers: GRMM3 and / or FGFR2; Gene markers used to identify the metabolites ASPSCR1, NTHL1, MADIL1, TERT, GADD45G, ACSF2, ACADVL, ACSF3 At least one of them, protein marker P16; Gene markers used to identify the immunomodulatory type REL, TGFBR2, COL3A1, COL1A1, SELE, CD274, PDCD1LG2, CCR4L, IL21R, IL7R, FAS At least one of the following, and at least one of the following protein markers: CD44, C1R, and Ki67.
8. The method as described in claim 1, characterized in that, The single-cell omics analysis data of the multiple tumor cell samples excluded data from IDH mutant samples.
9. The method as described in claim 1, characterized in that, The single-cell omics analysis data includes genomic, transcriptomic, epigenomic, and / or proteomic data of individual tumor cells.