Estimating proportions of cell types from DNA methylation microarray data to determine glioma immune microenvironment
GIMiCC addresses the limitations of existing methods by providing a CNS-specific DNAm deconvolution tool for glioma, enabling accurate immune microenvironment analysis and personalized treatment strategies through high-resolution cell type estimation.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- TRUSTEES OF DARTMOUTH COLLEGE THE
- Filing Date
- 2025-10-27
- Publication Date
- 2026-05-07
AI Technical Summary
Current techniques for characterizing the immune microenvironment in glioma, such as flow cytometry, histopathology, and single-cell sequencing, are challenging due to requirements for fresh samples, expensive equipment, and expert staff, while existing DNA methylation (DNAm) deconvolution libraries do not include CNS-specific cell types, limiting the ability to optimize immunotherapeutic approaches for glioma.
The Glioma Immune Microenvironment Composition Calculator (GIMiCC) is developed to deconvolute glioma DNAm microarray data, utilizing a hierarchical structure and machine learning to derive deconvolution libraries for CNS-specific cell types, providing a high-resolution analysis of the immune microenvironment.
GIMiCC accurately estimates the composition of glioma microenvironments, enabling personalized immunotherapeutic strategies by identifying composition-independent DNAm alterations associated with immune infiltration and improving clinical evaluation of glioma.
Smart Images

Figure US2025052585_07052026_PF_FP_ABST
Abstract
Description
Docket No.: 231 / 0021RClient Reference: 2024-026-02SYSTEM AND METHOD FOR ESTIMATING THE PROPORTIONS OF CELL TYPES FROM DNA METHYLATION MICRO ARRAY DATA TO DETERMINE GLIOMA IMMUNE MICROENVIRONMENTSTATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT
[0001] This invention was made with U.S. government support under Grant Numbers R01 CA207360, R01 CA216265 and P20 GM104416, awarded by the National Institutes of Health (NIH). The government has certain rights in this invention.RELATED APPLICATION
[0002] This application claims the benefit of co-pending U.S. Provisional Application Serial No. 63 / 712,935, entitled SYSTEM AND METHOD FOR ESTIMATING THE PROPORTIONS OF CELL TYPES FROM DNA METHYLATION MICRO ARRAY DATA TO DETERMINE GLIOMA IMMUNE MICROENVIRONMENT, filed October 28, 2024, the entire teachings of which, including Descriptions, Drawings, Claims and Appendices thereto, are expressly incorporated herein by reference.FIELD OF THE INVENTION
[0003] This invention relates to systems and methods for diagnosis of cancerous conditions from cellular samples based upon deconvolution of DNA methylation data, and more particularly to calculating a microenvironment related to glioma.BACKGROUND OF THE INVENTION
[0004] Adult diffuse glioma represents the most common primary malignancy within the central nervous system (CNS), affecting an estimated 16,000 Americans each year[L 2] (noting that any and all [bracketed] numbers refer to references provided in the References section hereinbelow by way of useful background information). Pathological and molecular markers have historically classified this class of CNS tumors. Pathologists now rely on criteria set by the World Health Organization CNS tumor classification (WHO CNS5) to make diagnoses[3-5]. Isocitrate dehydrogenase (I H) mutations were early adopted classification markers and strong markers of survival; lack of IDH mutation is associated with a dismal prognosis with a median survival of 14.6 months] 1 1. As of the most recent WHO CNS5 criteria, IDH mutations are used in combination with other findings to diagnose adult diffuse gliomas differentially and identify optimal treatment approaches [4, 5],Docket No.: 231 / 0021RClient Reference: 2024-026-02
[0005] The immune microenvironment is a critical component contributing to tumor growth and survival, as well as an avenue for the treatment of many cancers. The tumor microenvironment in adult diffuse glioma consists of CNS resident populations (neurons, oligodendrocytes, astrocytes, and microglia), infdtrating immune cells, tumor vasculature, and structural support cells. The immune compartment consists of immune cells attempting to eliminate cancer cells or necrotic tissue (tumor-specific and non-tumor-specific T-cells, “proinflammatory” macrophages / microglia, neutrophils, etc.) or cells that the tumor is "hijacking" to promote tumor survival (regulatory T cells, myeloid-derived suppressor cells, “anti-inflammatory ” macrophages / microglia, etc.)[6, 7],
[0006] There are predominant immunosuppressive factors in adult diffuse gliomas, particularly GBM, that impact the efficacy of immunotherapies[8]. This is driven by various factors, including the altered expression of cytokines and growth factors[9] and the upregulation of immunosuppressive myeloid populations
[0010] , These cells can differentially regulate cytokine and chemokine expression to recruit regulatory T cells, inhibit anti-tumor T cells, and stimulate exhaustion states in T cells
[0011] , The limited activity’ or presence of these T cells decreases the effectiveness of therapies that try’ to bolster antigen-specific responses, such as tumor vaccines or adoptive cell therapy[12, 13], Therapies that try to evade the tumor immune suppression, such as immune checkpoint inhibitors, are receiving mixed results in clinical trials) 14], Although various therapies have initial successes, there is currently no FDA-approved immunotherapy for treating glioma.
[0007] There are many’ hypotheses as to why immunotherapy is not working for adult diffuse gliomas. Firstly, the anatomy of the CNS lends itself to unique immunological phenomena that don't exist for other tumor types, such as the blood-brain barrier, meningeal inflammatory pathways, and glymphatic system[15-17]. Secondly, many tumors have a low total mutational burden (TMB), which limits the pool of tumor antigens that an immune response can be enacted upon
[0018] . Lastly, there is significant variability in the tumor immune microenvironment
[0019] , Not only across subtypes of glioma but across tumors with the same diagnoses, it is observed strong heterogeneity in the amount of vascularity, lymphocyte infiltration, and macrophage activity [20-22] . These underlying differences of the immune system, or how drugs can enter the tumor, may vary’ patient to patient. Understanding more about why immunotherapies fail or succeed on an individual patient level may allow us to conceptualize further how to optimize immunotherapeutic approaches[23, 24],Docket No.: 231 / 0021RClient Reference: 2024-026-02
[0008] Currently, the predominant techniques to characterize the immune microenvironment cells include flow / mass cytometry, histopathology, or single-cell sequencing. Many of these options are challenging due to requirements for fresh samples, expensive equipment / reagents, optimized protocols, and staff with expertise. One technique that is being rapidly adopted in clinical practice, and can be utilized for immune profding, is DNA methylation (DNAm) assays. DNAm refers to the epigenetic phenomena in which nucleotide bases are covalently modified by adding a methyl group; the most common loci for this are cytosine-guanine dinucleotides (CpGs)
[0025] , DNAm patterns in regulatory elements of genes can promote or inhibit the expression of genes. This mechanism can also lead to the activation of oncogenes or suppression of tumor suppressor genes in cancers and other disorders
[0025] , In addition to facilitating tumor progression, DNAm alterations are utilized during development to guide progenitor cells into terminally differentiated phenotypes [25- 28], Changes to DNAm allow for cell-ty pe specific genes to be expressed, allowing for the diversity of all human cell types to be derived from the same genetic information. Researchers have identified patterns of cell-type specific DNAm and use this information to deconvolve bulk DNAm samples; in other words, to estimate the relative proportions of each cell type in the sample[29-34],
[0009] Utilizing DNAm data for scalable CNS tumor immune profiling is highly compatible with current developments in using this data for tumor classification. Various iterations of classifiers have been developed as this is a rapid area of research[35-38], including in 2018, where Capper et al. utilized patterns of DNAm microarray data to establish a neuropathology7classifier that could predict tumor subtype
[0039] . This work shows the vast heterogeneity in epigenetic alterations in distinct glioma subty pes but has strong clinical utility’ and can stratify patient populations. Therefore, building a high-throughput, cost- effective platform that can predict the immune environment of each sample in parallel to the tumor subtype would further leverage the already existing clinical utility of DNAm profiling and promote further research into optimizing both aspects in unison.
[0010] Currently, there are limited deconvolution libraries for the deconvolution of glioma DNAm data. Carcinoma deconvolution libraries have been introduced, such as HiTIMED
[0030] , MethylResolver
[0040] , and methylCIBERSORT[41 , 42].' however, these methods do not include CNS-specific cell ty pes. Additionally, single-cell DNAm data has been collected on CNS tumors; however, the feature space of this data is not compatible with microarray data. A goal is to combine the best elements of all the previous methods with theDocket No.: 231 / 0021RClient Reference: 2024-026-02 recently validated brain deconvolution tool, HiBED
[0032] , to develop DNAm deconvolution libraries for glioma that are subtype-specific and maintain a high resolution of both CNS and immune cells.
[0011] More generally, the immune microenvironment is a critical component contributing to tumor grow th and survival. The glioma tumor microenvironment consists of resident immune cells such as microglia and infiltrating leukocytes from the periphery, all of which can have both pro and anti-tumor properties. Variability of the tumor microenvironment is apparent both betw een and within molecular subtypes of glioma. Personalized medicine approaches may be beneficial when treating glioma with immunotherapeutics; thus, there is a critical need to be able to detect the tumor-immune landscape in patients.
[0012] Proper detection of both tumor subclass and immune landscape is critical in the pursuit of personalized immunotherapeutic treatment strategies for glioma. DNAmethylation (DNAm) biomarkers are promising on both of these fronts. DNAm-based biomarkers for molecular classification have been developed for clinical use. However, the prediction of the cellular landscape of the tumor microenvironment is not as well defined.
[0013] It is therefore desirable to provide a computerized calculator for use in analyzing DMA methylation databases to determine the composition of glioma microenvironment markers / conditions in a variety of available cell types. SUMMARY OF THE INVENTION
[0014] This invention overcomes disadvantages of the prior art by providing a Glioma Immune Microenvironment Composition Calculator (GIMiCC), that defines and system and method for cellular deconvolution of glioma DNAm microarray data. By w ay of non-limiting example, using data from 17 isolated cell types, the system and method herein allows the derivation of the deconvolution libraries in the biological context of selected genomic regions and validate the results using independent datasets. GIMiCC is utilized to illustrate that DNAm-based glioma classification is unlikely to be biased by compositional variation. In addition, GIMiCC is utilized to identify composition-independent DNAm alterations that are associated with high immune infiltration. GIMiCC can be optimized to advance the clinical evaluation of glioma.
[0015] In an illustrative embodiment, a system and method for diagnosing glioma in a patient using a processor and a user interface is provided. A data input process receivesDocket No.: 231 / 0021RClient Reference: 2024-026-02 information related to the patient’s cancerous tissue based on DNA methylation. A data store, which that includes a plurality of layers of information, is arranged relative to types of cancerous conditions in tissue and predetermined types on non-cancerous tissue based on DNA methylation therein. An analysis process, including a trained machine learning process, compares DNA methylation characteristics in the patient’s tissue to the layers of information and performs a match so as to trace the cancerous conditions. A diagnostic process provides a user with information relative to the match. Illustratively, the layers can include (a) deconvolution results for at least four subtypes of brain tumors, and (b) deconvolution of the immune microenvironment of each type of the brain tumors, respectively. The analysis process can be constructed and arranged to deconvolve glioma DNAm microarray data from at least (e.g.) 17 isolated cell types. The data store can include publicly available information accessed through a public data communication network. The machine learning process can be trained using classifiers related to predetermined DNA methylation characteristics for related cancerous conditions. At least one of the data input process, the analysis process and the diagnostic process can be operated using a processor on a user-controlled computing device with a user interface. More generally, the system and method can be arranged for diagnosing and reporting upon cancerous conditions. The system and method can be operated by logging into, by a user, a public-network based subscription sendee, or using of validated credentials that are established for the user. An optimized treatment for the patient can be derived based upon results of the diagnostic process, and treatment of the patient can be performed using conventional and novel treatments therewith.BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The invention description below refers to the accompanying drawings, of which:
[0017] Fig. 1 is a diagram showing Structure of GIMiCC hierarchical deconvolution using DNAm data from isolated cell types, in which the hierarchical structure allows users to specify the resolution in which they want to investigate the microenvironment, such as in a broad manner at Layer 2 into ‘“myeloid” or “lymphoid” or in a specific manner at Layer 5, where details about B and T cell subtypes are distinguished;
[0018] Figs. 2A-2E are diagrams of CpGs selected for the L0 layer in the GIMiCC hierarchical deconvolution and accuracy of purity' predictions, with Fig 2A showing rows of heatmaps represent CpGs, the columns represent the average methylation value in the sampleDocket No.: 231 / 0021RClient Reference: 2024-026-02 group in the test set of the Capper et al. dataset
[0039] , and the color are representative of the methylation level of the sample at the specific CpG. The marks on the left of the figure denote which LO library the CpG belonged to, and in which rows were organized via hierarchical clustering, and Figs. 2B-2E showing correlation of GIMiCC estimated purity7and the purity estimated in the Capper et al. manuscript[39, 66], and circles denote samples from the test set, while crosses denote samples from the training set;
[0019] Figs. 3A-3I are diagrams showing Genomic context of CpGs utilized in GIMiCC, with Figs. 3A-3C showing the methylation level of representative CpGs from GIMiCC libraries at functionally relevant genes, Figs. 3D-3F showing the enrichment analysis testing for the odds of the DMCs of the GBM L0 (D), AST LO (E), and all other deconvolution layers (F) being in a CpG island, shore, shelf, or open sea region, and Figs. 3G- 31 showing enrichment analysis testing for the odds of the DMCs of the GBM L0 (G), AST LO (H), and all other deconvolution layers (I) being within a particular gene regulatory7element. N = north, S = south, UTR = untranslated region, TSS = transcriptional start site;
[0020] Figs. 4A-4D are diagrams showing GIMiCC predictions in Capper et al. test set and healthy controls, with Fig 4A showing boxplots of the distribution of the sample's predicted cell fractions, tumor cells. The x-axis represents the group of samples, whereas the boxplot color represents the different L0 library used to derive the purity, Fig. 4B showing GIMiCC predictions of the composition of the healthy control samples in the Capper et al. dataset, in which each column represents a single sample, the y-axis represents the scaled proportion of the non-tumor portion, and ADENOPIT = anterior pituitary gland, CEBM = cerebellum, HEMI = cortical hemisphere, HYPTHAL = hypothalamus, PINEAL = pineal gland, PONS = pons. WM = white matter. GIMiCC predictions of the composition of (as shown in Fig. 4C) the high immune infiltrate samples and (as shown in Fig. 4D) low-yield samples in the Capper et al. dataset.
[0021] Figs. 5A-5F are diagrams showing GIMiCC predictions in the TCGA glioma samples, in which each column represents a single sample, and the y-axis represents the proportion of cells within each tumor sample, and in which Figs. 5A-5C show lower layer deconvolution results for four subtypes of brain tumors, while Figs. 5D-5F show deeper layer deconvolution of the immune microenvironment of each tumor ty pe;
[0022] Figs. 6A-6E are diagrams showing semi-quantitative validation of GIMiCC predictions in the TCGA dataset, in which Figs. 6A-6C show correlation of consensus purity estimate (CPE)
[0072] with the GIMiCC predicted tumor proportion Fig. 6D shows theDocket No.: 231 / 0021RClient Reference: 2024-026-02 relationship between ESTIMATE immune scores
[0073] and by GIMiCC derived immune estimates, in which colors represent the three tumor types (red = GBM. green = OLG, blue = AST), and Fig. 6E depicts the same relationship, but including microglia into the microenvironment;
[0023] Figs. 7A-7F are diagrams showing DMCs independent of cellular composition associated with high tumor immune infiltration, in which Figs. 7A-7C show EWAS analysis comparing GBM glioma to AST glioma in the Capper et al. dataset using (A) Model 1 and (B) Model 5 (C), wherein summarized results of all models are used to compare tumor types, and Figs. 7D-7F show EWAS analysis comparing the highest immune infiltrating GBM glioma to the lowest immune infiltrating GBM glioma in the Capper et al. dataset using (D) Model 1 and (E) Model 5 (F). wherein summarized results of all models used to compare high and low infiltrating GBM gliomas; Model 1 : Univariable analysis; Model 2: Controlled for tumor proportion; Model 3: Controlled for proportions of tumor, angiogenic, glial and immune cells; Model 4: Controlled for proportions of tumor, angiogenic, astrocyte, microglia, oligodendrocyte, myeloid, and lymphoid cells; and Model 5: Controlled for proportions of tumor, angiogenic, astrocyte, microglia, oligodendrocyte, T cell, B cell, NK, neutrophil, and monocyte cells. Ref = reference group; DMCs = differentially methylated CpGs; Hyper = number of hypermethylated DMCs; Hypo = number of hypomethylated DMCs;
[0024] Figs. 8A-8C are diagrams showing higher levels of angiogenic cells within the tumor microenvironment are associated with worse 5-year survival in TCGA samples, in which Kaplan-Meier curves are stratified by the median percentile of angiogenic cells where if a sample is above the median, it is assigned to the “hot” (red) strata while the remaining is assigned to the “cold” (blue) strata, and individual models were run for each tumor type, in which Fig. 8 A shows GBM angiogenic hot (n = 96, median angiogenic proportion = 4.7) versus GBM angiogenic cold (n = 32, median angiogenic proportion = 1.5), Fig. 8B shows OLG angiogenic hot (n = 62, median angiogenic proportion = 3.3) versus OLG angiogenic cold (n = 104, median angiogenic proportion = 1.6), and Fig. 8C shows AST angiogenic hot (n = 57, median angiogenic proportion = 3.5) versus GBM angiogenic cold (n = 80, median angiogenic proportion = 1.5); and
[0025] Fig. 9 is a block diagram showing a generalized computing environment for performing the processes and steps of the system and method herein.DETAILED DESCRIPTIONDocket No.: 231 / 0021RClient Reference: 2024-026-02
[0026] Abbreviations
[0027] GIMiCC = Glioma Immune Microenvironment Composition Calculator; CNS= central nervous system; DNAm = DNA methylation; CpG = cytosine-guanine dinucleotide;DMC = differentially methylated CpG; GBM = glioblastoma; IDH = isocitrate dehydrogenase; TCGA = The Cancer Genome Atlas; CPE = consensus purity estimate;EWAS = epigenome-wide association study; AST-HG = high-grade IDH mutant astrocytoma;AST = IDH mutant astrocytoma; OLG = oligodendroglioma; VAMP2 = vesicle associated membrane protein 2; TREM2 = triggering receptor expressed on myeloid cells 2; HLA-DOB = major histocompatibility complex, class II, DO beta; GABA = GABAergic neurons; GLU = glutamatergic neurons; Oligo = oligodendrocytes; Ast = astrocytes; Micro = microglia; Endo= endothelial cells; Stromal = stromal cells; Neu = neutrophils; Mono = monocytes / macrophages; NK = natural killer cells, Treg = T regulatory cells; Bnv = naive B cells; Bmem = memory B cells; CD4nv = naive CD4+ T cells; CD8nv = naive CD8+ T cells; CD4mem = memory CD4+ T cells; CD8mem = memory' CD8+ T cells; OR = odd's ratio, SD = standard deviation
[0028] Materials and Methods
[0029] Isolated cell type and glioma datasets
[0030] In the analysis herein the following CNS reference samples are employed: human primary astrocytes from the post-mortem sub-ventricular deep white matter (Astro, n=6)
[0043] , endothelial (Endo, n=12) and stromal (Stromal, n=14) cells from umbilical cord tissue
[0044] , GABAergic neurons (GABA, n=5) and, glutamatergic neurons (GLU, n=5) from the post-mortem dorsolateral prefrontal cortex
[0045] , microglial cells from the post-mortem medial frontal gyrus, superior temporal gyrus, subventricular zone and thalamus (Micro, n=l 8)
[0046] , and oligodendrocytes from the post-mortem Brodmann area 46 (Oligo, n=20)
[0047] , The analysis herein utilizes purified cell types obtained via flow and magnetic sorting methods for the immune cells as described in a previous publication
[0029] . In brief, this included neutrophils (Neu), monocytes (Mono), B naive cells (Bnv), B memory' cells (Bmem), CD4 naive cells (CD4nv), CD4 memory cells (CD4mem), T regulatory cells (Treg), CD8 naive cells (CD8nv), CD8 memory' cells (CD8mem), natural killer cells (NK) as well as artificial mixtures of immune cell types. The healthy and anonymous donors included 41 males and 15 females, with a mean age of 32.2 years (SD = 12.2), and multiple race / ethnicities and further re-classified into broad genetic ancestries, including African (Sub-Saharan), East- Asian, Indo-European, and multiple / admixed. Horvath methylation age as inferred usingDocket No.: 231 / 0021RClient Reference: 2024-026-02Enmix software for any samples in which age was not provided[48, 49], For samples without chromosomal sex data, this was inferred using SeSAMe software
[0050] .
[0031] To integrate all of the reference data to compatible formats, i.e., converting whole genome bisulfite sequencing (WGBS) data to Illumina 450k or EPIC data, the analysis herein utilizes methylLiftover
[0051] . Subsequently, the analysis performs beta-mixture quantile normalization (BMIQ)
[0052] via ChAMP
[0053] , Probes that are known to be cross-reactive, closely associated with single nucleotide polymorphisms, exist on the X or Y chromosome, were on non-CpG sites, or had very large detection p-values
[0054] (p>0.01) were removed from the analysis, resulting in a dataset of 126 samples across 306,466 CpGs.
[0032] This study utilized two large databases of DNAm data on glioma samples. The first dataset originated from the construction of glioma subtype-specific DNAm-based classifiers derived by Capper et al.
[0039] , In brief, the samples utilized from this dataset consist of healthy brain samples across seven brain regions (n=72) as well as glioblastoma (GBM; n=671), grade 2 / 3 astrocytoma (AST-LG; n=172), grade 4 astrocytoma (AST-HG; n=87), and oligodendroglioma (OLG, n=163). The second dataset consisted of data produced by The Cancer Genome Atlas (TCGA), which has data on over 33 human cancer types that are publicly available (https: / / portal.gdc.cancer.aov / ). In brief, the samples utilized from this dataset consist of GBM (n=132), AST (n=139), OLG (n=169) tumor samples alongside a set of samples that could not be confidently assigned to one of these groups (n=216). These datasets were processed using minfl
[0055] via normal-exponential out-of-band background correction with dye bias normalization
[0056] . Probe filtering was done similarly to the cell-type- specific references.
[0033] Tumor diagnosis harmonization to WHO 2021 criteria
[0034] Since the datasets in this study were collected before the WHO 2021 revised classification of adult glioma implementation, previous annotations are harmonized as such.For the Capper et al. dataset, the analysis relies on the molecular classification of the tumors, which comprised four groups: glioblastoma, oligodendroglioma, astrocytoma IDH-mutant and astrocytoma IDH-mutant high grade. For the TCGA dataset, the availability of critical molecular markers is used to infer a WHO 2021 diagnosis in the legacy datasets. The TCGA offers glioma DNA methylation data under two large databases, one labeled “GBM” for glioblastoma and one labeled “LGG” for low-grade glioma. Samples from the GBM dataset that were confirmed IDH wildtype w ere considered GBMs (n=I32). OLG samples were categorized by the presence of IDH mutation and 1 p / 19q co-deletion in the LGG datasetDocket No.: 231 / 0021RClient Reference: 2024-026-02(n=169). AST samples were identified as IDH mutant tumors in the LGG dataset that were not 1 p / 19q co-deleted. These samples were also negative for TERT promoter mutations, EGFR amplification, and 7+ / 10- mutations (n=139). The remaining samples were unable to be confidently mapped to one of the groups (n=216).
[0035] GIMiCC library construction
[0036] L0 was constructed using the InfiniumPurify software
[0057] to identify the 1,000 most informative CpGs for discriminating healthy control brain samples from the glioma tumor samples from the Capper et al. dataset. The dataset was split 75:25 into training and testing; L0 was constructed using only the training set. Four libraries were produced for four subtypes of glioma according to Capper et al. defined molecular subtypes of adult diffuse glioma: GBM, OLG. AST. and AST-high grade (AST-HG). All subsequent layers are derived / developed using an adapted version of the meffll.cell.type.speciflc.methylation function from the meffil package
[0058] , 25 hybrid (hypo- and hyper-methylated) CpGs per cell ty pe per library layer are selected herein, as this was deemed most optimal in previous studies
[0032] ,
[0037] GIMiCC hierarchy was designed to blend previous hierarchies developed for the human brain
[0032] and solid tumors
[0030] , L0 separates the tumor from the non-tumor fraction. Library 1 was built to deconvolve the non-tumor fraction into neuronal cells, glial cells, angiogenic cells, and immune cells. Furthermore, Layers 2A, 2B, and 2C were used to separate these broader cell categories into cell subtypes such as GABA and GLU from the neuronal cells, Astro, Oligo, and Micro from the glial cells, and Endo and Stromal from the angiogenic cells. Layer 2D separates the immune cell fraction into myeloid and lymphoid compartments, and further libraries were used to investigate even more specific subpopulations. Layer 3A splits the myeloid cells into Mono and Neu, and Layer 3B splits the lymphoid cells into NK, B cells, and T cells. Layer 4 is used to separate T cells into CD4 or CD8-positive T cells. Lastly, Layers 5A, 5B, and 5C are used to separate the B cell, CD8T, and CD4T cell compartments into memory and naive subtypes, including Treg for CD4T cells.
[0038] GIMiCC hierarchical deconvolution
[0039] Firstly, the tumor and nontumor proportions are estimated from the probability density7distribution of the L0 CpGs using the InfiniumPurify pipeline
[0057] . The projections for Layer 1 were calculated using the constrained projection / quadratic programming (CP / QP) approach developed by Houseman et al.
[0059] , the sample is deconvolved into neuronal, glial,Docket No.: 231 / 0021RClient Reference: 2024-026-02 angiogenic, and immune cell proportions. These proportions are weighted by the proportion of non-tumor cells derived in L0 to develop the final deconvolution for Layer 1 into the tumor, neuronal, glial, angiogenic, and immune cell proportions. This process is then iterated for the rest of the hierarchical structure; project the cell proportions of layer n and weigh the estimates by the proportion of the parent node in the n-1 layer. Deconvolving to the deepest layer results in the full 18-cell type deconvolution.
[0040] Biological enrichment analyses
[0041] To map each CpG to associated genes, the analysis utilizes theInfiniumMethylation BeadChips Annotation file
[0060] . The UCSC Genome Browser was used to investigate the CpG location relative to associated gene(s)
[0061] ,
[0042] Gene set enrichment analysis was performed with the goMeth function from missMethyl software
[0062] . The input to the software is a list of differentially methylated CpGs and a total query CpG list. The output is a set of gene ontology (GO) terms with corresponding p-values for the test of enrichment. FDR values, generated with the Benjamini and Hochberg procedure, were used to select significantly enriched GO terms (FDR < 0.05).
[0043] Enrichment for relation to CpG islands was done with independent logistic regression models to calculate the odds of a query set of CpGs being in open sea, north shelves, north shores, islands, south shores, or south shelves. Similarly, enrichment in gene region is tested to calculate the odds of a query set of CpGs being in TSS1500, TSS200, 5’UTR, 3’UTR, 1stexon, or gene body regions.
[0044] Benchmarking GIMiCC with other DNAm deconvolution methods
[0045] methylCIBERSORT
[0041] and methylResolver
[0040] were implemented as described. The glioma signature matrix was utilized for methylCIBERSORT. The default blood signature matrix was utilized for methylResolver. These methods, alongside GIMiCC, were used to deconvolve a set of artificial mixtures of immune cells (GSE182379)
[0029] , The error was quantified as the difference in the true versus the predicted proportions. Cell types were aggregated into five categories to equate similar cell types across methods: granulocytes (Gran), NK, Mono, T cells, and B cells.
[0046] Epigenome-wide association studies (EWAS)
[0047] EWAS analysis was done using the minfi and Umma software[55, 63], Linear regressions are utilized herein to identify differentially methylated CpGs (DMCs) between two groups. Cell-type adjusted EWAS analysis includes the proportions of specified cell types, mean scaled and centered, as covariates in the linear model. DNAm proportions (0Docket No.: 231 / 0021RClient Reference: 2024-026-02 values) were converted to M-values by computing the logit of the (3 values in base 2. An empirical Bayes method was used to normalize CpG-wise residual variance. The Benjamini- Hochberg False Discovery Rate-FDR procedure was used to adjust for multiple hypothesis testing. Prior to the analysis, a DMC is defined by having an effect size (|A(3|) larger than 0.3 and an FDR less than 0.05.
[0048] Survival analysis
[0049] Cox-proportional hazards models were used to quantify the effect of cellular composition on survival in the TCGA dataset. Univariable and multivariable models were used; age and sex were included in the multivariable model. Additionally, separate models were created for each tumor type. The proportional hazard assumption was tested in each model; in cases where the assumption is not upheld in the multivariable model (adjusting for age and sex), results were derived from the univariable model. To visualize the results, the dataset is split based on the median value for a given cell type and compared the survival outcomes in the subpopulation with values above the median or “hof’ for a given cell type to the subpopulation with values below the median or "cold" for a given cell type.
[0050] Results
[0051] GIMiCC hierarchical tree development and library construction
[0052] It is desirable to develop a DNAm-based deconvolution method that can resolve the glioma microenvironment as a way to retrieve this information at epidemiological scales. Recent advances have allowed for the development of DNAm-based deconvolution libraries for whole blood (FlowSorted.BloodExtended.EPIC)['19}, brain samples (FfzB£’Z))
[0032] , and solid tumors (HiTIMED)[30\. The construction of these libraries required the collection of cell-type specific reference DNAm data from isolated CNS and immune cell types. Thus, these datasets are combined to develop cell-type-specific reference profiles for significant expected cell types in glioma. These datasets are aggregated to create a database of 132 samples for 17 cell types, including GABAergic (GABA) and glutamatergic (GLU) neurons, oligodendrocytes (Oligo), astrocytes (Ast), microglia (Micro), endothelial cells (Endo), stromal cells (Stromal), neutrophils (Neu), monocytes / macrophages (Mono), natural killer cells (NK), T regulatory cells (Treg), naive B cells (Bnv), memory B cells (Bmem), naive CD4+ T cells (CD4nv), naive CD8+ T cells (CD8nv), memory CD4+ T cells (CD4mem), and memory CD8+ T cells (CD8mem). When aggregating these data, it is noted that the age distribution for microglia samples was much higher than the other samples. Because of the known epigenetic alterations associated with age in microglia and other phagocytic immuneDocket No.: 231 / 0021RClient Reference: 2024-026-02 cells[64, 65], the samples are stratified into those coming from younger individuals (< 75) and those from older individuals (> 75). Because there was significant epigenetic variation between these two groups of samples, only microglial samples from younger individuals are utilized in the deconvolution library development.
[0053] Hierarchical deconvolution resolves major cell t pes in more shallow layers and takes the results of those layers to scale the output of deeper layers. This approach forces the algorithm herein to leverage shared lineage marks across cell types independently of the markers used to resolve unique subsets, allowing us to resolve more cell types than previously
[0030] , The hierarchical tree for CNS cells from HiBED is combined with the hierarchical tree for the tumor-immune microenvironment from HITIMED to generate a 6- layered hierarchical tree consisting of 17 cell types and a tumor cell proportion (Fig. 1). Layer 0 (L0) will first deconvolve the sample into the tumor and non-tumor cell proportions. In Layer 1 (LI), the non-tumor cell fraction is deconvolved into four broad categories: neuronal, glial, angiogenic, and immune. Subsequent layers then deconvolve each one of these categories into more specific cell types.
[0054] To generate the deconvolution library for L0, the Capper et al. dataset
[0039] is utilized, which consists of DNAm-based classes of glioma in addition to healthy control samples across seven brain regions. The dataset is split into training and testing datasets in a 3: 1 ratio, and each L0 was derived only using the training set. An L0 library specific to distinct tumor types is generated / computed according to the most recent WHO CNS5 criteria: glioblastoma (GBM), oligodendroglioma (OLG), IDH mutant astrocytoma (AST), and highgrade IDH mutant astrocytoma (AST-HG). Prototype versions of GIMiCC that included a pan-glioma L0 were outperformed by the glioma-subtype-specific approach (data not shown). Each library consists of the 1,000 most informative differentially methylated CpGs (DMCs) when comparing the tumor samples to the healthy control (Fig. 2A). It is recognized that the libraries for IDH mutant tumors were mainly hypermethylated in the tumor samples.Specifically, 99.7% of the OLG, 97.9% of the AST-HG, and 99.6% of the AST L0 CpG sites were hypermethylated in the tumors. In contrast, only 46.5% of the tumors' GBM L0 CpG sites were hypermethylated. In a subset of the Capper et al. dataset, the authors provided a tumor purity estimate using the Cancer Genome Atlas (TCGA) pan-glioma DNAm model[39, 66], It is recognized that GIMiCC tumor purity estimates correlate with the TCGA-basedDocket No.: 231 / 0021RClient Reference: 2024-026-02 estimates (Figs. 2B-2E). A sensitivity analysis is additionally performed, in which the AST LO library is derived using ten random splits of training and testing data for the tumors and controls. For the ten folds, 541 of the 1000 L0 CpGs were consistently identified in each iteration. Additionally, the tumor purity estimates in the test samples are shown to be consistent with each iteration.
[0055] To generate the deconvolution libraries for all subsequent layers, the limma approach is employed in the Meffll software|58J to select the top 25 hyper and hypomethylated cell-type-specific DMCs per cell type for each layer using linear models. Minimal overlap occurs betw een the CpGs utilized in each L0 library' compared to the other deconvolution layers.
[0056] The implementation of GIMiCC utilizes these derived libraries in a hierarchical fashion. Firstly, the tumor and nontumor proportions are estimated in the distribution of the L0 CpGs using the InfiniumPurify pipeline
[0057] . For all subsequent layers, the constrained projection quadratic programming approach
[0059] was used to project the proportions of the cell types in the subsequent layer by weighing their projections by the values of the previous projection. For instance, the outputs of using Layer 2A are weighted by the result of the neuronal cell proportion from using Layer 1. In this manner, this approach is iterates and estimates the proportions of all 18 cell ty pes.
[0057] Biological context of CpGs selected in GIMiCC libraries.
[0058] After developing the libraries for GIMiCC. which genome regions were utilized and connect these patterns to known biological functions is ascertained. For instance, VAMP 2 is highly expressed in the brain as it is involved in synaptic vesicle fusion
[0067] , A CpG associated with VAMP2 is included in the Layer 1 library' and is hypomethylated in neurons compared to other cell types (Fig. 3A). Similarly, hypomethylation of a CpG near TREM2 in microglia is used in Layer 2 of the deconvolution (Fig. 3B). TREM2 is a marker of microglia and macrophages[68, 69]; however, on the epigenetic level at this site, it seems to be specific to microglia. Lastly , a CpG near HLA-DOB, a major histocompatibility complex gene highly expressed on B cells
[0070] , is used in Layer 3 of the deconvolution (Fig. 3C).
[0059] The compatibility of GIMiCC is examined across the three most recent Illumina DNAm profiling arrays: 450k. EPIC, and EPICv2. It is identified that most, but not all, CpGs were conserved across platforms. To determine if this would impact the results of GIMiCC, the analysis herein compares the deconvolution results of the Capper el al. datasetDocket No.: 231 / 0021RClient Reference: 2024-026-02 with the entire probe set and with the probes only available on EPICv2. As such, these results are strongly correlated, with a median correlation coefficient of 0.99 and a minimum of 0.96.
[0060] To get a pathway-level perspective of the CpGs utilized in the GIMiCC libraries, a gene set enrichment is performed with the missMethyl software
[0062] , For all LO libraries, the significantly enriched pathways and ontologies involved the plasma membrane and cell adhesion, whereas the GBMLO library was also enriched with immune-related processes. For the libraries associated with the other layers of deconvolution, the enriched pathways and ontologies involved immunological pathways and processes.
[0061] Next, the analysis tests for enrichment in the GIMiCC library CpGs for contexts of the CpG island methylation region (island, shore, shelf, and open sea)
[0071] . Opensea CpGs were most enriched in the GBM LO library, OR = 1.3, 95% CI [1.2-1.5] (Fig. 3D). However, CpGs within CpG islands were most enriched in the AST LO library, OR = 3.7 [3.2- 4.2], AST-HG LO library, OR = 7.7 [6.6-8.9], and OLG LO library, OR = 4.1 [3.6-4. 7], (Fig. 3E). The CpGs in cell-type specific layers (L1-L5) were enriched on open sea CpGs, OR = 1.5 [1.4-1.8], andwere unlikely to be on CpG Islands. OR = 0.59 [0.51-0.69], (Fig. 3F).
[0062] Lastly, the analysis tests for enrichment of GIMiCC library CpGs within gene regulatory regions (TSS1500, TSS200, 5 ’UTR, 1stexon, Body, 3 ’UTR). All LO libraries are enriched with CpGs on the 1stexon of genes; GBM OR = 1.7 [1.4-2. 1], AST OR = 1.5 [1.2- 1.8], AST-HG OR = 1.7 [1.4-2.1], and OLG OR = 1.4 [1.2-1. 7], (Figs. 3G-3H). There was enrichment for CpGs within gene bodies in the remaining L1-L5 layers OR = 1.5 [1.3-1. 7](Fig. 31).
[0063] GIMiCC tumor type specificity
[0064] Using the test subset of samples, the tumor purity is projected using all four LO libraries and compared these estimates (Fig. 4A). The AST-HG, AST, and OLGLO libraries yield similar estimates. Additionally, using an AST-HG, AST, or OLG LO to deconvolve GBM tumors results in lower tumor purity estimates. Similarly, the distribution of purity estimates is much lower using a GBMLO to deconvolve OLG or AST tumors. Interestingly, AST-HG tumors deconvolved with any of the four LO libraries yield similar estimates.
[0065] Qualitative and semi-quantitative validation of GIMiCC
[0066] Low tumor purity occurs when GIMiCC was applied to non-tumor samples(Fig. 4A). Additionally, when the projections are performed on the other cell types, cell proportion estimates reflect the physiolog}- of the tissue type, i.e., a high proportion of oligodendrocytes in the white matter (WM) samples, the highest levels of neurons in corticalDocket No.: 231 / 0021RClient Reference: 2024-026-02 samples (HEMI), and high levels of angiogenic and immune cells in more vascularized regions of the brain such as the anterior pituitary gland (ADENOPIT), pineal gland (PINEAL) and the pons (PONS).
[0067] The ability of GIMiCC to detect immune infiltration in brain tissue can be further validated by the analysis herein. As such, additional samples are employed from the Capper et al. dataset that are annotated to have ‘“high immune infiltrate"’ and “low-yield” (Figs. 4C-4D) For the former, these samples have a high granulocytic infiltration associated with necrosis or intense hemorrhage. The latter is comprised of samples that had very low tumor cell content. GIMiCC correctly predicts high fractions of immune cells within these samples compared to the healthy control samples.
[0068] To validate GIMiCC with an independent dataset, the analysis employs the DNA methylation data from the Cancer Genome Atlas (TCGA, https: / / www.cancer.gov / tcga). Genetically confirmed tumors are selected and harmonized their diagnoses to WHO 2021 criteria. Using GIMiCC, the analysis estimates these samples’ cellular composition, including the immune microenvironment (Fig. 5). The OLG tumors had much lower immune infiltration than the other tumor types, and the GBM tumors had higher levels of infiltration. The predominant infiltrates were myeloid, including microglia, neutrophils, and monocytes.
[0069] To further validate GIMiCC, estimates of tumor purity are compared to the TCGA a consensus purity estimate (CPE); this estimate integrates predictions based on gene expression data, somatic copy-number data, DNAm data, and immunohistochemistry to generate a single estimate for tumor purity
[0072] . Tumor purity estimates are highly correlated with the TCGA CPE (Figs. 6A-6C). The immune composition output of GIMiCC is compared to the ESTIMATE immune score, which uses expression data to identify the level of immune cells in the TCGA tumors
[0073] . The samples deemed highly inflamed by GIMiCC have an elevated ESTIMATE immune score compared to the others (Fig. 6D). This relationship is emphasized further when incorporating microglia into the immune composition (Fig. 6E)
[0070] The ideal experiment to test the performance of DNAm-based deconvolution methods is to use a validated sorting method such as flow cytometry to isolate each cell of interest and to develop artificial mixtures of know n proportions to deconvolve them computationally. No dataset as such is presently available. Thus, a dataset of artificial DNA mixtures of immune cells from human blood is utilized to test the ability of GIMiCC to deconvolve the immune microenvironment. GIMiCC is compared to two validatedDocket No.: 231 / 0021RClient Reference: 2024-026-02 deconvolution methods: methylCIBERSORT
[0041] and methylResolver
[0040] . GIMiCC performed at the same level as these other methods.
[0071] Determining the impact of cellular heterogeneity on DNAm-based tumor classification
[0072] Previous work has shown that glioma subtypes can be identified via distinctDNAm patterns; however, it is possible that these classifiers may be confounded by differences in cellular composition across tumor types when constructing tumor classifiers
[0037] . To test this, the analysis performs epigenome-wide association studies (EWAS) comparing different tumor ty pes from the Capper et al. dataset to generate lists of differentially methylated CpGs (DMCs) that can be used for tumor classification. To identify the effects of controlling for cellular heterogeneity’, five different EWAS models are performed: 1) unadjusted for cell type, 2) adjusted only for the tumor fraction, 3) adjusted for broad cell categories {Tumor, Angiogenic, Glial, Immune} 4) adjusted for the immune microenvironment {Tumor, Angiogenic, Astrocyte, Microglia, Oligodendrocyte, Myeloid, Lymphoid} and 5) adjusted for a deeper layer deconvolution of the immune microenvironment {Tumor, Angiogenic, Astrocyte, Microglia, Oligodendrocyte, Tcell, Bcell, NK, Neu, Mono}. Significant heterogeneity7in cellular composition occurs within tumor ty pes in the Capper et al. dataset. In an EWAS comparing GBM to AST gliomas, more than 30,000 DMCs occur in the unadjusted model; however, there was no significant change to the number of identified DMCs when adjusting for cell type (Figs. 7A-7C). This analysis is repeated by comparing other types of tumors to each other and found similar results. However, identifying CpGs for tumor classification is not improved via controlling for cellular composition.
[0073] Identifying compositionally independent alterations in DNAm associated with highly infiltrative glioma
[0074] Because a high variation of immune infiltration occurs within tumor types, the analysis aims to identify DNAm patterns associated with increased inflammation in the Capper et al. GBM tumor samples. Thus, an EWAS analysis comparing the highest and low est deciles of immune infiltrated samples is performed. Using the five models described previously, is appears that controlling for cell type in this analysis dramatically' reduces the number of DMCs identified (Figs. 7D-7F). The same trend occurs in the other tumor.
[0075] Assessing the clinical relevance of GIMiCC estimatesDocket No.: 231 / 0021RClient Reference: 2024-026-02
[0076] Lastly, the analysis determines whether GIMiCC-derived immune cell proportions were associated with patient survival. Cox proportional hazard models are, thus, used to test the survival effects in the TCGA samples. The analysis constructs univariable and multivariable models for each cell type to test the relationship between immune cell level and patient survival after 5 and 10 years. The multivariable model is adjusted for age and sex. The models are also stratified by tumor type as including tumor type in a global model produced large violations of proportional hazards assumptions.
[0077] This analysis showed that higher levels of angiogenic cells were most strongly associated with w orse survival outcomes in OLG and AST tumors. For GBM, the presence of many immune cell types, including CD8nv, Mono, and Microglia, is associated with worse survival. Notably, microglia are associated with survival but only in the multivariable model adjusting for age and sex (HRuni = 1 .06 [1 .00-1.13] compared to HRmuiti = 1.16 [1.08-1.25]).
[0078] Discussion
[0079] In view of previous DNAm-based deconvolution libraries developed for whole blood
[0029] , brain
[0032] . and solid tumors
[0030] , the analysis aims to integrate the cell references to have a set of deconvolution libraries for human gliomas. Cell-type-specific methylation patterns occur ini 8 unique cell types and used these signatures to predict the composition of glioma samples hierarchically across six total layers. This new method was validated in an excluded subset of the training data (Capper et al.) and an independent dataset (TCGA) and compared to other DNAm-based deconvolution methods. GIMiCC recapitulates the known biology of glioma inflammation, and its predictions can be used to determine more information about the diversity of glioma immune microenvironments and to perform celltype adjusted EWAS analyses.
[0080] GIMiCC prediction of tumor purity
[0081] Four versions of L0 for subclasses of adult diffuse glioma are defined byDNAm profile (AST, AST-HG, OLG, and GBM). The analysis can create glioma subtypespecific deconvolution, as it produces more accurate results than a pan-cancer tumor purity estimation.
[0030]
[0082] Mutations to IDH result in metabolic reprogramming of the cell, causing an accumulation of D-2-hydroxy glutarate (D-2-HG), which can disrupt the demethylation of histones and DNA
[0074] , This has been linked to a distinct hypermethylated phenotype, particularly in CpG islands[75, 76], The results herein show that the CpGs used in L0 for the IDH mutant gliomas were all hypermethylated, and an enrichment analysis showed highDocket No.: 231 / 0021RClient Reference: 2024-026-02 enrichment for L0 CpGs to be on CpG islands. This further validates the approach herein as the LO libraries reproduce the known biology of these tumors.
[0083] Using the test set samples from the Capper et al. dataset, it is shown that the LO libraries predict low tumor purity when applied to the incorrect tumor type. However, subsets of samples exhibit high purity regardless of the library utilized. This may reflect the nature of gliomas to exhibit a heterogeneity of tumor cell types that may not be fully captured by a single cell-type profile[77-79J.
[0084] The resources of the TCGA database are highly useful to validate GIMiCC LO predictions. The CPE-predicted tumor purity is a well-established benchmark for tumor purity' estimates in this database
[0072] . This method integrates tumor purity estimation via four platforms: RNA expression data
[0073] , copy number alteration data
[0080] , DNAm data
[0072] , and immunohistochemistry
[0072] , The results herein are highly correlated with this estimate, with the highest correlation for GBM LO. The IDH mutant LO libraries tended to underpredict tumor purity compared to CPE.
[0085] GIMiCC prediction of immune composition
[0086] Layers 1 through 5 of GIMiCC are used to predict the remaining composition of the tumor samples. The analysis can test the accuracy of this prediction with artificial mixtures of immune cells constructed independently. Compared to other methods of DN Am- based deconvolution, GIMiCC performs at the same level, if not better, than the other approaches. Additionally, the performance of GIMiCC is stable using the probes available on the next iteration of the Illumina DNAm sequencing platform, EPICv2
[0081] ,
[0087] Some of the CpGs included in GIMiCC are associated with genes expected to be functionally active in only particular cell types. Firstly, a CpG associated with VAMP2 was included in Layer 1 to distinguish neurons from other cell types. VAMP 2 plays a role in vesicle fusion at the synapse during neurotransmission
[0067] ; thus, the analysis can expect to find the region around this gene hypomethylated in neurons. Secondly, a CpG is highlighted near TREM2 that is included in Layer 2 to separate microglia from other cell types. TREM2 is expressed on microglia and tissue-resident macrophages and is a significant target for current research about microglial immunology and neuroinflammation
[0082] . Although peripheral monocytes / macrophages express TREM2
[0083] , this particular CpG is only hypomethylated in microglia. The results herein suggest that TREM2 may be differentially regulated epigenetically across peripheral (monocyte) versus local (microglia) cell types. Lastly, a B cell-specific hypomethylation of a CpG near HLA -DOB known to be overexpressed on B cellsDocket No.: 231 / 0021RClient Reference: 2024-026-02 is identified herein
[0070] . These results, taken together with the gene ontology data, show that the libraries used in GIMiCC are connected to known biological processes of the cell types.
[0088] GIMiCC is additionally validated, using it on inflammatory and low-yield control samples from Capper et al. Verbatim descriptions of these groups can be found in the original publication! 39|. Both groups of samples were noted to be high in leukocytes, either due to necrosis, hemorrhage, or tumor infiltration. GIMiCC can detect the immune signal in these samples compared to true healthy controls, further validating the method.
[0089] Our last form of qualitative validation comes from comparing the results herein to what is currently known about the immunological heterogeneity across brain tumors. Across methods including single-cell RNA-sequencing[20, 84-88], imaging mass spectroscopy
[0087] , flow cytometry[89, 90]. and mass cytometry
[0091] , there is significant variation in the immune composition across gliomas, which are observed in both Capper et al. and TCGA datasets. Even within distinct tumor types, there is known heterogeneity [92, 93], One potential source of this variation in these datasets is tumor progression, as it has been shown that the tumor microenvironment alters over time, increasing with signatures of oligodendrocytes and myeloid cells
[0088] . Another potential bias could be in which section of the tumor was used for DNAm array, as the tumor core and periphery have distinct immune microenvironments[87, 90],
[0090] Notably, the analysis has also replicated known findings about the differences in immune microenvironments between tumor subtypes. For example, glioblastoma samples were generally more infiltrative than the other tumor types, consisting primarily of infiltrating monocytes and T cells that likely contribute to local immunosuppression[20, 84, 87, 94], Microglia is the largest immune microenvironment contributor in the IDH mutant tumors, whereas in the GBMs high monocyte infiltration occurs. This aligns with previous findings indicating bone marrow-derived macrophages are increased in higher grade and immunosuppressed tumors, but microglia are present at all grades[84, 90, 95], A limitation of GIMiCC is the ambiguity of tumor-associated macrophages and whether these cells are classified as "‘microglia’" or ‘“monocytes.” The ability of cells to preserve epigenetic markers of origin has previously been identified[96, 97], This can occur because these cells derive from two distinct lineages (microglia from yolk-sac progenitors versus monocytes from myeloid progenitors in the bone marrow)
[0098] , distinct epigenetic marks would exist that maintain information about the source progenitor. Thus, the analysis can expect that monocytes that infiltrate and differentiate into tumor-associated macrophages in the brainDocket No.: 231 / 0021RClient Reference: 2024-026-02 would be captured in the monocytes (‘‘Mono’') proportion of GIMiCC,' however, this has yet to be shown.
[0091] When looking at the adaptive immune cells, higher levels of T and B cells occur in GBM tumors compared to the other subtypes, which have been replicated in other studies[20, 85], In particular, other’s findings of higher levels of mature and regulatory' T and B cells in GBMs[85, 91, 99, 100] are corroborated herein. Note that these cell types are generally only detectable in a subpopulation of patients within each tumor type, which may explain the variability in the efficacy of certain immunotherapies.
[0092] Impact
[0093] Using GIMiCC, two main hypotheses about the cellular heterogeneity of glioma and its impact on DNAm analyses can be addressed herein.
[0094] Firstly, the analysis examines whether the composition of the microenvironment could influence the DNAm-based classification of glioma. Others have shown that resolving the non-tumor fraction of the sample aids in the correct assignment of tumor class when using DNAm biomarkers
[0037] , This is tested by performing EWAS analyses between distinct tumor types that did and did not control the tumor’s cellular heterogeneity’. Adding this information to the model does not change the calling of DMCs. Thus, for strong molecular alterations, such as IDH mutations, DNAm-based classifiers can select CpGs that are less likely to be impacted by compositional variation.
[0095] Secondly, the analysis attempts to identify specific DNAm alterations associated with higher immune infiltration. Similarly, multiple EWAS analyses are performed, comparing high to loyv infiltrating tumors yvhile adjusting or not adjusting for cell ty pe. In this case, adjusting for cell heterogeneity greatly reduces the number of DMCs being called. Without controlling for cell proportions, many of the DMCs are associated with the proportion of a given immune cell rather than the distinct biological alteration that is causing the downstream recruitment of leukocytes
[0101] . By running a model that can fully adjust for this, the analysis can more readily point to the molecular alterations associated with recruitment without the confounding of having more cells in the “high infiltrate” group.
[0096] Lastly, GIMiCC is utilized to identify levels of cell proportions that are associated with 5 and 10-year-long survival. Notably, higher levels of stromal and / or endothelial cell proportions in the tumor are associated with w orse survival in OLG and AST, but not GBM. Marks of angiogenesis and microvascular proliferation are included in the diagnostic criteria of GBM; thus, this association was not observed, possibly due to loyverDocket No.: 231 / 0021RClient Reference: 2024-026-02 variance in the presence of angiogenic cells (Fig. 5). Nonetheless, this finding has been replicated in independent modalities in the TCGA set of tumors[102, 103], Previous work in solid tumors has also show n a similar effect in head and neck squamous cell carcinoma, stomach adenocarcinoma, and thyroid carcinoma; higher levels of angiogenic cells in these tumors were associated with worse 5-year survival
[0030] . Other work on glioma has replicated this finding
[0102] , including the observation that increased Angiogenin (A NG) expression correlates with worse survival outcomes and more immune infiltration) 104], These results suggest that therapies targeted tow ard preventing angiogenesis within these tumors may improve survival by restricting nutrients or immunosuppressive cells from the tumor.
[0097] Note that the first is that cell-type specific references were aggregated from various sources across many different DNAm profiling platforms. Also, proper implementation of GIMiCC typically requires prior knowledge of tumor classification. Using the wrong library' to deconvolve a glioma subtype can potentially deflate the tumor purity estimation. Thus, any misclassification of tumors can greatly impact the interpretation of the results, which may also explain the presence of outliers in the datasets that were assessed.
[0098] Furthermore, GIMiCC was built using data where tumor type was annotated via DNAm profile as defined in the Capper et al. dataset. This can limit resolution in tumor grade being able to distinguish grade 2 / 3 / 4 astrocytoma individually. As additional annotations for DNAm data are developed, novel, more accurate classification schemes can be incorporated readily into the analysis pipeline herein.
[0099] Moreover, the cell types included in GIMiCC are built off pre-identified cell populations w ith established markers and isolation protocols. As knowledge of tumor immunology grows and the analysis herein can develop profiles for more specific cell types (such as myeloid-derived suppressor cells), these can be incorporated into the analysis algorithm to efficiently and systematically screen various large databases for these new cell types.
[0100] Lastly , GIMiCC is validated herein using publicly available datasets with silver-standard purity estimates and artificial mixtures of immune cells. It is expected that direct matching of DNAm data to other modalities for validation (single-cell RNA sequencing, chromatin accessibility, or histology) will bias estimates of cell proportion for tw o reasons. One is that the microenvironment cells that get captured will vary' across separate regions of the tumor or serial slices of the tumor
[0020] . The second reason is the lack of overlap in feature spaces between DNAm and these other modalities. Single-cell approaches do notDocket No.: 231 / 0021RClient Reference: 2024-026-02 guarantee the detection of marker genes across all cell types, which may bias cell ty pe identification. When collecting single-cell information, technical and biological variations due to cell cycle and cell states may increase the chance of classical measurement errors, biasing the results to the null and reducing precision. In contrast, deconvolution approaches assume that the average of the weighted signal represents a layer of a cell hierarchy. This increases the power of the analysis but changes the estimates, introducing what is known as a Berkson error (increased variability but less systematic biases when used correctly)ll05J. The strongest form of validation of GIMiCC would require the development of artificial mixtures of DNA from each cell ty pe across a robust number of patients, including neuronal, immune, and tumor cell types. As more specific DNAm profiles become available in the literature, the analysis platform herein is flexible for the addition of updated cell types to the library or calibration to a true validation set.
[0101] GIMiCC Web-based Application and Computing Environment
[0102] The computing environment and associated tool are shown in the computing and processing arrangement 900 of Fig. 9. which shows a generalized computing environment / system 900 for performing the tasks of the system and method herein. The system 900 includes at least one computing device 910 in the form of a general purpose computer (e.g., a PC, laptop, tablet, server, cloud computing arrangement, etc.) that includes an interface screen (e.g., touchscreen) 912, and various user interface devices (e.g. keyboard 914 and mouse 916). The computing device 910 instantiates a process(or) 920 that operates the data handling and diagnostic tasks relevant to the GIMiCC (and related processes) herein. The computing device 910 receives patient tumor (and other related) data 930 — via manual input, network based-inputs from patient records and / or from appropriate medical devices.The computing device 910 is further connected, via an appropriate wired and / or wireless link to a public and / or private data network (such as the Internet) 940 that allows access to one or more data stores 950 having relevant DNA methylation data 953 related to tumor types, the various cell ty pes (both healthy and cancerous as described above), etc. as described herein. Access consists of requests 954 for particular information, which result in the return of relevant data 956 for use in the process(or) 920. The date store 950 can be constructed using any appropriate data structure, including well-known database arrangements, and can be distributed among a plurality of data stores managed by one or multiple entities. Requests 954 are directed to the appropriate store based upon a known addressing scheme.Docket No.: 231 / 0021RClient Reference: 2024-026-02
[0103] The process(or) 920 can be arranged in any acceptable configuration clear to those of skill, and the functional processes / ors or modules depicted are by way of non-limiting example. The process(or) 920 includes a data access process(or) 922 that handles patient data on tumors and user inputs to issue appropriate requests 954 to the data store 950 and retrieve relevant data 956. The data is used by the analysis process(or) 924 to perform a relevant processing using (e.g.) machine learning and trained classifiers to operate on presented data. This can be facilitated by appropriate comparison routines, including those supported by commercially available (or custom) Artificial Intelligence (Al) based systems, including, but not limited to Neural Networks, Convolutional Neural Networks (CNNs), and similarly functioning systems. The results of the analysis can be presented as a diagnosis with associated data on the condition by a diagnostic process(or) 926 using various stored and / or derived (via programmed algorithms / processes) that interoperate with results from the analysis process(or) 924.
[0104] By way of further background, Python Dash (See WorldWideWeb URL address https: / / dash.plotly.com / dash-html-components / cite) and Heroku (See WorldWideWeb URL address https: / / www.heroku.com / ) are used to develop a user-friendly web-based GIMiCC tool. Python Dash is a frame ork developed by Plotly, which is based on Python, and used for building and deploying data applications with a customized interface. Heroku, which is (e g.) a cloud platform, is next employed to host and deploy the GIMICC web application. The GIMiCC web application running in the process(or) 920 contains two major parts. The first part is a user guide. Users should follow the instructions to finish the prediction process. An exemplary7input data csv file is available to the user for demonstration. The second part includes data upload, model running, and output download. After constructing the input data as instructed, users (operating the interface 912. 914 and 916) can either click the data upload box to choose the file or drag the file to the box from the local end to upload the input data. Then the algorithm will automatically compute the output and show up on the right side of the panel. Users can also dow nload the output result as a csv file by clicking the export box.
[0105] In operation, the application can include appropnate APIs that allow it to interoperate with the operating system and / or web brow ser of the users computing device(s).Appropriate security7functionality can be installed in the user’s device (e.g. SSL-based communication frameworks) the ensure confidentiality of patient and user information. Availability of the software application (e.g. for download and / or installation) can be providedDocket No.: 231 / 0021R Client Reference: 2024-026-02 via an appropriate web-based distribution source that operates a server for such downloads (e.g. Google Play, Apple Store, etc.). The revenue model, if applicable, can be based on a subscription service and / or use of validated credentials that are established for each user and employed when logging into the application to enter, manipulate or view data. Such arrangements should be clear to those of skill in the art.
[0106] More generally, it is noted that commonly assigned, published International Application No. WO 2024 / 187092, entitled SYSTEM AND METHOD FOR HIERARCHICAL TUMOR ARTIFICIAL INTELLIGENCE CLASSIFIER TRACES TISSUE OF ORIGIN AND TUMOR TYPE USING DNA METHYLATION, filed March 8, 2024 describes a system and method that employs similar techniques and computing processors / es to carry out a diagnostic task on tumor cells. This application, and its related family of applications, is expressly incorporated here by reference as useful background information.
[0107] Conclusion
[0108] have introduced GIMiCC, a computational tool that allows users to calculate the cellular composition of glioma samples from DNAm data. This method is valid and benchmarked to previous DNAm deconvolution methods while increasing the output’s resolution and developing glioma-subtype specificity. As DNAm data is being collected more frequently for tumor identification, the analysis herein can continue to optimize GIMiCC and implement it to understand more about the immunological interactions in glioma and potentially add information regarding optimizing treatment strategies.
[0109] References[1] A. M. Molinaro, J. W. Taylor, J. K. Wiencke, and M. R. Wrensch, "Genetic and molecular epidemiology of adult diffuse glioma," (in eng), Nat Rev Neurol, vol. 15, no. 7, pp. 405-417, Jul 2019, doi: 10.1038 / s41582-019-0220-2.[2] Q. T. Ostrom, G. Cioffi, K. Waite, C. Kruchko, and J. S. Barnholtz-Sloan, "CBTRUS Statistical Report: Primary Brain and Other Central Nervous System Tumors Diagnosed in the United States in 2014- 2018," (in eng), Neuro Oncol, vol. 23, no. 12 Suppl 2, pp. iiil-iiil05, Oct 05 2021, doi: 10.1093 / neuonc / noab200. [3] R. Chen, M. Smith-Cohn, A. L. Cohen, and H. Colman, "Glioma Subclassifications and Their ClinicalSignificance," (in eng), Neurotherapeutics, vol. 14, no. 2, pp. 284-297, Apr 2017, doi: 10.1007 / sl3311- 017-0519-x.[4] D. N. Louis et al., "The 2021 WHO Classification of Tumors of the Central Nervous System: a summary," (in eng), Neuro Oncol, vol. 23, no. 8, pp. 1231-1251, Aug 02 2021, doi: 10.1093 / neuonc / noabl06. [5] W. C. o. T. E. Board, World Health Organization Classification of Tumours of the Central NervousSystem, 5th Edition ed. International Agency for Research on Cancer, 2021.[6] B. Lv et al., "Immunotherapy: Reshape the Tumor Immune Microenvironment," (in eng), Front Immunol, vol. 13, p. 844142, 2022, doi: 10.3389 / fimmu.2022.844142.Docket No.: 231 / 0021R Client Reference: 2024-026-02[7] X. X. Wang et al., "Immune Gene Signatures and Immunotypes in Immune Microenvironment Are Associated With Glioma Prognose," (in eng), Front Immunol, vol. 13, p. 823910, 2022, doi: 10.3389 / fimmu.2022.823910.[8] B. T. Himes, P. A. Geiger, K. Ayasoufi, A. G. Bhargav, D. A. Brown, and I. F. Parney, "Immunosuppression in Glioblastoma: Current Understanding and Therapeutic Implications," (in eng), Front Oncol, vol. 11, p. 770561, 2021, doi: 10.3389 / fonc.2021.770561.[9] I. Burghardt et al., "Endoglin and TGF-fS signaling in glioblastoma," (in eng), Cell Tissue Res, vol. 384, no. 3, pp. 613-624, Jun 2021, doi: 10.1007 / s00441-020-03323-5.
[0010] J. Qian et al., "TLR2 Promotes Glioma Immune Evasion by Downregulating MHC Class II Molecules in Microglia," (in eng), Cancer Immunol Res, vol. 6, no. 10, pp. 1220-1233, Oct 2018, doi: 10.1158 / 2326- 6066.CIR-18-0020.
[0011] R. S. Andersen, A. Anand, D. S. L. Harwood, and B. W. Kristensen, "Tumor-Associated Microglia and Macrophages in the Glioblastoma Microenvironment and Their Implications for Therapy," (in eng), Cancers (Basel), vol. 13, no. 17, Aug 24 2021, doi: 10.3390 / cancersl3174255.
[0012] M. W. Yu and D. F. Quail, "Immunotherapy for Glioblastoma: Current Progress and Challenges," (in eng), Front Immunol, vol. 12, p. 676301, 2021, doi: 10.3389 / fimmu.2021.676301.
[0013] U. Sener, M. W. Ruff, and J. L. Campian, "Immunotherapy in Glioblastoma: Current Approaches and Future Perspectives," (in eng), Int J Mol Sci, vol. 23, no. 13, Jun 24 2022, doi: 10.3390 / ijms23137046.
[0014] P. C. Gedeon et al., "Checkpoint inhibitor immunotherapy for glioblastoma: current progress, challenges and future outlook," (in eng), Expert Rev Clin Pharmacol, vol. 13, no. 10, pp. 1147-1158, Oct 2020, doi: 10.1080 / 17512433.2020.1817737.
[0015] Y. Wang et al., "Remodelling and Treatment of the Blood-Brain Barrier in Glioma," (in eng), Cancer Manag Res, vol. 13, pp. 4217-4232, 2021, doi: 10.2147 / CMAR.S288720.
[0016] C. H. Toh and T. Y. Siow, "Factors Associated With Dysfunction of Glymphatic System in Patients With Glioma," (in eng), Front Oncol, vol. 11, p. 744318, 2021, doi: 10.3389 / fonc.2021.744318.
[0017] D. Xu, J. Zhou, H. Mei, H. Li, W. Sun, and H. Xu, "Impediment of Cerebrospinal Fluid Drainage Through Glymphatic System in Glioma," (in eng), Front Oncol, vol. 11, p. 790821, 2021, doi: 10.3389 / fonc.2021.790821.
[0018] M. Gromeier et al., "Very low mutation burden is a feature of inflamed recurrent glioblastomas responsive to cancer immunotherapy," (in eng), Nat Commun, vol. 12, no. 1, p. 352, Jan 13 2021, doi: 10.1038 / s41467-020-20469-6.
[0019] Z. Chen and D. Hambardzumyan, "Immune Microenvironment in Glioblastoma Subtypes," (in eng), Front Immunol, vol. 9, p. 1004, 2018, doi: 10.3389 / fimmu.2018.01004.
[0020] N. Abdelfattah et al., "Single-cell analysis of human glioma and immune cells identifies S100A4 as an immunotherapy target," (in eng), Nat Commun, vol. 13, no. 1, p. 767, Feb 09 2022, doi: 10.1038 / s41467-022-28372-y.
[0021] J. J. D. Moffet et al., "Spatial architecture of high-grade glioma reveals tumor heterogeneity within distinct domains," (in eng), Neurooncol Adv, vol. 5, no. 1, p. vdadl42, 2023, doi: 10.1093 / noajnl / vdadl42.
[0022] J. G. Nicholson and H. A. Fine, "Diffuse Glioma Heterogeneity and Its Therapeutic Implications," (in eng), Cancer Discov, vol. 11, no. 3, pp. 575-590, Mar 2021, doi: 10.1158 / 2159-8290.CD-20-1474.
[0023] C. I. Ene and E. C. Holland, "Personalized medicine for gliomas," (in eng), Surg Neurol Int, vol. 6, no. Suppl 1, pp. S89-95, 2015, doi: 10.4103 / 2152-7806.151351.
[0024] K. Kiyotani, Y. Toyoshima, and Y. Nakamura, "Personalized immunotherapy in cancer precision medicine," (in eng), Cancer Biol Med, vol. 18, no. 4, pp. 955-65, Aug 09 2021, doi:10.20892 / j.issn.2095-3941.2021.0032.
[0025] M. V. C. Greenberg and D. Bourc'his, "The diverse roles of DNA methylation in mammalian development and disease," (in eng), Nat Rev Mol Cell Biol, vol. 20, no. 10, pp. 590-607, Oct 2019, doi: 10.1038 / s41580-019-0159-6.
[0026] M. Farlik et al., "DNA Methylation Dynamics of Human Hematopoietic Stem Cell Differentiation," (in eng), Cell Stem Cell, vol. 19, no. 6, pp. 808-822, Dec 01 2016, doi: 10.1016 / j.stem.2016.10.019.
[0027] M. Okano, D. W. Bell, D. A. Haber, and E. Li, "DNA methyltransferases Dnmt3a and Dnmt3b are essential for de novo methylation and mammalian development," (in eng), Cell, vol. 99, no. 3, pp. 247- 57, Oct 29 1999, doi: 10.1016 / s0092-8674(00)81656-6.Docket No.: 231 / 0021RClient Reference: 2024-026-02
[0028] H. Jeong et al., "Evolution of DNA methylation in the human brain," (in eng), Nat Commun, vol. 12, no. 1, p. 2021, Apr 01 2021, doi: 10.1038 / s41467-021-21917-7.
[0029] L. A. Salas et al., "Enhanced cell deconvolution of peripheral blood using DNA methylation for high- resolution immune profiling," (in eng), Nat Commun, vol. 13, no. 1, p. 761, Feb 09 2022, doi: 10.1038 / S41467-021-27864-7.
[0030] Z. Zhang, J. K. Wiencke, K. T. Kelsey, D. C. Koestler, B. C. Christensen, and L. A. Salas, "HiTIMED: hierarchical tumor immune microenvironment epigenetic deconvolution for accurate cell type resolution in the tumor microenvironment using tumor-type-specific DNA methylation data," (in eng), J Transl Med, vol. 20, no. 1, p. 516, Nov 08 2022, doi: 10.1186 / sl2967-022-03736-6.
[0031] A. J. Titus, R. M. Gallimore, L. A. Salas, and B. C. Christensen, "Cell-type deconvolution from DNA methylation: a review of recent applications," (in eng), Hum Mol Genet, vol. 26, no. R2, pp. R216-R224, Oct 01 2017, doi: 10.1093 / hmg / ddx275.
[0032] Z. Zhang et al., "Hierarchical deconvolution for extensive cell type resolution in the human brain using DNA methylation," 2023, doi: 10.21203 / rs.3.rs-2679515 / vl.
[0033] M. E. Muse, C. D. Carroll, L. A. Salas, M. R. Karagas, and B. C. Christensen, "Application of Novel Breast Biospecimen Cell-Type Adjustment Identifies Shared DNA Methylation Alterations in Breast Tissue and Milk with Breast Cancer-Risk Factors," (in eng), Cancer Epidemiol Biomarkers Prev, vol. 32, no. 4, pp. 550-560, Apr 03 2023, doi: 10.1158 / 1055-9965. EPI-22-0405.
[0034] T. Zhu et al., "A pan-tissue DNA methylation atlas enables in silico decomposition of human tissue methylomes at cell-type resolution," (in eng), Nat Methods, vol. 19, no. 3, pp. 296-306, Mar 2022, doi: 10.1038 / S41592-022-01412-7.
[0035] J. Yang et al., "DNA methylation-based epigenetic signatures predict somatic genomic alterations in gliomas," (in eng), Nat Commun, vol. 13, no. 1, p. 4410, Jul 29 2022, doi: 10.1038 / s41467-022-31827-x.
[0036] S. Ferreyra Vega, T. Olsson Bontell, A. Corell, A. Smits, A. S. Jakola, and H. Caren, "DNA methylation profiling for molecular classification of adult diffuse lower-grade gliomas," (in eng), Clin Epigenetics, vol. 13, no. 1, p. 102, May 03 2021, doi: 10.1186 / sl3148-021-01085-7.
[0037] Z. Wu et al., "Impact of the methylation classifier and ancillary methods on CNS tumor diagnostics," (in eng), Neuro Oncol, vol. 24, no. 4, pp. 571-581, Apr 01 2022, doi: 10.1093 / neuonc / noab227.
[0038] A. Wenger and H. Caren, "Methylation Profiling in Diffuse Gliomas: Diagnostic Value and Considerations," (in eng), Cancers (Basel), vol. 14, no. 22, Nov 18 2022, doi: 10.3390 / cancersl4225679.
[0039] D. Capper et al., "DNA methylation-based classification of central nervous system tumours," (in eng), Nature, vol. 555, no. 7697, pp. 469-474, Mar 22 2018, doi: 10.1038 / nature26000.
[0040] D. Arneson, X. Yang, and K. Wang, "MethyIResolver-a method for deconvoluting bulk DNA methylation profiles into known and unknown cell contents," (in eng), Commun Biol, vol. 3, no. 1, p. 422, Aug 03 2020, doi: 10.1038 / s42003-020-01146-2.
[0041] A. Chakravarthy et al., "Pan-cancer deconvolution of tumour composition using DNA methylation," (in eng), Nat Commun, vol. 9, no. 1, p. 3220, Aug 13 2018, doi: 10.1038 / s41467-018-05570-l.
[0042] O. Singh, D. Pratt, and K. Aidape, "Immune cell deconvolution of bulk DNA methylation data reveals an association with methylation class, key somatic alterations, and cell state in glial / glioneuronal tumors," (in eng), Acta Neuropathol Commun, vol. 9, no. 1, p. 148, Sep 082021, doi: 10.1186 / s40478-021- 01249-9.
[0043] P. G. Weightman Potter et al., "Attenuated Induction of the Unfolded Protein Response in Adult Human Primary Astrocytes in Response to Recurrent Low Glucose," (in eng), Front Endocrinol (Lausanne), vol. 12, p. 671724, 2021, doi: 10.3389 / fendo.2021.671724.
[0044] X. Lin et al., "Cell type-specific DNA methylation in neonatal cord tissue and cord blood: a 850K- reference panel and comparison of cell types," (in eng), Epigenetics, vol. 13, no. 9, pp. 941-958, 2018, doi: 10.1080 / 15592294.2018.1522929.
[0045] A. Kozlenkov et al., "A unique role for DNA (hydroxy)methylation in epigenetic regulation of human inhibitory neurons," (in eng), Sci Adv, vol. 4, no. 9, p. eaau6190, Sep 2018, doi: 10.1126 / sciadv.aau6190.
[0046] L. D. de Witte et al., "Contribution of Age, Brain Region, Mood Disorder Pathology, and Interindividual Factors on the Methylome of Human Microglia," (in eng), Biol Psychiatry, vol. 91, no. 6, pp. 572-581, Mar 15 2022, doi: 10.1016 / j. biopsych.2021.10.020.
[0047] I. Mendizabal et al., "Cell type-specific epigenetic links to schizophrenia risk in the brain," (in eng), Genome Biol, vol. 20, no. 1, p. 135, Jul 09 2019, doi: 10.1186 / sl3059-019-1747-7.Docket No.: 231 / 0021RClient Reference: 2024-026-02
[0048] Z. Xu, L. Niu, L. Li, and J. A. Taylor, "ENmix: a novel background correction method for Illumina HumanMethylation450 BeadChip," (in eng), Nucleic Acids Res, vol. 44, no. 3, p. e20, Feb 182016, doi: 10.1093 / nar / gkv907.
[0049] Z. Xu, L. Niu, and J. A. Taylor, "The ENmix DNA methylation analysis pipeline for Illumina BeadChip and comparisons with seven other preprocessing pipelines," (in eng), Clin Epigenetics, vol. 13, no. 1, p. 216, Dec 09 2021, doi: 10.1186 / sl3148-021-01207-l.
[0050] W. Zhou, T. J. Triche, P. W. Laird, and H. Shen, "SeSAMe: reducing artifactual detection of DNA methylation by Infinium BeadChips in genomic deletions," (in eng), Nucleic Acids Res, vol. 46, no. 20, p. el23, Nov 162018, doi: 10.1093 / nar / gky691.
[0051] A. J. Titus, E. A. Houseman, K. C. Johnson, and B. C. Christensen, "methyLiftover: cross-platform DNA methylation data integration," (in eng), Bioinformatics, vol. 32, no. 16, pp. 2517-9, Aug 15 2016, doi: 10.1093 / bioinformatics / btwl80.
[0052] A. E. Teschendorff et al., "A beta-mixture quantile normalization method for correcting probe design bias in Illumina Infinium 450 k DNA methylation data," (in eng), Bioinformatics, vol. 29, no. 2, pp. 189- 96, Jan 15 2013, doi: 10.1093 / bioinformatics / bts680.
[0053] Y. Tian et al., "ChAMP: updated methylation analysis pipeline for Illumina BeadChips," (in eng), Bioinformatics, vol. 33, no. 24, pp. 3982-3984, Dec 15 2017, doi: 10.1093 / bioinformatics / btx513.
[0054] J. A. Heiss and A. C. Just, "Improved filtering of DNA methylation microarray data by detection p values and its impact on downstream analyses," (in eng), Clin Epigenetics, vol. 11, no. 1, p. 15, Jan 24 2019, doi: 10.1186 / S13148-019-0615-3.
[0055] M. J. Aryee et al., "Minfi: a flexible and comprehensive Bioconductor package for the analysis of Infinium DNA methylation microarrays," (in eng), Bioinformatics, vol. 30, no. 10, pp. 1363-9, May 15 2014, doi: 10.1093 / bioinformatics / btu049.
[0056] T. J. Triche, D. J. Weisenberger, D. Van Den Berg, P. W. Laird, and K. D. Siegmund, "Low-level processing of Illumina Infinium DNA Methylation BeadArrays," (in eng), Nucleic Acids Res, vol. 41, no.7, p. e90, Apr 2013, doi: 10.1093 / nar / gkt090.
[0057] X. Zheng, N. Zhang, H. J. Wu, and H. Wu, "Estimating and accounting for tumor purity in the analysis of DNA methylation data from cancer studies," (in eng), Genome Biol, vol. 18, no. 1, p. 17, Jan 252017, doi: 10.1186 / S13059-016-1143-5.
[0058] J. L. Min, G. Hemani, G. Davey Smith, C. Relton, and M. Suderman, "Meffil: efficient normalization and analysis of very large DNA methylation datasets," (in eng), Bioinformatics, vol. 34, no. 23, pp. 3983- 3989, Dec 01 2018, doi: 10.1093 / bioinformatics / bty476.
[0059] E. A. Houseman et al., "DNA methylation arrays as surrogate measures of cell mixture distribution," (in eng), BMC Bioinformatics, vol. 13, p. 86, May 082012, doi: 10.1186 / 1471-2105-13-86.
[0060] W. Zhou, P. W. Laird, and H. Shen, "Comprehensive characterization, annotation and innovative use of Infinium DNA methylation BeadChip probes," (in eng), Nucleic Acids Res, vol. 45, no. 4, p. e22, Feb 28 2017, doi: 10.1093 / nar / gkw967.
[0061] W. J. Kent et al., "The human genome browser at UCSC," (in eng), Genome Res, vol. 12, no. 6, pp. 996- 1006, Jun 2002, doi: 10.1101 / gr.229102.
[0062] J. Maksimovic, A. Oshiack, and B. Phipson, "Gene set enrichment analysis for genome-wide DNA methylation data," (in eng), Genome Biol, vol. 22, no. 1, p. 173, Jun 08 2021, doi: 10.1186 / sl3059-021- 02388-x.
[0063] M. E. Ritchie et al., "limma powers differential expression analyses for RNA-sequencing and microarray studies," (in eng), Nucleic Acids Res, vol. 43, no. 7, p. e47, Apr 20 2015, doi: 10.1093 / nar / gkv007.
[0064] L. Wang, C. C. Yu, X. Y. Liu, X. N. Deng, Q. Tian, and Y. J. Du, "Epigenetic Modulation of Microglia Function and Phenotypes in Neurodegenerative Diseases," (in eng), Neural Plast, vol. 2021, p. 9912686, 2021, doi: 10.1155 / 2021 / 9912686.
[0065] I. Shchukina et al., "Enhanced epigenetic profiling of classical human monocytes reveals a specific signature of healthy aging in the DNA methylome," (in eng), Nat Aging, vol. 1, no. 1, pp. 124-141, Jan 2021, doi: 10.1038 / s43587-020-00002-6.
[0066] M. Ceccarelli et al., "Molecular Profiling Reveals Biologically Discrete Subsets and Pathways of Progression in Diffuse Glioma," (in eng), Cell, vol. 164, no. 3, pp. 550-63, Jan 282016, doi: 10.1016 / j. cell.2015.12.028.Docket No.: 231 / 0021RClient Reference: 2024-026-02
[0067] S. Hussain and S. Davanger, "Postsynaptic VAMP / Synaptobrevin Facilitates Differential Vesicle Trafficking of GluAl and GluA2 AMPA Receptor Subunits," (in eng), PLoS One, vol. 10, no. 10, p. e0140868, 2015, doi: 10.1371 / journal.pone.0140868.
[0068] J. Hou, Y. Chen, G. Grajales-Reyes, and M. Colonna, "TREM2 dependent and independent functions of microglia in Alzheimer's disease," (in eng), Mol Neurodegener, vol. 17, no. 1, p. 84, Dec 23 2022, doi: 10.1186 / sl3024-022-00588-y.
[0069] D. Khantakova, S. Brioschi, and M. Molgora, "Exploring the Impact of TREM2 in Tumor-Associated Macrophages," (in eng), Vaccines (Basel), vol. 10, no. 6, Jun 14 2022, doi: 10.3390 / vaccinesl0060943.
[0070] X. Chen and P. E. Jensen, "The expression of HLA-DO (H2-O) in B lymphocytes," (in eng), Immunol Res, vol. 29, no. 1-3, pp. 19-28, 2004, doi: 10.1385 / IR:29:l-3:019.
[0071] R. Edgar, P. P. Tan, E. Portales-Casamar, and P. Pavlidis, "Meta-analysis of human methylomes reveals stably methylated sequences surrounding CpG islands associated with high gene expression," (in eng), Epigenetics Chromatin, vol. 7, no. 1, p. 28, 2014, doi: 10.1186 / 1756-8935-7-28.
[0072] D. Aran, M. Sirota, and A. J. Butte, "Systematic pan-cancer analysis of tumour purity," (in eng), Nat Commun, vol. 6, p. 8971, Dec 04 2015, doi: 10.1038 / ncomms9971.
[0073] K. Yoshihara et al., "Inferring tumour purity and stromal and immune cell admixture from expression data," (in eng), Nat Commun, vol. 4, p. 2612, 2013, doi: 10.1038 / ncomms3612.
[0074] S. Han et al., "IDH mutation in glioma: molecular mechanisms and potential therapeutic targets," (in eng), BrJ Cancer, vol. 122, no. 11, pp. 1580-1589, May 2020, doi: 10.1038 / s41416-020-0814-x.
[0075] B. C. Christensen et al., "DNA methylation, isocitrate dehydrogenase mutation, and survival in glioma," (in eng), J Natl Cancer Inst, vol. 103, no. 2, pp. 143-53, Jan 19 2011, doi: 10.1093 / jnci / djq497.
[0076] H. Noushmehr et al., "Identification of a CpG island methylator phenotype that defines a distinct subgroup of glioma," (in eng), Cancer Cell, vol. 17, no. 5, pp. 510-22, May 18 2010, doi: 10.1016 / j.ccr.2010.03.017.
[0077] A. P. Patel et al., "Single-cell RNA-seq highlights intratumoral heterogeneity in primary glioblastoma," (in eng), Science, vol. 344, no. 6190, pp. 1396-401, Jun 202014, doi: 10.1126 / science.1254257.
[0078] C. P. Couturier et al., "Single-cell RNA-seq reveals that glioblastoma recapitulates a normal neurodevelopmental hierarchy," (in eng), Nat Commun, vol. 11, no. 1, p. 3406, Jul 082020, doi: 10.1038 / s41467-020-17186-5.
[0079] Q. Wang et al., "Tumor Evolution of Glioma-Intrinsic Gene Expression Subtypes Associates with Immunological Changes in the Microenvironment," (in eng), Cancer Cell, vol. 32, no. 1, pp. 42-56. e6, Jul 10 2017, doi: 10.1016 / j.ccell.2017.06.003.
[0080] S. L. Carter et al., "Absolute quantification of somatic DNA alterations in human cancer," (in eng), Nat Biotechnol, vol. 30, no. 5, pp. 413-21, May 2012, doi: 10.1038 / nbt.2203.
[0081] A. Noguera-Castells, C. A. Garcia-Prieto, D. Alvarez-Errico, and M. Esteller, "Validation of the new EPIC DNA methylation microarray (900K EPIC v2) for high-throughput profiling of the human DNA methylome," (in eng), Epigenetics, vol. 18, no. 1, p. 2185742, Dec 2023, doi: 10.1080 / 15592294.2023.2185742.
[0082] M. Colonna, "The biology of TREM receptors," (in eng), Nat Rev Immunol, pp. 1-15, Feb 07 2023, doi: 10.1038 / S41577-023-00837-1.
[0083] K. Nakamura and M. J. Smyth, "TREM2 marks tumor-associated macrophages," (in eng), Signal Transduct Target Ther, vol. 5, no. 1, p. 233, Oct 092020, doi: 10.1038 / s41392-020-00356-8.
[0084] S. Muller et al., "Single-cell profiling of human gliomas reveals macrophage ontogeny as a basis for regional differences in macrophage activation in the tumor microenvironment," (in eng), Genome Biol, vol. 18, no. 1, p. 234, Dec 20 2017, doi: 10.1186 / sl3059-017-1362-4.
[0085] N. D. Mathewson et al., "Inhibitory CD161 receptor identified in glioma-infiltrating T cells by single-cell analysis," (in eng), Cell, vol. 184, no. 5, pp. 1281-1298. e26, Mar 042021, doi:10.1016 / j. cell.2021.01.022.
[0086] K. C. Johnson et al., "Single-cell multimodal glioma analyses identify epigenetic regulators of cellular plasticity and environmental stress response," (in eng), Nat Genet, vol. 53, no. 10, pp. 1456-1468, Oct 2021, doi: 10.1038 / s41588-021-00926-8.
[0087] E. Karimi et al., "Single-cell spatial immune landscapes of primary and metastatic brain tumours," (in eng), Nature, vol. 614, no. 7948, pp. 555-563, Feb 2023, doi: 10.1038 / s41586-022-05680-3.Docket No.: 231 / 0021RClient Reference: 2024-026-02
[0088] Y. Hoogstrate et al., "Transcriptome analysis reveals tumor microenvironment changes in glioblastoma," (in eng), Cancer Cell, vol. 41, no. 4, pp. 678-692. e7, Apr 102023, doi: 10.1016 / j.ccell.2023.02.019.
[0089] K. A. Schalper et al., "Neoadjuvant nivolumab modifies the tumor immune microenvironment in resectable glioblastoma," (in eng), Nat Med, vol. 25, no. 3, pp. 470-476, Mar 2019, doi: 10.1038 / s41591-018-0339-5.
[0090] L. Pinton et al., "The immune suppressive microenvironment of human gliomas depends on the accumulation of bone marrow-derived macrophages in the center of the lesion," (in eng), J Immunother Cancer, vol. 7, no. 1, p. 58, Feb 27 2019, doi: 10.1186 / s40425-019-0536-x.
[0091] W. Fu et al., "Single-Cell Atlas Reveals Complexity of the Immunosuppressive Microenvironment of Initial and Recurrent Glioblastoma," (in eng), Front Immunol, vol. 11, p. 835, 2020, doi: 10.3389 / fimmu.2020.00835.
[0092] Q. Feng, L. Li, M. Li, and X. Wang, "Immunological classification of gliomas based on immunogenomic profiling," (in eng), J Neuroinflammation, vol. 17, no. 1, p. 360, Nov 27 2020, doi: 10.1186 / sl2974-020- 02030-w.
[0093] F. Wu et al., "Immunological profiles of human oligodendrogliomas define two distinct molecular subtypes," (in eng), EBioMedicine, vol. 87, p. 104410, Jan 2023, doi: 10.1016 / j.ebiom.2022.104410.
[0094] K. Yu et al., "Surveying brain tumor heterogeneity by single-cell RNA-sequencing of multi-sector biopsies," (in eng), Natl Sci Rev, vol. 7, no. 8, pp. 1306-1318, Aug 2020, doi: 10.1093 / nsr / nwaa099.
[0095] A. Buonfiglioli and D. Hambardzumyan, "Macrophages and microglia: the cerberus of glioblastoma," (in eng), Acta Neuropathol Commun, vol. 9, no. 1, p. 54, Mar 25 2021, doi: 10.1186 / s40478-021-01156- z.
[0096] L. A. Salas, J. K. Wiencke, D. C. Koestler, Z. Zhang, B. C. Christensen, and K. T. Kelsey, "Tracing human stem cell lineage during development using DNA methylation," (in eng), Genome Res, vol. 28, no. 9, pp. 1285-1295, Sep 2018, doi: 10.1101 / gr.233213.117.
[0097] Z. Zhang, Y. Lu, S. Vosoughi, J. J. Levy, B. C. Christensen, and L. A. Salas, "erarchical tumor artificial intelligence classifier traces tissue of origin and tumor type in primary and metastasized tumors using DNA methylation," (in eng), NAR Cancer, vol. 5, no. 2, p. zcad017, Jun 2023, doi: 10.1093 / narcan / zcad017.
[0098] M. Andoh and R. Koyama, "Comparative Review of Microglia and Monocytes in CNS Phagocytosis," (in eng), Cells, vol. 10, no. 10, Sep 27 2021, doi: 10.3390 / cellsl0102555.
[0099] C. M. Laumont, A. C. Banville, M. Gilardi, D. P. Hollern, and B. H. Nelson, "Tumour-infiltrating B cells: immunological mechanisms, clinical impact and therapeutic opportunities," (in eng), Nat Rev Cancer, vol. 22, no. 7, pp. 414-430, Jul 2022, doi: 10.1038 / s41568-022-00466-l.
[0100] E. C. Cordell, M. S. Alghamri, M. G. Castro, and D. H. Gutmann, "T lymphocytes as dynamic regulators of glioma pathobiology," (in eng), Neuro Oncol, vol. 24, no. 10, pp. 1647-1657, Oct 03 2022, doi: 10.1093 / neuonc / noac055.
[0101] A. E. Jaffe and R. A. Irizarry, "Accounting for cellular heterogeneity is critical in epigenome-wide association studies," (in eng), Genome Biol, vol. 15, no. 2, p. R31, Feb 04 2014, doi: 10.1186 / gb-2014- 15-2-r31.
[0102] Q. Zhang et al., "Intra-tumoral angiogenesis correlates with immune features and prognosis in glioma," (in eng), Aging (Albany NY), vol. 14, no. 10, pp. 4402-4424, May 17 2022, doi: 10.18632 / aging.204079.
[0103] T. Hu et al., "Construction and validation of an angiogenesis-related gene expression signature associated with clinical outcome and tumor immune microenvironment in glioma," (in eng), Front Genet, vol. 13, p. 934683, 2022, doi: 10.3389 / fgene.2022.934683.
[0104] J. Wang, A. Shan, F. Shi, and Q. Zheng, "Molecular and clinical characterization of ANG expression in gliomas and its association with tumor-related immune response," (in eng), Front Med (Lausanne), vol. 10, p. 1044402, 2023, doi: 10.3389 / fmed.2023.1044402.
[0105] I. M. Heid, H. Kuchenhoff, J. Miles, L. Kreienbrock, and H. E. Wichmann, "Two dimensions of measurement error: classical and Berkson error in residential radon exposure assessment," (in eng), J Expo Anal Environ Epidemiol, vol. 14, no. 5, pp. 365-77, Sep 2004, doi: 10.1038 / sj.jea.7500332.
[0106] Code utilized in the manuscript entitled "Glioma Immune Microenvironment Composition Calculator (GIMiCC): a method of estimating the proportions of eighteen key cell types from glioma DNA methylation microarray data" DOI: 10.5281 / zenodo.10093224.Docket No.: 231 / 0021RClient Reference: 2024-026-02
[0110] The foregoing has been a detailed description of illustrative embodiments of the invention. Various modifications and additions can be made without departing from the spirit and scope of this invention. Features of each of the various embodiments described above may be combined with features of other described embodiments as appropriate in order to provide a multiplicity of feature combinations in associated new embodiments.Furthermore, while the foregoing describes a number of separate embodiments of the apparatus and method of the present invention, what has been described herein is merely illustrative of the application of the principles of the present invention. For example, as used herein, the terms “process” and / or “processor” should be taken broadly to include a variety of electronic hardware and / or software-based functions and components (and can alternatively be termed functional “modules” or “elements”). Moreover, a depicted process or processor can be combined with other processes and / or processors or divided into various sub-processes or processors. Such sub-processes and / or sub-processors can be variously combined according to embodiments herein. Likewise, it is expressly contemplated that any function, process and / or processor herein can be implemented using electronic hardware, software consisting of a non-transitory computer-readable medium of program instructions, or a combination of hardware and software. Additionally, as used herein various directional and dispositional terms such as “vertical”, “horizontal”, “up”, "down", “bottom”, “top”, “side”, “front”, “rear”, “left”, “right”, and the like, are used only as relative conventions and not as absolute directions / dispositions with respect to a fixed coordinate space, such as the acting direction of gravity'. Additionally, where the term “substantially” or “approximately” is employed w ith respect to a given measurement, value, or characteristic, it refers to a quantity' that is within a normal operating range to achieve desired results, but that includes some variability due to inherent inaccuracy and error within the allowed tolerances of the system (e.g., 1-5 percent). Accordingly, this description is meant to be taken only by way of example, and not to otherwise limit the scope of this invention.
[0111] What is claimed is:
Claims
Docket No.: 231 / 0021RClient Reference: 2024-026-02CLAIMS1. A system for diagnosing glioma in a patient using a processor and a user interface comprising: a data input process that receives information related to the patient’s cancerous tissue based on DNA methylation; a data store that includes a plurality of layers of information arranged relative to types of cancerous conditions in tissue and predetermined types on non-cancerous tissue based on DNA methylation therein; an analysis process, including a trained machine learning process, that compares DNA methylation characteristics in the patient’s tissue to the layers of information and performs a match so as to trace the cancerous conditions; and a diagnostic process that provides a user with information relative to the match.
2. The system as set forth in claim 1, wherein the layers include (a) deconvolution results for at least four subty pes of brain tumors, and (b) deconvolution of the immune microenvironment of each type of the brain tumors, respectively.
3. The system as set forth in claim 2, wherein the analysis process is constructed and arranged to deconvolve glioma DNAm microarray data from at least 17 isolated cell types.
4. The system as set forth in claim 1, wherein the data store includes publicly available information accessed through a public data communication network.
5. The system as set forth in claim 1, wherein the machine learning process is trained using classifiers related to predetermined DNA methylation characteristics for related cancerous conditions.
6. The system as set forth in claim 1, wherein at least one of the data input process, the analysis process and the diagnostic process is operated using a processor on a user-controlled computing device with a user interface.Docket No.: 231 / 0021R Client Reference: 2024-026-027. A method for diagnosing and reporting upon cancerous conditions that operates the system of claim 1.
8. A method for diagnosing glioma in a patient using a processor and a user interface comprising the steps of: receiving information related to the patient’s cancerous tissue based on DNA methylation; accessing a data store that includes a plurality of layers of information arranged relative to types of cancerous conditions in tissue and predetermined types on non-cancerous tissue based on DNA methylation therein; comparing, with a trained machine learning process, DNA methylation characteristics in the patient’s tissue to the layers of information and performs a match so as to trace the cancerous conditions; and providing a user with information relative to the match.
9. The method as set forth in claim 8, wherein the layers include (a) deconvolution results for at least four subty pes of brain tumors, and (b) deconvolution of the immune microenvironment of each type of the brain tumors, respectively.
10. The method as set forth in claim 9, wherein the step of comparing includes deconvolving glioma DNAm microarray data from at least 17 isolated cell types.
11. The method as set forth in claim 8. wherein the step of accessing data store includes accessing publicly available information through a public data communication network.
12. The method as set forth in claim 8, further comprising, training the machine learning process using classifiers related to predetermined DNA methylation characteristics for related cancerous conditions.
13. The method as set forth in claim 8, wherein at least one of the steps of accessing, comparing and providing is operated using a processor on a user-controlled computing device with a user interface.Docket No.: 231 / 0021R Client Reference: 2024-026-02 14. The method as set forth in claim 13, further comprising, logging into, by a user, a public-network based subscription service, or using of validated credentials that are established for the user.
15. The method as set forth in claim 8, further comprising, deriving an optimized treatment for the patient based on the step of providing.
Citation Information
Patent Citations
Cancer detection and classification using methylome analysis
US20220251665A1
Optical access system and monitoring method
US20240187092A1
System and method for hierarchical tumor immune microenvironment epigenetic deconvolution
WO2023196051A1