Predicting DNA methylation and tumor types from histopathology
A deep learning framework for CNS tumor classification using histopathology images addresses diagnostic challenges by integrating methylation, direct, and demographic models, enhancing accuracy and reducing infrastructure needs.
Patent Information
- Application Number
- PCT/US2025/013654
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-29
- Filing Date
- 2025-01-29
- Publication Date
- 2025-08-07
AI Technical Summary
Current methods for diagnosing CNS tumors rely heavily on interobserver variability and require expensive, time-consuming genome-wide DNA methylation profiling, which is not feasible in underresourced areas, delaying treatment decisions for aggressive tumors.
A deep learning framework that predicts DNA methylation and classifies tumor types using histopathology images, integrating methylation, direct classification, and demographic models to enhance diagnostic accuracy and reduce reliance on costly infrastructure.
Provides rapid, accurate tumor classification and treatment recommendations, improving diagnostic precision and clinical outcomes in resource-limited settings.
Smart Images

Figure US2025013654_07082025_PF_FP_ABST
Abstract
Description
PREDICTING DNA METHYLATION AND TUMOR TYPES FROM HISTOPATHOLOGYGOVERNMENT INTEREST STATEMENT
[0001] The present subject matter was made with U.S. government support under project numbers ZIA BC 011802 and ZIA BC 011850 by the National Institutes of Health, National Cancer Institute. The U.S. government has certain rights in the invention.FIELD
[0002] The present disclosure relates to systems and methods for predicting DNA methylation and tumor types from histopathology. More specifically, the disclosure relates to integrated deep learning models for predicting DNA methylation and tumor types from histopathology in tumors.BACKGROUND
[0003] Pathologic diagnosis plays a critical role in guiding the clinical care of patients with tumors of the central nervous system (CNS). Over 100 CNS tumor types are recognized by the World Health Organization WHO), and the diagnostic process starts with an examination of hematoxylin and eosin (H&E)-stained slides. CNS tumor diagnoses can be subject to interobserver variability using standard practices of conventional histopathology and technologies such as immunohistochemistry. Nextgeneration sequencing can serve as an adjunct, but only for those tumors with tumor type-defining genomic alterations. Genome-wide DNA methylation profiling has been proposed as an important diagnostic modality, where it has been found to reclassify a proportion of CNS tumors even if they have been reviewed by expert neuropathologists. DNA-methylation-based classifiers for CNS tumors reflect both cel l-of-orig in as well as changes acquired during the neoplastic process, and while powerful, the assay is currently performed in only a few centers, and is not currently feasible in underresourced areas. The commonly used methylation platform utilizes a genome-wide array that provides data on CpG site-specific methylation levels (beta values) for approximately 850,000 sites across the genome.
[0004] Progress in deep learning holds great promise in achieving accurate performance in diagnostic medicine. Deep learning techniques have demonstrated the ability to identify characteristics that may not be easily recognized by human observers, as evidenced by various studies. A recent study examined deep learning on histology slides in a study of CNS tumor diagnosis and biomarker assessment. Moreover, deep learning methods have shown potential in predicting genomic features in tumors based on H&E images, such as genetic mutations, bulk mRNAseq expression, and spatial mRNAseq expression.
[0005] With respect to DNA methylation, previous work has utilized morphometric features to classify hypermethylated and hypomethylated levels in gliomas, but no study has been conducted to predict DNA methylation beta values on a large genomic scale from histopathology images using deep learning models, and then utilize those predicted methylation values to classify tumor types. If deep learning models could predict the methylation status of many CpG sites across the genome, it might contribute to diagnostic accuracy without the need for methylation profiling, which involves cost and infrastructure that is not currently available at the majority of centers. The turnaround time (often several weeks) required for methylation testing in clinical laboratories can delay diagnoses and therefore treatment decisions for patients who may have highly aggressive tumors that require prompt therapeutic intervention. In addition, deep learning approaches may shed light on the spatial organization of the epigenomic landscape of an individual, thus potentially providing further biological and clinical insights.
[0006] Therefore, there is a need for a method of using deep learning framework that predicts DNA methylation and classifies tumor subtypes using histopathology to distinguish and classify tumors to provide a more accurate diagnosis and improved treatment protocols.SUMMARY
[0007] This disclosure provides systems and methods of predicting DNA methylation and tumor types from histopathology.
[0008] Provided herein are systems and methods for classifying a tumor. The system may be configured to receive a histopathology image of a sample of the tumor, generate a first, second, and / or third classification of the tumor using trained models / classifiers and sets of prediction scores, and generate an integrated classification of the tumor. The classifications and prediction scores from the various trained models may be generated simultaneously or sequentially, in any order.
[0009] The first classification of the tumor may be generated by: providing the histopathology image to a trained methylation model to predict a methylation profile for the tumor, providing the predicted methylation profile to a trained methylation classifier, and generating a first set of prediction scores based on the trained methylation classifier. The second classification of the tumor may be generated by: providing the histopathology image to a trained direct classification model, and generating a second set of prediction scores based on the trained direct classification model. The third classification of the tumor may be generated by: providing demographic data to a trained demographic model and generating a third set of prediction scores based on the trained demographic model. The integrated classification may be generated by averaging the first set of prediction scores, the second set of prediction scores, and the third set of prediction scores and selecting the classification with the highest averaged prediction score. The classification is selected from a set of tumor families, types, and sub-types.
[0010] In some aspects, the histopathology image is preprocessed by dividing the histopathology image into a plurality of non-overlapping tiles, omitting any nonoverlapping tiles which include more than half background, extracting features from the non-overlapping tiles, and compressing the histopathology image. The first set of prediction scores and the second set of prediction scores may include a set of prediction scores for each of the non-overlapping tiles in the histopathology image, thereby facilitating a spatial classification of the histopathology image using the methylation model, the methylation classifier, the direct classification model, or a combination thereof.
[0011] In some aspects, the system may further be configured to update the trained methylation model and the trained methylation classifier with the methylationprofile and the integrated classification, update the trained direct classification model with the histopathology image and the integrated classification, and update the trained demographic model with the demographic data and the integrated classification.
[0012] In an aspect, the trained methylation model comprises a deep learning model trained using labeled data comprising matched slide images and DNA methylation profiles. The trained methylation classifier may include at least two machine learning models trained using labeled data comprising DNA methylation data and tumor types. Non-limiting examples of the at least two machine learning models include logistic regression, support vector machine, k-nearest neighbors, random forest, and combinations thereof. The predicted methylation profile may include DNA methylation beta values.
[0013] In an aspect, the trained direct classification model comprises a deep learning model trained using labeled data comprising histopathology images and tumor types. In an aspect, the trained demographic model comprises a machine learning model trained using labeled data comprising demographic data and tumor types. The demographic data may include age, sex, and location of the tumor. In some aspects, the integrated classification comprises at least one tumor family, type, sub-type or more than one tumor sub-type.
[0014] In various aspects, the system may further be configured to train the methylation model, the methylation classifier, the direct model, and / or the demographic model prior to providing the histopathology image.
[0015] In additional aspects, the system may be configured to diagnose the tumor using the generated integrated classification. A treatment plan may be formed specific to the diagnosis of the tumor.
[0016] In some aspects, the histopathology image is an H&E image.BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
[0018] The description will be more fully understood with reference to the following figures and data graphs, which are presented as various embodiments of the disclosure and should not be construed as a complete recitation of the scope of the disclosure. It is noted that, for purposes of illustrative clarity, certain elements in various drawings may not be drawn to scale. Understanding that these drawings depict only exemplary embodiments of the disclosure and are not therefore to be considered to be limiting of its scope, the principles herein are described and explained with additional specificity and detail through the use of the accompanying drawings in which:
[0019] FIG. 1 is an example classification method in one embodiment.
[0020] FIG. 2 illustrates example system embodiments.
[0021] FIG. 3 illustrates an example machine learning environment.
[0022] FIG. 4A shows the integrated model framework comprises several key components. Initially, the whole-slide images (WSIs) were partitioned into small tiles. These tiles were then subjected to feature extraction using the pre-trained ResNet50 model. Subsequently, an autoencoder was employed to compress the 2,048 ResNet features to a lower-dimensional representation consisting of 512 features. To construct a DNA methylation model, MLP regression was utilized based on the autoencoder features. For tumor classification, the integrated model incorporates three distinct models. The direct model, displayed in the orange square, employs an MLP classifier that uses features compressed by an autoencoder to directly predict tumor types. The Indirect / methylation Model, denoted by green squares, employs four classical machine learning algorithms to classify tumors based on predicted DNA methylation levels. Complementing these, the Demographic Model, displayed in the light blue square, utilizes a Random Forest classifier to incorporate age, sex, and tumor location for tumor type classification. The integrated model averages prediction scores across these three methods by computing their mean without any additional training or fine-tuning.
[0023] FIG. 4B shows patient cohorts: the integrated model was trained and cross-validated using an in-house (NCI) dataset, which comprises matched slides and DNA profiles from 1 ,796 patients. For external validation, slides from 1 ,522 patients from the Digital Brain Tumor Atlas (DBTA), 348 patients from the Children's Brain Tumor Network (CBTN), and 286 patients from the NCI-Prospective were used.
[0024] FIG. 4C shows while the class labels are at the whole slide level, the direct and indirect models are able to provide predictions at tile-level resolution, which, at times, predicts the co-existence of different tumor types that are depicted in a spatially organized manner within a single Whole-Slide Image (WSI).
[0025] FIG. 5A are histograms illustrating the distribution of Pearson correlation coefficients between predicted and actual methylation beta-values for each probe. On the left is the cross-validation result from the NCI cohort, while the middle and right histograms illustrate independent validation in the CBTN and NCI-Prospective cohorts, respectively.
[0026] FIG. 5B are curves that show the number of probes (y-axis) that exceed a specific threshold predicted / measured correlation (x-axis). Again, the left graph relates to the NCI cohort, the middle to the CBTN cohort, and the right to the NCI-Prospective cohort.
[0027] FIG. 5C is the Venn diagram illustrating the intersection among the three cohorts of probes exhibiting a predicted / measured correlation coefficient exceeding 0.4.
[0028] FIG. 5D are bar plots comparing the number of hypermethylated and hypomethylated sites in the NCI IDH-mutant gliomas versus IDH-wildtype Glioblastomas, based on actual (measured, left) and DEPLOY-predicted (right) beta values. Both clearly match the known global hypermethylation characteristic of IDH- mutant gliomas.
[0029] FIG. 5E shows Venn diagrams illustrating the overlap between differentially methylated sites in IDH mutants that are called according to actual and DEPLOY-predicted beta values. Overlaps of hypermethylated sites are shown on the right, while those for hypomethylated sites are depicted on the left. P-values were calculated using the Fisher exact test.
[0030] FIG. 5F shows cancer hallmarks pathway enrichment analysis for gene promoter methylation. Each row represents a cancer hallmark pathway, while each column corresponds to a different brain tumor type. The tumor type names in the columns are abbreviated. Values inside the grid denote the mean normalized enrichment score for the specific tumor subtype and pathway combination. Analysesbased on actual and predicted beta-values from the NCI cohort are displayed on the left and right, respectively.
[0031] FIG. 5G is a similar plot to FIG. 5F but for gene body methylation.
[0032] FIG. 5H shows scatter plots for the average enrichment score per brain cancer subtype in the NCI cohort, displaying actual values on the x-axes and predicted values on the y-axes. Each row of the plots corresponds to a distinct cancer hallmark pathway, following the order presented in (FIG. 5F) and (FIG. 5G). The left column shows promoter analysis results and the right column shows gene body analysis results. Pearson correlation coefficients between actual and predicted values across the different cancer subtypes are also shown in each plot.
[0033] FIG. 6A shows micro-averaged precision-recall curves, the area under the precision-recall curve (AUPRC) values are depicted for the four models: Demographic (light blue), Direct (orange), Indirect (green), and Integrated (DEPLOY, dark blue) for each cohort: NCI, DBTA, CBTN, and NCI-Prospective.
[0034] FIG. 6B shows top-k accuracy metrics depicted for the four models: Demographic (light blue), Direct (orange), Indirect (green), and Integrated (DEPLOY, dark blue) for each cohort: NCI, DBTA, CBTN, and NCI-Prospective.
[0035] FIG. 6C are confusion matrices obtained by the integrated model for the NCI, DBTA, CBTN, and NCI-Prospective cohorts. Each cell's color intensity indicates the fraction of DEPLOY-assigned classifications in relation to the total sample size for each actual tumor type in the corresponding column. Notably, GBM, SE and O-IDH are not present in the CBTN cohort.
[0036] FIG. 6D are graphs where the vertical axis plots both accuracy and coverage as a function of DEPLOY'S integrated model prediction score threshold, shown on the horizontal axis. Blue and gray curves illustrate the top-1 accuracy and the proportion of samples that surpass the indicated prediction score threshold (coverage), respectively. Red and orange vertical lines denote the prediction score thresholds of 0.39 and 0.46 exceeded by approximately two-thirds and approximately half of the NCI cohort samples.
[0037] FIG. 6E shows top-1 and top-2 class accuracies for tumors with high predictive score matches (>0.39).
[0038] FIG. 7A shows the diagnostic changes suggested by DEPLOY in a subcohort of diagnostically challenging cases. Tumors included from the NCI cohort in which the DEPLOY predictions differed from initial diagnosis and were consistent with the methylation class. Proportion / numbers of tumors in each of the 3 clinical categories (simple change of diagnosis, establishment of a definitive diagnosis and clinically impactful change in diagnosis (potential change in management and / or tumor grade)).
[0039] FIG. 7B is a Sankey plot showing diagnostic changes suggested by DEPLOY (and verified by the methylation classification).
[0040] FIG. 7C are treemaps contrasting DEPLOY'S classifications with original pathologists' diagnoses. The treemaps show the counts and proportion of cases where DEPLOY'S top-1 prediction corresponds with the methylation class; DEPLOY'S top-2 prediction matches the methylation class; Both DEPLOY'S top-1 and top-2 predictions diverge from the methylation class. The treemaps are shown for all 309 cases where the DEPLOY predicted class differed from the initial pathologist’s diagnosis (top section) and specifically for the subset of diagnostically challenging gliomas (n=157) (bottom section).
[0041] FIG. 8A delineates the marked differences in DEPLOY'S prediction scores for tumors from the 10 predefined types versus scores for other tumor types. The samples come from three distinct cohorts (NCI, DBTA, and CBTN) presented from left to right. The number of samples in each category is indicated in brackets. Statistical significance between the groups is evaluated using p-values from the two-tailed Mann- Whitney U test.
[0042] FIG. 8B is a graph showing DEPLOY'S accuracy in classifying tumors across the NCI, DBTA, and CBTN cohorts when considering samples with high prediction scores (higher than 0.56) from all tumor types, both the 10 predefined types and others. Error bars denote the 95% confidence intervals derived from bootstrapping.
[0043] FIG. 9A is an H&E-stained tumor with areas of oligodendroglioma and astrocytoma histology, with specific regions of interest marked by Box 1 and Box 2. Scale bar: 4mm bar depicted in FIG. 9A.
[0044] FIG. 9B shows DEPLOY tumor type predictions depicted at the tiles level, with blue representing the prediction of oligodendroglioma, and red representing the prediction of astrocytoma. Scale bar: 4mm bar depicted in FIG. 9A.
[0045] FIG. 9C shows immunohistochemistry for IDH1 -R132H, the IDH mutation present in the glioma. Scale bar: 4mm bar depicted in FIG. 9A.
[0046] FIG. 9D shows immunohistochemistry for ATRX; ATRX is normally present in all cells, but lost in most IDH-mutant astrocytomas, and retained in oligodendrogliomas. Scale bar: 4mm bar depicted in FIG. 9A.
[0047] FIG. 9E shows a higher magnification view of Box 1 in FIG. 9A demonstrating oligodendroglial morphology. Scale bar: 200-micron bar depicted in FIG. 9E.
[0048] FIG. 9F shows a higher magnification view of Box 1 in FIG. 9B demonstrating DEPLOY tile-level predictions of oligodendroglioma. Scale bar: 200- micron bar depicted in FIG. 9E.
[0049] FIG. 9G shows a higher magnification view of Box 1 in FIG. 9C demonstrating IDH1 -R132H positivity. Scale bar: 200-micron bar depicted in FIG. 9E.
[0050] FIG. 9H shows a higher magnification view of Box 1 in FIG. 9D demonstrating retention of ATRX. Scale bar: 200-micron bar depicted in FIG. 9E.
[0051] FIG. 9I shows a higher magnification of Box 2 FIG. 9A demonstrating the tumor morphology. Scale bar: 200-micron bar depicted in FIG. 9E.
[0052] FIG. 9J shows a higher magnification of Box 2 FIG. 9B demonstrating DEPLOY predictions of IDH-mutant astrocytoma. Scale bar: 200-micron bar depicted in FIG. 9E.
[0053] FIG. 9K a higher magnification of Box 2 FIG. 9C demonstrating tumor cells with astrocytic morphology positive for the IDH mutation. Scale bar: 200-micron bar depicted in FIG. 9E.
[0054] FIG. 9L shows a higher magnification of Box 2 FIG. 9D demonstrating loss of ATRX expression in tumor nuclei, with retention in nonneoplastic elements, such as endothelial cells. Scale bar: 200-micron bar depicted in FIG. 9E.
[0055] FIG. 10A provides an illustration of the training and cross-validation.
[0056] FIG. 10B provides an illustration of the testing on external data.
[0057] FIG. 11 shows an example deep learning model trained using matched slide images and DNA methylation profiles from the training cohort.
[0058] FIG. 12 is a confusion matrix of CNS tumor families showing the correlation of predicted versus actual values.
[0059] FIG. 13A and FIG. 13B show Cohen’s kappa statistics calculated from contingency tables constructed for each pair of raters.
[0060] FIG. 14A and FIG. 14B show confusion matrices for test cohorts.
[0061] FIG. 15A, FIG. 15B show overall accuracy and sum correct for family for 30 higher-confidence samples. FIG. 15C and FIG. 15D show overall accuracy and sum correct for class for 30 higher-confidence samples.
[0062] FIG. 16A, FIG. 16B show overall accuracy and sum correct for family for 96 samples, no CS cutoff. FIG. 16C and FIG. 16D show overall accuracy and sum correct for class for 96 samples, no CS cutoff.
[0063] Reference characters indicate corresponding elements among the views of the drawings. The headings used in the figures do not limit the scope of the claims.DETAILED DESCRIPTION
[0064] Various embodiments of the disclosure are discussed in detail below.While specific implementations are discussed, it should be understood that this is done for illustration purposes only. A person skilled in the relevant art will recognize that other components and configurations may be used without parting from the spirit and scope of the disclosure. Thus, the following description and drawings are illustrative and are not to be construed as limiting. Numerous specific details are described to provide a thorough understanding of the disclosure. However, in certain instances, well-known or conventional details are not described in order to avoid obscuring the description.References to one or an embodiment in the present disclosure can be references to the same embodiment or any embodiment; and such references mean at least one of the embodiments.
[0065] Reference to “one embodiment”, “an embodiment”, or “an aspect” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the disclosure. The appearancesof the phrase “in one embodiment” or “in one aspect” in various places in the specification are not necessarily all referring to the same embodiment, nor are separate or alternative embodiments mutually exclusive of other embodiments. Moreover, various features are described which may be exhibited by some embodiments and not by others.
[0066] As used herein, “about” refers to numeric values, including whole numbers, fractions, percentages, etc., whether or not explicitly indicated. The term “about” generally refers to a range of numerical values, for instance, ± 0.5-1 %, ± 1 -5% or ± 5-10% of the recited value, that one would consider equivalent to the recited value, for example, having the same function or result.
[0067] The terms “model”, “trained model”, “classifier”, “trained classifier”, “DEPLOY”, and “REDEPLOY” may be used interchangeably to refer to all or part of software executed by a processor that is trained using machine learning techniques to generate a tumor classification based on a histopathology image of a tumor sample.
[0068] The terms used in this specification generally have their ordinary meanings in the art, within the context of the disclosure, and in the specific context where each term is used. Alternative language and synonyms may be used for any one or more of the terms discussed herein, and no special significance should be placed upon whether or not a term is elaborated or discussed herein. In some cases, synonyms for certain terms are provided. A recital of one or more synonyms does not exclude the use of other synonyms. The use of examples anywhere in this specification including examples of any terms discussed herein is illustrative only and is not intended to further limit the scope and meaning of the disclosure or of any example term.Likewise, the disclosure is not limited to various embodiments given in this specification.
[0069] Additional features and advantages of the disclosure will be set forth in the description which follows, and in part will be obvious from the description, or can be learned by practice of the herein disclosed principles. The features and advantages of the disclosure can be realized and obtained by means of the instruments and combinations particularly pointed out in the appended claims. These and other features of the disclosure will become more fully apparent from the following description and appended claims or can be learned by the practice of the principles set forth herein.
[0070] Precision in diagnosis of diverse tumor types, such as central nervous system (CNS) tumor types is crucial for optimal patient treatment. DNA methylation profiles, which capture the methylation status of thousands of individual CpG sites, are data-driven means to enhance diagnostic accuracy, but this technique is expensive, time-consuming, and not yet routinely available.
[0071] Provided herein is a novel deep learning framework that predicts DNA methylation and classifies tumor subtypes using histopathology that address the challenges described above. First, a model was trained to accurately predict DNA methylation beta values from H&E images. The predicted beta values can then be used to classify tumor types using a DNA methylation-based classifier or gene expression classifier (the ‘indirect’ model). In parallel, a separate trained model can be used to directly classify tumor types from H&E images without any intermediate molecular level predictions (the ‘direct’ model). Augmenting these two models, a demographic model can be used to classify tumors by leveraging age, sex, and biopsy location. The integrated model harmonizes these indirect, direct, and demographic models. The models can be trained on tumor samples that may be subject to tumor-based methylation profiling, so as to have an objective label to compare the predicted versus actual tumor classes. The trained models can accurately predict methylation levels using tumors on a subset of CpG sites and use these highly predicted methylation levels to accurately classify samples into one of 10 tumor types. The models use deep learning to improve diagnostic accuracy from H&E-stained slides, providing a rapid assessment that may assist the pathologist in diagnostically challenging cases and advance brain cancer treatment in resource-limited areas.
[0072] Provided herein are systems and methods of classifying a tumor and methods of treatment thereof using histopathology images and one or more trained models. The trained models can be trained such that the classification method is performed automatically upon input of a histopathology image. The systems and methods of classifying the tumor can include multiple trained models to improve the speed and / or accuracy of the classification. For example, the systems and methods may include an integrated model of two or more trained models. The systems and methods described herein provide for classification of a tumor without relying onsubjectivity from a pathologist or other physician. A more accurate identification of the tumor can then lead to better clinical outcomes, e.g., classification specific treatment. The overall framework of the method 100 is shown in FIG. 1.
[0073] At step 102, the method 100 can include receiving a histopathology image of a sample of the tumor. The tumor can be any tumor associated with cancer. Nonlimiting examples of tumors include but are not limited to CNS tumors, renal tumors, hematolymphoid (HL) neoplasms, tumors of the lung, Gl tract tumors, reproductive system tumors, or any other tumors. For example, the trained models can be used in pan-cancer analysis.
[0074] In some examples, a biopsy (sample) of the tumor can be acquired from a patient. The sample can be formalin fixed paraffin embedded tissues. The histopathology image can be acquired from the sample using any techniques known in the art. For example, the histopathology image may be an H&E-stained image of the sample.
[0075] In some embodiments, the histopathology image may be preprocessed prior to being received by a processor of a system for classifying the tumor or prior to being provided to one or more trained models. The histopathology image can be preprocessed by dividing the histopathology image into non-overlapping tiles, omitting tiles comprised of more than half background, extracting features from the tiles, and / or compressing the image.
[0076] Each tile size is standardized to a set pixel size. In some examples, the tile size can be selected from about 224x224 to about 1 ,024x1 ,024 pixels. In at least one example, the tile size is about 512x512 pixels. The image can be at a set magnification level for all images to allow for consistency between images. In some examples, the magnification level may be selected from 5X, 10X, 20X, or 40X. The number of tiles in each histopathology image may vary from image to image.Consequently, each whole slide image can be represented by several thousand tiles, depending on the slide's dimensions. In some examples, tiles predominantly comprised more than half of the background may be omitted to improve efficiency of the models. Color normalization techniques can be employed to mitigate staining discrepancies across slides.
[0077] The tiles can then be subjected to feature extraction. The total number of features extracted from the tiles may range from at least about 500 to at least about 2,500 features. In one example, about 2,048 features can be extracted from the tiles. An autoencoder can then be used to compress the features to a lower-dimensional representation. In various examples, the compression may result in at least about 200 features to at least about 1 ,000 features. In one example, the compression may result in about 512 features.
[0078] At step 104, the method 100 can include providing the histopathology image to one or more trained models. In various examples, the histopathology image provided to the one or more trained models is a preprocessed histopathology image as described above.
[0079] The one or more trained models may include a methylation (methylation- indirect) model, gene expression model (expression-indirect), a direct model, and / or a demographic model. The models can be used independently to predict a tumor classification or can be combined to form an integrated model. The trained models may be combined in any combination for the integrated model. For example, the integrated model may include the methylation-indirect model, the expression-indirect model, the direct model, and the demographic model, the methylation model and the direct model, the methylation model and the demographic model, the direct model and the demographic model, or the methylation model, the direct model, and the demographic model, or other combinations thereof.
[0080] The methylation model may first be trained to predict a methylation profile of the tumor. For example, the trained methylation model may be a deep learning model trained using labeled data of matched slide images and DNA methylation profiles. In some examples, predicting the methylation profile includes predicting DNA methylation beta values.
[0081] In some examples, a Multi-Layer Perceptron (MLP) regression can be used to establish a relationship between the auto-encoded features and methylation beta values. The model can include three layers, for example, a 512-node input layer, a 512-node hidden layer, and a 2,000-node output layer. A multi-task learning strategy can be employed by clustering methylation sites based on similar median beta values.In some examples, an MLP regression model can be developed for each cluster, with each model possessing a 2,000-node output layer.
[0082] The methylation model may then be used with a trained methylation classifier to predict a classification of the tumor. In some examples, “trained methylation model” may include both the first model to predict the methylation profile and the methylation classifier to predict the classification based on the predicted methylation profile. For example, to leverage the inferred methylation beta values for tumor-type classification, machine-learning methylation classifiers were used. The inferred beta values can be first normalized to a range of 0 - 1 to ensure a consistent standard across all sites. The top 1 ,000 features (sites) can then be selected bearing the highest ANOVA F-values relative to tumor class. In some examples, up to four machine learning algorithms can be implemented for the classification task, including but not limited to Logistic Regression, Support Vector Machine, K-Nearest Neighbor, and Random Forest. The final prediction score is the average of the prediction scores of the individual models.
[0083] The gene expression model may first be trained to predict a gene expression profile e.g. mRNA expression. For example, the trained gene expression model may be a deep learning model trained using labeled data of matched slide images and RNA expression profiles. Similar methods as the methylation model may be used with the gene expression model.
[0084] To predict gene expression levels based on WSI features, tile embeddings for each WSI may be passed through an MLP designed to learn relationships between the compressed image feature representations and a subset of gene transcripts per million reads (TPM) values selected from the transcriptom ic profile of the sample represented in the WSI. A set of 20,000 genes may be selected for MLP training to maximize coverage and TPM variance across samples in the training cohort. In some examples, the MLP may include three layers: a 1 ,024-node input layer, where each node corresponded to a compressed image feature; a 1 ,024-node hidden layer; and a 2,000-node output layer, where each node corresponded to a gene. Relationships between image features and genes may be learned in sets of 2,000 genes until all 20,000 selected genes are covered.
[0085] The direct model may be trained to predict the classification of the tumor directly from the histopathology image. In the direct model, a Multi-Layer Perceptron (MLP) classifier can be used to directly link the auto-encoded features and tumor classes. This component parallels the MLP regression structure previously described, with a significant difference present in the output layer. In an example, the output layer may consist of the number nodes matching the number of tumor types that can be the classification output from the model. The output layer may include 5, 10, 15, 20, 25, 30, 35, 40, 35, 50 or more nodes. For example, when the tumor is a CNS tumor, the output layer may include 10 nodes, matching the number of the brain tumor families or types.
[0086] The demographic model may be trained to predict the classification of the tumor based on demographic data. The demographic model takes the patient's age, sex and surgical location (cerebral hemisphere, posterior fossa, dural based, ventricle, spinal cord, lumbar spinal cord) as input and predicts tumor family or types as output. Similar to the indirect model, four traditional machine learning algorithms, including but not limited to Logistic Regression, Support Vector Machine, K-Nearest Neighbor, and Random Forest can be used.
[0087] At step 106, the method may include generating one or more prediction scores. For example, the method may include generating a first, second, third, and / or fourth set of prediction scores. In some examples, only some prediction scores may be used. Each set of prediction scores from each model may be generated in parallel or sequentially. In some examples, prediction scores may not be generated for one or more models, for example the trained demographic model. Each set of prediction scores may include a prediction score for the whole image for each tumor family or type that could be selected as the classification. The tumor family or type with the highest prediction score is the final generated classification for that model. Each set of prediction scores can also include a subset of prediction scores for each tile within the image. This may allow for identifying more than one classification if more than one tumor type has a prediction score above a set threshold. In some examples, the first set of prediction scores are generated from the methylation indirect model based on the trained methylation classifier. In some examples, the second set of prediction scores are generated from the expression indirect model based on the trained geneexpression classifier. In some examples, the third set of prediction scores is based on the trained direct classification model. In some examples, the fourth set of prediction scores are generated based on the trained demographic classifier.
[0088] Because treatment plans may vary based on classification, a physician would find it beneficial to have as accurate a classification as possible. The prediction score may provide information as to the confidence of the classification. For example, a prediction score of 0.39 or greater may indicate a high confidence in the classification of the tumor. In various examples, a prediction score of greater than or equal to 0.35, 0.36, 0.37, 0.38, 0.39, 0.40, 0.41 , 0.42, 0.43, 0.44, 0.45, 0.46, 0.47, 0.48, 0.49, or 0.50 indicates a high confidence in the classification of the tumor.
[0089] At step 108, the method can include generating a first, second, third, and / or fourth classification of the tumor. Each classification from each model may be generated in parallel or sequentially. The classification can be performed automatically once the histopathology image is provided to one or more of the trained models. The models may be previously trained using labeled histopathology images or demographic data as further described below. In some examples, step 108 may be optional. For example, the first, second, third, and / or fourth prediction scores may be used to directly generate an integrated classification (step 110) without first generating a first, second, third, or fourth classification.
[0090] The classification of the tumor may be generated based on the set of prediction scores. For example, the generated classification may be the tumor family or type with the highest prediction score within the set of prediction scores for the particular model. In other examples, the classification may include more than one tumor type if more than one tumor type has a prediction score above a set threshold, as described above.
[0091] The tumor types used in the trained models may include sub-classes or be divided into sub-classes. In some aspects, the method can further include identifying a tumor sub-class / subtype based on the classification of the tumor.
[0092] The classification may be selected from a set of tumor families, types or sub-types for a particular cancer. In some examples, the classification may be selected from 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, or greater tumor types or sub-types. Insome examples, when the tumor is a CNS tumor, tumor sub-types may include but are not limited to glioblastoma, medulloblastoma, ependymoma, pilocytic astrocytoma, meningioma, astrocytoma IDH-mutant, choroid plexus, subependymoma, myxopapillary ependymoma, oligodendroglioma, and combinations thereof.
[0093] At step 110, the method may include generating an integrated classification. One or more models / prediction scores may be used to generate the integrated classification. The average of the prediction scores of the individual models represents the final prediction score. In some examples, the final prediction score may be the average of 1 , 2, 3, 4, or more prediction scores. The final prediction score may then identify the integrated classification, which may be selected from a tumor family, a set of tumor types, or a set of tumor sub-types. The tumor family, type, or sub-type with the highest averaged prediction score may be the integrated classification. In some examples, the integrated classification may include at least one tumor family, type, or sub-type or two tumor sub-types if more than one tumor type has an averaged prediction score above a threshold. In some examples, the threshold may be greater than or equal to 0.35, 0.36, 0.37, 0.38, 0.39, 0.40, 0.41 , 0.42, 0.43, 0.44, 0.45, 0.46, 0.47, 0.48, 0.49, or 0.50.
[0094] The integrated classification may be selected from the same set of tumor sub-types for the classification by each model. In some examples, when the tumor is a CNS tumor, the integrated classification may be a tumor sub-type selected from glioblastoma, medulloblastoma, ependymoma, pilocytic astrocytoma, meningioma, astrocytoma IDH-mutant, choroid plexus, subependymoma, myxopapillary ependymoma, oligodendroglioma, and combinations thereof. In other examples, the classification can include additional sub-types. The possible sub-types for selection for the classification may be included in the set of tumor sub-types.
[0095] In some aspects, one or more of steps 104 to 110 may be repeated one or more times. The method may include one or more recursive passes through the CNS tumor classifiers when multiple tumor subtypes can be distinguished within the higher- level tumor type predicted with highest confidence. For example, a first pass may identify a tumor family, with subsequent repetitions identifying a tumor type or sub-type. The method executes additional recursive steps until a tumor subtype without additionalsubtypes is reached or the confidence score for the integrated prediction does not reach the relevant threshold.
[0096] The models can be updated and improved by adding new patient data with each use and / or refining class designations based on new information. In some aspects, updating the models can include re-training the model using the updated classification database / reference set. In some examples, the method may further include updating the trained methylation model and the trained methylation classifier with the DNA methylation beta values and the integrated classification. In additional examples, the method may further include updating the trained direct classification model with the histopathology image and the integrated classification. In further examples, the method may further include updating the trained demographic model with the demographic data and the integrated classification.
[0097] Further provided herein is a method of recommending a treatment plan for the patient by diagnosing the tumor using the classification generated from the one or more trained models. The method can include forming a treatment plan specific to the diagnosis of the tumor.
[0098] The method can further include treating the patient based on the recommended treatment plan or classification. The treatment can be performed over a period of time sufficient to reduce or remove the tumor. The treatment can include treatment steps specific for the family, class or sub-class of tumor. For example, the treatment can include administering a chemotherapy drug, surgery, radiation, hormone therapy, immunotherapy, bone marrow transplantation, and / or any cancer therapy known in the art. The treatment can include combining various treatments over a period of time sufficient to reduce or remove the tumor.
[0099] Further provided herein is a method of training a model for classifying a tumor. As described herein, the methylation model may include a deep learning model trained using labeled data that includes matched slide images and DNA methylation profiles. The methylation classifier may include at least two machine learning models trained using labeled data that includes DNA methylation data and tumor types. The at least two machine learning models may be selected from logistic regression, supportvector machine, k-nearest neighbors, random forest, other known classifiers in the art, and combinations thereof.
[0100] In some examples, the direct classification model may include a deep learning model trained using labeled data comprising histopathology images and tumor types. In additional examples, the demographic model may include a machine learning model trained using labeled data comprising demographic data and tumor types. The demographic data may include age, sex, and location of the tumor. The machine learning model may include but is not limited to logistic regression, support vector machine, k-nearest neighbors, random forest, other known classifiers in the art, and combinations thereof.
[0101] The disclosure now turns to the example system illustrated in FIG. 2 which may be used to implement the methods for classifying a tumor, diagnosing a tumor, and / or training the models. FIG. 2 shows an example of computing system 200 in which the components of the system are in communication with each other using connection 205. Connection 205 can be a physical connection via a bus, or a direct connection into processor 210, such as in a chipset or system-on-chip architecture. Connection 205 can also be a virtual connection, networked connection, or logical connection.
[0102] In some examples computing system 200 is a distributed system in which the functions described in this disclosure can be distributed within a datacenter, multiple datacenters, a peer network, throughout layers of a fog network, etc. In some examples, one or more of the described system components represents many such components each performing some or all of the function for which the component is described. In some examples, the components can be physical or virtual devices.
[0103] Example system 200 includes at least one processing unit (CPU or processor) 210 and connection 105 that couples various system components including system memory 215, read only memory (ROM) 220 or random access memory (RAM) 225 to processor 210. Computing system 200 can include a cache of high-speed memory 212 connected directly with, in close proximity to, or integrated as part of processor 210.
[0104] Processor 210 can include any general purpose processor and a hardware service or software service, such as services 232, 234, and 236 stored instorage device 230, configured to control processor 210 as well as a special-purpose processor where software instructions are incorporated into the actual processor design. Processor 210 may essentially be a completely self-contained computing system, containing multiple cores or processors, a bus, memory controller, cache, etc. A multi-core processor may be symmetric or asymmetric.
[0105] To enable user interaction, computing system 200 includes an input device 245, which can represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, keyboard, mouse, motion input, speech, etc. Computing system 200 can also include output device 235, which can be one or more of a number of output mechanisms known to those of skill in the art. In some instances, multimodal systems can enable a user to provide multiple types of input / output to communicate with computing system 200. Computing system 200 can include communications interface 240, which can generally govern and manage the user input and system output, and also connect computing system 200 to other nodes in a network. There is no restriction on operating on any particular hardware arrangement and therefore the basic features here may easily be substituted for improved hardware or firmware arrangements as they are developed.
[0106] Storage device 230 can be a non-volatile memory device and can be a hard disk or other types of computer readable media which can store data that are accessible by a computer, such as magnetic cassettes, flash memory cards, solid state memory devices, digital versatile disks, cartridges, battery backed random access memories (RAMs), read only memory (ROM), and / or some combination of these devices.
[0107] The storage device 230 can include software services, servers, services, etc., that when the code that defines such software is executed by the processor 210, it causes the system to perform a function. In some examples, a hardware service that performs a particular function can include the software component stored in a computer- readable medium in connection with the necessary hardware components, such as processor 210, connection 205, output device 235, etc., to carry out the function.
[0108] The disclosure now turns to FIG. 3, which illustrates an example machine learning environment 300. The machine learning environment can be implemented onone or more computing devices 302 (e.g., cloud computing servers, virtual services, distributed computing, one or more servers, etc.). The computing device(s) 302 can include training data 304 (e.g., one or more databases or data storage device, including cloud-based storage, storage networks, local storage, etc.). In some examples, the training data 304 may be used as an input into a machine learning model or deep learning model. Training data 304 can be labeled data (e.g., one or more tags associated with the data). In some examples, the training data 304 may include data from labeled histopathology images, tumor family or types, sex, age, histological labels, and / or methylation profiles. For example, training data can be one or more histopathology images and a label (e.g., tumor classification, methylation profile, etc.) can be associated with each histopathology image. The training data 304 for each of the models may differ for each model. The training data 304 of the computing device 302 can be populated by one or more data sources 306 (e.g., data source 1 , data source 2, data source n, etc.) over a period of time (e.g., t, t+1 , t+n, etc.). The computing device(s) 302 can continue to receive data from the one or more data sources 306 until the neural network 308 (e.g., convolution neural networks, visiontransformer model, deep convolutional neural networks, artificial neural networks, learning algorithms, etc.) of the computing device(s) 302 are trained (e.g., have had sufficient unbiased data to respond to new incoming data requests and provided an autonomous or near autonomous tumor classification). In some examples, the neural network can be a convolutional neural network, for example, utilizing five layer blocks, including convolutional blocks, convolutional layers, and fully connected layers. While example neural networks are realized, neural network 308 can be one or more neural networks of various types that are not specifically limited to a single type of neural network or learning algorithm. In other examples, one or more machine learning models may be used. The machine learning models may include but are not limited to logistic regression, support vector machine, k-nearest neighbors, random forest, other known classifiers in the art, and combinations thereof.
[0109] In some examples, while not shown here, the training data 304 can be checked for biases, for example, by checking the data source 306 (and corresponding user input) versus previously known unbiased data. Other techniques for checking databiases are also realized. The data sources can be any of the sources of data for providing the input histopathology image as described above in this disclosure.
[0110] The computing device(s) 302 can receive user (e.g., physician) input 310 related to the data source. In some examples, the user input 310 and the data source 306 can be temporally related (e.g., by time t, t+1 , t+n, etc.). That is, the user input 310 and the data source 306 can be synchronous in that the user input 310 corresponds and supplements the data source 306 in a manner of supervised or reinforced learning. For example, a data source 306 can provide a histopathology image at time t and corresponding user input 310 can be tumor family, type / class, sub-class, methylation profile, or patient demographics at time t. While time t may actually be different in real- world time, they may be synchronized in time with respect to the data provided to the training data. In other examples, the data source 306 may be used in a manner of unsupervised learning without user input.
[0111] The training data 304 can be used to train a neural network 308 or learning algorithms (e.g., convolutional neural network, artificial neural network, etc.). The neural network 308 can be trained, over a period of time, to automatically (e.g., autonomously) determine what the user input 310 would be, based only on received data 312 (e.g., DNA data, methylation profiles, etc.). For example, by receiving a plurality of unbiased data and / or corresponding user input for a long enough period of time, the neural network will then be able to determine what the user input would be when provided with only the data. For example, a trained neural network 308 will be able to receive a histopathology image of a sample tumor (e.g., 312) and based on the histopathology image determine the classification of the tumor that a physician would manually identify (and that could have been provided as user input 310 during training). In some examples, this can be based on labels associated with the data as described above. The output 314 from the trained neural network can be the classification of the tumor, one or more prediction scores, or both. The classification may include more than one classification. The classification and / or prediction scores can then be used for treating a patient.
[0112] Trained neural network system 316 can include a trained neural network 308, received data 312, and output 314. The received data 312 can be information related to a patient, as previously described above. The received data 312 can be used as input to the trained neural network 308. Trained neural network 308 can then, based on the received data 312, classify the received data and / or determine a recommendedcourse of action for treating the patient, based on how the neural network was trained (as described above). The recommended course of action or output 314 of trained neural network 308 can include a prediction of the classification of the tumor of the patient to which the received data 312 corresponds. In other instances, the output 314 from the trained neural network can be provided in a human readable form, for example, to be reviewed by a physician to determine a course of action.
[0113] For clarity of explanation, in some instances the present technology may be presented as including individual functional blocks including functional blocks comprising devices, device components, steps or routines in a method embodied in software, or combinations of hardware and software.
[0114] In some embodiments the computer-readable storage devices, mediums, and memories can include a cable or wireless signal containing a bit stream and the like. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals per se.
[0115] Methods according to the above-described examples can be implemented using computer-executable instructions that are stored or otherwise available from computer readable media. Such instructions can comprise, for example, instructions and data which cause or otherwise configure a general purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. Portions of computer resources used can be accessible over a network. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, firmware, or source code. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during methods according to described examples include magnetic or optical disks, flash memory, USB devices provided with non-volatile memory, networked storage devices, and so on.
[0116] Devices implementing methods according to these disclosures can comprise hardware, firmware and / or software, and can take any of a variety of form factors. Typical examples of such form factors include laptops, smart phones, small form factor personal computers, personal digital assistants, rackmount devices,standalone devices, and so on. Functionality described herein also can be embodied in peripherals or add-in cards. Such functionality can also be implemented on a circuit board among different chips or different processes executing in a single device, by way of further example.
[0117] The instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are means for providing the functions described in these disclosures.ExamplesExample 1: I tegrated deep learning model for predicting DNA methylation and tumor types from histopathology in central nervous system tumors
[0118] Deep learning was employed to accurately classify CNS tumors and it was reasoned that the development and integration of independently derived and complementary models would outperform single models applied in isolation. FIG. 4A summarizes the workflow. First methylation profiling was conducted on 1 ,796 tumor samples during the course of clinical case consultation of CNS tumor diagnosis at the National Cancer Institute (NCI). Study inclusion required the presence of a matched H&E image, and the cohort consisted of samples with high-confidence methylation matches for 10 distinct CNS tumor categories. The deep learning models were trained on this dataset of H&E slide images and their corresponding genome-wide methylation profiles. For external validation, three independent datasets were harnessed from the publicly available Digital Brain Tumor Atlas (DBTA) and the Children's Brain Tumor Network (CBTN), and the NCI-Prospective collected following the development and creation of the model. While DNA methylation data were not available for the DBTA set, DNA methylation profiling was performed on 492 CBTN and 286 NCI-Prospective samples to develop robust methylation profiling-defined accuracy for these datasets. Each case was reviewed by at least one of two board-certified neuropathologists (M.P.N. and / or K.A.) to verify the accuracy of their diagnoses. The ten tumor type categories selected for inclusion in the study were based on sample size in the datasets and resulted in a total of tumor samples from 3,952 patients (1 ,796 patients from theNCI; 1 ,522 patients from the DBTA; 348 patients from the CBTN; 286 patients from the NCI-Prospective). FIG. 4B summarizes the cohorts.
[0119] To analyze these datasets, three specialized class / label predictors / models were built, an “indirect” predictor / model that makes predictions from methylation values inferred from slides, a “direct” predictor / model that makes predictions directly from the slides, and a “demographic” predictor / model that makes predictions from demographic variables. The individual scores of these three predictors are subsequently integrated to formulate the final, integrated predictor / model (“DEPLOY”).
[0120] Specifically, (i) to construct a DNA methylation predictor from the H&E slides, a deep learning model (sketched in FIG. 4A and described in detail in Examples 2-5) was trained using matched slide images and DNA methylation profiles from the inhouse NCI cohort. The indirect model then utilizes four distinct classical machine learning algorithms to train and predict tumor types using the inferred DNA methylation data, (ii) The direct model employs another deep learning classifier to directly train and predict tumor types based on slide images, (iii) Lastly, the demographic model takes three specific variables, age, sex, and tumor location, to classify tumor types. The prediction scores from the three models were then integrated by taking their mean scores to create a unified predictive model, which outperforms each of the three individual class prediction models. In some examples, the mean scores do not include any further training or weighting. To assess the performance of the models, a 5x5 nested cross-validation was employed for both the direct and indirect models. The NCI dataset was initially partitioned into five outer folds. Each of these folds was further subdivided into five inner folds. The stratification of these folds was performed according to cancer type, thereby maintaining a balanced representation of all classes. The training and evaluation process was iterative. Each outer fold, accounting for 20% of the total data, was successively held out as a test set, and the remaining four outer folds, representing 80% of the data, were used for model training. These training data of each fold was used to train 5 different models, where for each model a different inner fold was utilized as a validation set. FIG. 10A provides an illustration of the training and cross-validation. Second and importantly, to validate the model, the DEPLOY classifier was then evaluated on external H&E images from 1 ,522 patients, 348 patients and 286patients in the DBTA and CBTN cohorts, respectively. The outputs from the 5 models, each trained excluding a different inner fold, were averaged for the samples of a given outer fold. For the external validation sets (DBTA, CBTN, NCI-Prospective), the predictions from all 25 models (five models for each of the five outer folds) were averaged to obtain the final prediction. FIG. 10B provides an illustration of the testing on external data. For traditional machine learning algorithms used to infer cancer types from predicted beta values, the same outer folds as those used for deep learning models were employed. However, as these models do not necessitate a validation set, the inner folds were not utilized.Predicting DNA methylation from H&E slides
[0121] Leveraging the NCI dataset's matched H&E slides and methylation profiles, the indirect model was constructed to predict DNA methylation from H&E slides. Model input probes were chosen based on the variance of their methylation levels, yielding a pool of 65,591 probes for DEPLOY methylation prediction whose variance on the NCI training set exceeds 0.2. A five-fold cross-validation approach was employed, iteratively training the model on 80% of the dataset and assessing its performance on the remaining 20% (see above for description of the cross-validation scheme). The model's prediction accuracy was evaluated by computing the Pearson correlation between actual and predicted methylation values for each probe across all samples. The median correlation was 0.44 (FIG. 5A). 42,231 probes displayed a correlation above 0.4, and 4,027 probes surpassed a correlation of 0.6 (FIG. 5B). Of interest, these findings have higher accuracy than prior findings on transcriptom ics prediction from H&E images.
[0122] To further substantiate the model's performance and showcase its generalizability, methylation profiling was performed on tumors from the CBTN and NCI-Prospective datasets, which consisted of available tumor DNA and H&E images on childhood CNS tumors. The model — trained on the NCI cohort — was employed to predict methylation profiles from the images within this distinct dataset. Remarkably, the median Pearson correlation between the actual and predicted beta-values for each probe, computed across all samples within each cohort, was 0.44 and 0.47 for the CBTN and NCI-Prospective cohorts, respectively (FIG. 5A). Importantly, the number ofwell-predicted probes were similar in these validation datasets, with 40,133 and 45,755 probes exceeding a correlation coefficient of 0.4 in the CBTN and NCI-Prospective cohorts, respectively. Moreover, 8,293 probes in the CBTN cohort and 8,400 in the NCI- Prospective cohort surpassed a correlation threshold of 0.6 (FIG. 5B). A pronounced overlap was observed among the cohorts in relation to highly predicted probes: Of the 52,962 probes that attained a correlation exceeding 0.4 in any cohort, 84% (44,442) surpassed this threshold in at least two of the cohorts. Moreover, 63% (33,215) of these probes consistently exceeded this correlation benchmark across all three cohorts (FIG. 5C). Such consistent performance across distinct datasets underlines the robustness and reliability of the model, affirming its potential for application across diverse datasets.
[0123] It was examined whether the predicted methylation values recapitulate known global relationships of the methylation landscape across common CNS tumor types. It was found that the predicted methylation data demonstrates the expected differential hypermethylation in IDH-mutant gliomas (n = 233) compared to IDH-wildtype glioblastoma (n = 379) in the NCI cohort (FIG. 5C). Remarkably, the analysis based on the predicted values resulted in an 85% accuracy rate in identifying sites differentially methylated between the two tumor types as determined by the actual measured methylation arrays (FIG. 5D), highlighting the model's ability to accurately capture established patterns of DNA methylation.
[0124] The alignment between the predicted and actual methylation values was then evaluated with respect to key cancer hallmarks. A pathway enrichment analysis was performed using established cancer hallmark pathways. The probes were linked with their associated gene bodies or promoters, guided by the annotations provided by Illumina. Beta values were then aggregated for each gene promoter or gene body on a per-sample basis within the NCI cohort, generating promoter-level methylation values for 6,316 unique genes, along with gene body-level methylation values for 8,401 unique genes (considering the divergent impacts of promoter and gene-body methylation on gene expression, these values were analyzed separately). Building on these values, a single-sample gene set enrichment analysis (ssGSEA) was carried out to gauge the methylation levels of cancer hallmark pathways whose genes were represented within the gene set. Remarkably, the predicted beta values yielded highly consistent pathwayactivation patterns with those inferred from the measured values across diverse cancer types (FIG. 5F, 5G), further testifying to the biologically meaningful strong correlation between them (FIG. 5H). These findings underscore DEPLOY'S ability to predict crucial cancer-related methylation events from H&E images with good accuracy.Classifying the CNS tumor type from H&E slides
[0125] The four classification models (demographic, direct, indirect and integrated) were initially aimed at CNS tumor type stratification. Examining the entire NCI cohort, the direct model accuracy, as evaluated via the area under the precisionrecall curve (ALIPRC, aka average precision) was 0.77, with the demographic model a slightly lower AUPRC of 0.76. The indirect model outperformed both, achieving an AUPRC of 0.82, testifying to the value of harnessing predicted CpG site methylation levels for tumor diagnosis and classification. Importantly, integrating the predictions of the demographic, direct and indirect models yielded a further marked improvement, yielding an AUPRC of 0.92 (FIG. 6A).
[0126] The output of the DEPLOY model provides a ranked order of diagnostic possibilities for each tumor, based on the scores assigned by the model. Model performance was further evaluated using a top-k differential diagnosis accuracy, counting how often the methylation profiling result label was found in the k highest confidence predictions of the model, using the top-1 and top-2 accuracies. Consistently, the indirect model performs better than the direct model and the demographic model, and the integration of the three produced the best results: achieving an accuracy of 85% for top-1 and 94% for top-2 across all samples, regardless of prediction score (FIG. 6B). These top-tier predictions reduce the range of 10 possible diagnoses to a more manageable set of potential classes, thereby facilitating a focused differential diagnosis that could be directly evaluated and tested by the pathologist.
[0127] To assess the performance of DEPLOY models in an independent external validation, they were applied as is to predict CNS tumor types on the slides from the DBTA cohort. Surprisingly, a very similar performance was obtained for this cohort, achieving an AUPRC of 0.91 (FIG. 6A) and overall accuracies of 84% for top-1 , 93% for top-2 (FIG. 6B) with the integrated model. Applying DEPLOY to the CBTN dataset yielded an AUPRC of 0.93 (FIG. 6A), accompanied by top-1 and top-2accuracies of 83% and 94%, respectively (FIG. 6B). Applying DEPLOY to the NCI- Prospective dataset yielded an AUPRC of 0.87 (FIG. 6A), accompanied by top-1 and top-2 accuracies of 79% and 92%, respectively (FIG. 6B). Further granularity of the types of errors made between different class labels from the integrated model are presented in the confusion matrices in FIG. 6C. Of note, the CBTN dataset is focused on pediatric tumors and thus does not include adult tumor types, as observed within the CBTN confusion matrix. To reiterate, these findings on the external validation sets were achieved without any training on this specific dataset, attesting to the generalizability of DEPLOY.
[0128] Beyond providing classifications for tumor types, DEPLOY also provides a numerical prediction score that serves to enhance diagnostic precision. Prediction scores are an integral component of methylation-based diagnosis of CNS tumors in the clinic, where threshold scores are set to mark a reasonable cutoff for a specific prediction. Methylation-based confidence thresholds are at levels that are reached by approximately 2 / 3 of the tested samples as a practical trade-off of coverage versus prediction accuracy. Importantly, there is a monotonic increase in DEPLOY'S accuracy as its prediction scores increase, as evidenced in FIG. 6D. To further elucidate the relationship between prediction scores and accuracy, score thresholds were identified (0.39 and 0.46) that capture the top two-thirds and top 50% of samples within the NCI cohort, respectively. Remarkably, when focusing on the 0.39 threshold, accuracies of the top-1 predicted class in the DBTA, CBTN, and NCI-Prospective cohorts are 96%, 94%, and 95% respectively. Results using this threshold are shown for the individual and integrated models for all 3 datasets (FIG. 6E). Recognizing that higher thresholds could further improve prediction accuracy but at a cost of lower overall coverage of the dataset, when the threshold is set to include top 50th percentiles of prediction scores, accuracies in the DBTA, CBTN, and NCI-Prospective datasets are 99%, 96% and 98%, respectively (FIG. 6D). This feature underscores DEPLOY’S capability to provide reliable, high-confidence diagnosis for CNS tumors within the 10- classes under study, for those cases that receive high predictive scores.
[0129] To further study the potential impact for DEPLOY on a sub-cohort that is enriched for diagnostically difficult cases, cases where the DEPLOY predicted classdiffered from the original diagnosis rendered by the original case pathologists were examined in individual cases. Within the NCI cohort, there were 311 such cases, where a high-confidence (>0.39) DEPLOY prediction differed from the original diagnosis. Among these 309 cases, 261 were concordant with the methylation class (which differed from the original classification given by the pathologist, as mentioned). For each of these 263 cases, the potential clinical impact of the DEPLOY prediction was estimated by the type of diagnostic change (FIG. 7A-7B), ranging from a simple change in diagnosis (limited clinical impact) to a clarification of diagnosis (change from a general ‘descriptive’ label to a definitive label) or a clinically impactful change in diagnosis (involving a change in tumor grade and / or clinical management plan). For 42 of these 263 cases the diagnostic changes recommended by DEPLOY were estimated to have a narrow clinical impact (for example ependymoma to myxopapillary ependymoma). However, a majority of these cases (n=205) involved a clarification of a diagnosis from descriptive terms to a more definitive tumor class (for example “Embryonal neoplasm” to medulloblastoma), a category of change expected to have clinical relevance. Finally, for 1 cases, the change suggested by DEPLOY would have been expected to have clinically importance, by virtue of a change in tumor grade and / or patient management plan (for example GBM (grade 4) to O-IDH (grade 2-3), or PA (grade 1 ) to GBM (grade 4).
[0130] For the remaining 48 cases where the top-1 results of DEPLOY were discordant with the initial pathology diagnosis, it was found that the Top-2 (second highest score) matched the methylation profiling result in 35 / 48 cases and was misleading in the remaining 13 cases (Figure 4S). All in all, the Top 1 -2 predictions made by DEPLOY would have correctly narrowed down the possible labels to consider in this cohort of diagnostically difficult samples in 298 / 309 cases (96%). Overall, among the 309 cases where DEPLOY made high-confidence calls that diverged from the pathologist's, DEPLOY'S top-1 prediction aligned with the methylation class in 84% of the cases (261 out of 309). Encouragingly, for the remaining 48 cases, where the pathologist's diagnoses were consistent with the methylation class, DEPLOY'S top-2 prediction matched the methylation class in 35 instances (see FIG. 7C).Assessing DEPLOY’S performance in out-of-training-scope settings
[0131] In the previous sections, there was focus on the 10 most prevalent tumor types, as there was sufficient sample sizes for both training and validation (at least 60 samples in the NCI and 15 samples in the DBTA). Next, the generalizability of DEPLOY to a full neuro-oncology practice was evaluated, where some samples could come from less frequent tumor types not included among the 10 types. The analysis was extended to encompass all available samples across the three cohorts, regardless of their sample size (notably, the 10-class predictors already covered approximately 70% of the samples in these cohorts). To this end, models that were trained on the 10 tumor types were applied to compute the confidence score for the samples belonging to the “other types”. Reassuringly, tumors from “other types” tended to receive significantly lower DEPLOY prediction scores than that from the 10 tumor types. This finding highlights the discriminative capability of DEPLOY'S prediction scores in identifying the 10 predefined tumor types from other tumor types (FIG. 8A).
[0132] Building on this observation, the evaluation was expanded to include all samples (from both the 10 predefined tumor types and all other classes), provided they exceeded a DEPLOY prediction score threshold of 0.56. This threshold was established based on the top tercile of prediction scores within the NCI training cohort. DEPLOY demonstrated robust prediction accuracies across all cohorts: 96% in the NCI, 92% in DBTA, and 94% in CBTN (FIG. 8B). Evidently, as there is currently not sufficiently available data to train classifiers on these rare tumor types, DEPLOY can currently maintain high accuracy in this general case by providing coverage of about one third cases. However, these results testify to its promise when more such data becomes available.Detecting and depicting spatial heterogeneity of diagnostic class in individual tumors
[0133] DEPLOY provides prediction scores at the tile level, which are then aggregated to make the final prediction at the slide level. This approach puts DEPLOY in a quite intriguing position, where it is potentially able to uncover spatial heterogeneity in tumors and in some cases diagnose multiple tumor types that are present within a single slide. Notably, this feat is not possible with measured methylation-basedclassifiers, which base their predictions on the values measured in bulk across the sample as a whole.
[0134] To investigate this potential capability, DEPLOY models trained on the NCI cohort were used to predict tumor type at tile level for a molecularly-proven dualgenotype oligoastrocytoma. The top 1 prediction by the integrated model was A-IDH (IDH-mutant astrocytoma), with a score of 0.32, and the top 2 prediction was O-IDH (oligodendroglioma), with a score of 0.24. Although the top 2 predictions matched the true diagnosis, the scores do not reach the threshold of 0.39, potentially indicative of the appropriate inability of the model to choose one tumor diagnosis. However, at a tile level, DEPLOY’S predictions spatially match the histologic and molecular ‘ground truth’, as demonstrated by comparison to the hematoxylin and eosin-stained tissue section used for prediction and immunohistochemical staining patterns that correlate with gene variants (FIG. 9A-9L). Despite some ambiguity of the morphology (FIG. 9I), which shows some perinuclear haloes and fine vasculature as commonly seen in oligodendrogliomas, DEPLOY predicts astrocytoma (FIG. 9J). The direct model, without the integration of the indirect and demographic model, gave a top-1 prediction of A-IDH, but a top-2 prediction of PA (pilocytic astrocytoma), rather than O-IDH, which was the top-3 prediction. This could be due to some overlap in histologic features between PAs and O-IDH. However, the indirect model provided correct predictions, as did the integrated model. This example highlights the performance of the indirect model, as observed above.
[0135] Identifying heterogeneity can be useful for several different tumor types. For example, multiple targeted therapies are in development for glioblastoma, which are notorious for tumor heterogeneity and limited response to treatment. Molecular testing of these tumors is generally performed on a single tissue sample, and, hence, will not capture the full complexity of the molecular profile of the tumor. DEPLOY has the potential to analyze the entire tissue for tumor subtypes and therapeutic targets, allowing a patient’s treatment to be precisely designed for greater efficacy. Similarly, grading of meningiomas is challenging, with the field converging on molecular assessment as the most accurate prognostic method. With the addition of meningioma subtypes to the model, DEPLOY will be able to add spatial detail to the histologic andmolecular assessment of these tumors, which will aid clinicians in difficult treatment decisions, such as whether to use radiation.Discussion
[0136] DEPLOY is a deep learning algorithm developed to identify the specific type of tumor based on digitized routine histopathology slides. The determination of CNS tumor diagnoses can be difficult, based in part on the considerable number of diagnostic possibilities. Accurate diagnosis of individual cases can challenge even the most expert of neuropathologists. Prior studies have determined high inter-observer variability, especially with less common tumor types in both adults and children.
[0137] Genome-wide DNA methylation-based tumor classification is based on the principle that distinct CNS tumor types have unique methylation signatures, by virtue of their cell of origin and / or additional epigenomic changes that occur during oncogenesis. The combination of methylation-based classification along with conventional histopathology and ancillary tests can refine tumor diagnoses and not infrequently lead to unexpected changes in the final diagnosis. However, widespread use of methylationbased classification has several challenges. First, the time required for methylationbased diagnostic testing can be a significant drawback, often taking several weeks or more, depending on laboratory workflows. Patients with high-grade CNS neoplasms often require therapeutic decisions sooner than the time frame required for testing. Second, and most importantly, resource-constrained regions suffer from a lack of availability of these molecular tests, at nearly all hospitals across the globe, providing a very strong impetus to improve tumor diagnostic accuracy from the widely available H&E slides. DEPLOY leverages new insights gained by DNA methylation profiling for CNS tumor diagnostics, without itself requiring methylation profiling to be performed, thereby removing a potential barrier inherent to resource-limited areas. Therefore, DEPLOY allows access to the benefits of molecular testing while bypassing these challenges.
[0138] In In most cases, tumors within 1 of the 10 diagnostic categories currently in DEPLOY can readily be distinguished from each other by an experienced neuropathologist, and many diagnostic questions can be resolved with the wide range of ancillary immunohistochemical tests that are currently available at many moderncenters. However, specific cases elude definitive diagnosis, a problem magnified at centers without an experienced neuropathologist and / or lacking in availability of specific ancillary tests. Cases remain where a diagnosis can be narrowed to a general tumor family (for example ‘Embryonal neoplasm” or “Glioma, NOS”) where an additional tool could contribute to diagnostic specificity (FIG. 7B). In addition, occasional cases are encountered where a classifier such as DEPLOY could prompt a clinically significant change in diagnosis from one tumor type to another; indeed, DEPLOY would have prompted clinically significant changes in WHO tumor grade where, for example isolated cases initially designated as GG and PA (both grade 1 ), were found to represent GBM (grade 4). DEPLOY may be used as an additional tool in the diagnostic armamentarium that, depending on the context, provides confirmation of a suspected diagnosis or prompts reconsideration of a diagnosis, similar to the current role of DNA methylation profiling in neuropathology.
[0139] The present Examples leverage existing DNA methylation data in two important ways. First, since methylation-based classification is an objective process (in contrast to standard pathology diagnoses in the clinic, which can be subject to error a model built from labels derived from DNA methylation classes) that may be more reproducible across centers. In this regard while two datasets had labels based on methylation profiling, the third (DBTA) did not. However, in the DBTA data, each case was reviewed by at least two (and sometimes three) neuropathologists to obtain diagnostic consensus, thereby reducing errors in labeling, a process that is not practical in many centers. Second, the existence of such DNA methylation data allowed a model to be constructed and tested to predict individual DNA methylation beta values from the histopathology image, which could then, in turn, be used to build a classifier.
[0140] The model integrates direct, indirect, and demographic classifiers. Perhaps surprisingly, the indirect model (classifier result based on inferred methylation data) performance exceeds that of the direct method with better average precision and accuracy for the three analyzed cohorts and shows promise as a means to pursue additional refinements in an H&E-based CNS tumor classifier, as well as for potential expansion of this method to future classifier projects in additional organ systems. Furthermore, the integrated model outperforms each of the individual models. Thesefindings underscore that the indirect approach is not only robust as a standalone solution but also synergizes seamlessly with the more conventional techniques.
[0141] Inherent in this approach is the requirement for a matched histopathology image with methylation data for the development of predictors involving new tumor types. It is worth noting that the pace of advancements in methylation-based classifiers for additional tumor types is accelerating suggesting that such datasets will likely become more widely available in the near future on which such classifiers could be developed. An additional attribute of the model was that it was trained on H&E’s from a wide variety of medical centers across the USA, and internationally. Since tissue processing and staining techniques can show variation across centers, the use of a dataset from highly diverse sources expands the likelihood that it will be generally reliable, and not dependent on specific tissue processing protocols that may be specific to individual pathology laboratories. The fact that accuracies remained high (FIG. 6E) on validation / unseen datasets from independent sources in the USA (CBTN) as well as Europe (DBTA) adds to the robust nature of the model.
[0142] DEPLOY’S unique ability to provide tile-level prediction scores, which are then aggregated for slide-level diagnoses, presents an opportunity to explore spatial heterogeneity. While existing methylation-based classifiers generate predictions from bulk measurements, thereby obfuscating intratumoral variability, DEPLOY can potentially identify multiple tumor types or subtypes within a single slide. This was demonstrated when DEPLOY was used to classify a molecularly confirmed dualgenotype oligodendroglioma-astrocytoma (FIG. 9A-9L). DEPLOY holds potential for shedding light on intratumoral variability, a critical yet often overlooked aspect of cancer diagnostics.
[0143] The present example includes 2 external validation sets, where DEPLOY was applied to datasets that were unseen during model development, where accuracy levels were achieved that could contribute to a clinical decision, if implemented. DEPLOY was designed to achieve very high accuracies in a top-1 model, as this is most relevant to the clinical need of a pathologist to come to a specific diagnosis. Such high accuracies may have been achieved based upon several factors: 1 ) the use of methylation-based objective labels in model development; 2) the finding that inferredmethylation levels could be applied to a predictive model (the ‘indirect’ model) as well as the combination of multiple models (direct, indirect and demographic) into a single integrated model that achieves accuracies higher than any of its single component models. Overall, DEPLOY fulfills a previously unmet need through the development of a diagnostic model showing high predictive accuracies (94-96%) of 10 major CNS tumor types on external datasets that could be an assistive tool for pathologists in specific real-world settings.
[0144] Overall, DEPLOY provides a first of its kind classifier of CNS tumors at an unprecedented high resolution. It leverages advances in methylation-based classification of CNS tumors along with deep learning from histopathology images to potentially improve diagnostic accuracy with the enhancement of including spatial information, in a time and cost-efficient manner. DEPLOY is positioned to serve as a complementary pathology reader to reassure and confirm the pathologist’s initial diagnosis, or prompt revaluation if a disparity is found. The number of validated tumor types may be expanded, both within and beyond the CNS to extend the clinical potential of DEPLOY. In some examples, DEPLOY may be used by those working in underserved areas, by sending a scanned image of a tumor slide to a DEPLOY platform to get an Al-based “second opinion” for their consideration, wherever they are located. In this context, the architecture of DEPLOY leverages new insights gained by DNA methylation profiling for CNS tumor diagnostics, without itself requiring methylation profiling to be performed on new samples, thereby removing a potential barrier inherent to resource-limited areas.Example 2: Methods Data Collection
[0145] DEPLOY was trained using both newly generated and publicly available datasets.
[0146] The NCI histological images, their corresponding methylation profiles, and demographics were newly generated in the Laboratory of Pathology at the NCI. This dataset consists of 1 ,796 patients comprising 10 CNS tumor types.
[0147] The publicly available Digital Brain Tumor Atlas (DBTA) histopathological images and demographics were downloaded from Roetzer-Pejrimovsky et al. 2022. Only patients with diagnostic labels from the 10 CNS tumor types were selected, making a total of 1 ,522 patients.
[0148] Samples on methylation arrays from The Children's Brain Tumor Network (CBTN) were profiled and profiled samples were matched with available histopathological images. Only tumors with diagnostic labels from the 10 CNS tumor types were selected, for a total of 348 patients with matched methylation profiles and histopathology.
[0149] The NCI-Prospective cohort consisted of histopathological images, methylation profiles and demographics from 286 patients. This dataset was collected during the clinical consultation practice at the NCI, following the development and creation of the DEPLOY model.Normalization
[0150] For the normalization of data acquired from the Illumina Human Methylation EPIC 850K DNA methylation array, a stringent normalization strategy was implemented based on guidelines put forth by Capper et al. Firstly, the raw intensity data gleaned from the microarrays with the "minfi" library were processed. Specifically, the "preprocesslllumina" function was utilized to conduct background correction and initial normalization via control probes. This foundational step was pivotal for reducing inherent noise in the dataset. To enhance material type prediction, thereby improving normalization quality, the "MNPgetFFPE" function was employed. This additional layer of granularity further solidified the robustness of the normalization process. Addressing batch effects, commonly encountered in large-scale experimental setups, was the next crucial step. The "MNPbatchadjust" function was used for this purpose, ensuring that variations arising from different data batches were appropriately mitigated. After these steps, DNA methylation levels were extracted — represented as beta values — using the "getBeta" function. These beta values served as the basis for all subsequent analyses. Lastly, additional quality control was performed by filtering out a total of 124,955 probes that were linked with known artifacts affecting performance.
[0151] To concentrate on probes with highly variable beta values, 51 ,052 of the initial 865,859 probes were filtered out due to missing measurements. Probes were then selected demonstrating a standard deviation of at least 0.2, reducing the pool to 130,285 probes. In addition, a balance criterion was implemented to secure the selection of probes that exhibited a relatively balanced distribution of methylation levels across the samples. Specifically, the proportion of samples with methylation beta-values falling below 0.5 and above 0.5 for each probe was computed. Probes were deemed balanced, and thus appropriate for further analysis, if both proportions fell beneath a predefined imbalance threshold set at 0.7. This led to a final selection of 65,591 probes. This strategic selection allowed us to exclude probes that were consistently hypomethylated or hypermethylated across the sample set. Consequently, the investigation was concentrated on genomic regions displaying substantial variability in methylation patterns among samples.Example 3: DEPLOY computational framework
[0152] The DEPLOY architecture consists of six main components (FIG. 4A): image processing, feature extraction, feature compression, indirect model, direct model, and demographic model.Image processing
[0153] The image pre-processing phase was initiated by dividing the entire slide images into non-overlapping small images known as tiles. Each tile size was standardized to 512x512 pixels across all cases. A magnification level of 20X was chosen for analysis. Tiles that predominantly comprised more than half of the background were omitted. Consequently, each whole slide image could be represented by several thousand tiles, depending on the slide's dimensions. Color normalization techniques were employed to mitigate staining discrepancies across slides.Feature extraction
[0154] Tile images were processed using the pre-trained ResNet50 model to distill image features. Each tile was thus represented by a 2,048-dimensional ResNet feature vector. After preprocessing, each whole slide image was represented as a matrix of size (n_tiles, 2,048).Feature compression
[0155] Given that the ResNet50 model was initially trained on natural images, it is likely that some features may not contribute significantly to pathological image analysis and remain unused. To mitigate this, an autoencoder was employed to reduce the 2,048 features derived from ResNet50 to a more manageable and information-rich 512- dimensional space. This approach not only reduces noise and mitigates overfitting but also considerably diminishes computational demands. The employed autoencoder is composed of an encoder and a decoder. The encoder, a fully connected layer followed by a ReLU activation function, maps the input vector of 2,048 dimensions down to a 512-dimensional space. Subsequently, the decoder applies another fully connected layer followed by a ReLU function, translating the 512-dimensional representation back into the original 2,048-dimensional space.
[0156] After these steps, the resulting auto-encoded features are used as input for both the indirect and direct models, as detailed below. The model yields tile-level predictions. Since the targets (beta values, tumor classes) for both indirect and direct models are available at the slide level, the tile-level predictions were averaged to produce slide-level predictions for training and evaluations. The original tile-level predictions were utilized to investigate spatial heterogeneity (FIG. 9B).Indirect model
[0157] The first phase in this model is to use a Multi-Layer Perceptron (MLP) regression to establish a relationship between the auto-encoded features and methylation beta values. The model comprises three layers: a 512-node input layer, a 512-node hidden layer, and a 2,000-node output layer. A multi-task learning strategy was employed by clustering methylation sites based on similar median beta values. An MLP regression model was developed for each cluster, with each model possessing a 2,000-node output layer.
[0158] Subsequently, to leverage the inferred methylation beta values for tumortype classification, traditional machine-learning classifiers were used. The procedure began within the training fold, where the inferred beta values were first normalized to a range of 0 - 1 using sklearn's MinMaxScaler. This step ensured a consistent standard across all sites. Next, within the same training fold, the top 1 ,000 features (sites) wereselected bearing the highest ANOVA F-values relative to tumor class (sklearn's SelectKBest and f_classif). Four traditional machine learning algorithms were implemented for the classification task: Logistic Regression, Support Vector Machine, K-Nearest Neighbor, and Random Forest. The final prediction score is the average of the prediction scores of the individual models.Direct model
[0159] In this model, a Multi-Layer Perceptron (MLP) classifier was used to directly link the auto-encoded features and tumor classes. This component parallels the MLP regression structure previously described, with a significant difference present in the output layer. Consisting of 10 nodes, the output layer matches the number of the brain tumor types present in the NCI dataset.Demographic model
[0160] The demographic model takes the patient's age, sex and surgical location (cerebral hemisphere, posterior fossa, dural based, ventricle, spinal cord, lumbar spinal cord) as input and predicts tumor types as output. Similar to the indirect model, four traditional machine learning algorithms, including Logistic Regression, Support Vector Machine, K-Nearest Neighbor, and Random Forest were employed. The average of the prediction scores of the individual models represents the final prediction score. Before feeding the data into the classifier, the patient’s ages were normalized using MinMaxScaler.Example 4: Model training procedureData consolidation and classification schemas
[0161] To get a higher resolution and more samples for training, the classification models were trained with the entire NCI cohort, consisting of 3,088 samples from 76 methylation-based classes. The methylation-based classes were then aggregated to the 29 WHO-defined classes for evaluation purposes.Cross-validation approach
[0162] A 5x5 nested cross-validation was employed for both the direct and indirect models. The NCI dataset was initially partitioned into five outer folds. Each ofthese folds was further subdivided into five inner folds. The stratification of these folds was performed according to cancer type, thereby maintaining a balanced representation of all classes.
[0163] The training and evaluation process was iterative. Each outer fold, accounting for 20% of the total data, was successively held out as a test set, and the remaining four outer folds, representing 80% of the data, were used for model training. These training data of each fold was used to train 5 different models, where for each model a different inner fold was utilized as a validation set.
[0164] To generate predictions for the NCI dataset, the outputs from the 5 models, each trained excluding a different inner fold, were averaged for the samples of a given outer fold. For the external validation sets (DBTA, CBTN, and NCI-Prospective), the predictions from all 25 models (five models for each of the five outer folds) were averaged to obtain the final prediction.
[0165] For traditional machine learning algorithms used to infer cancer types from predicted beta values, the same outer folds as those used for deep learning models were employed. However, as these models do not necessitate a validation set, the inner folds were not utilized.Model training and optimization protocols
[0166] All deep-learning models were trained using a strategy consisting of stochastic gradient descent coupled with mini-batches. For the optimization process, Adam optimizer was used, specifying an initial learning rate of 0.0001.
[0167] The Mean Square Error between predicted and actual values serves as the loss function for the autoencoder and indirect models, while the cross-entropy loss was employed for the direct model.
[0168] For both the indirect and direct models, multi-layer perceptron (MLP) regressors and classifiers underwent training for 500 epochs. To mitigate the risk of overfitting, a dropout rate of 20% was applied at the initial layer. Additionally, an early stopping strategy was employed to optimize computational efficiency and prevent overfitting. This mechanism halted the training process if no improvement was observed in the mean correlation between predicted and actual beta values on the validation setfor the indirect model, or in the accuracy for the direct model, after 50 epochs (for the indirect model) and 30 epochs (for the direct model).
[0169] The hyperparameters of classifiers that infer tumor class from the predicted beta values were fine-tuned using 5-fold cross-validation within the training fold. This optimization was guided by the micro-averaged F1 score, facilitated by sklearn's GridSearchCV.Example 5: Analysis of predicted methylation levels and evaluation of classification models
[0170] In the process of pathway enrichment analysis, the first step involved mapping the collection of 65,591 probes onto their corresponding gene body or promoter, making use of Illumina's annotation. The outcome of this mapping task yielded 23,941 probes that were linked to the body of 8,401 unique genes, along with 10,709 probes associated with the promoters of 6,316 distinct genes.
[0171] Subsequently, the mean beta values were computed for each gene promoter and gene body across all samples, deriving methylation values at both the promoter and gene body levels. Following this, a single-sample gene set enrichment analysis (ssGSEA) was conducted on these values. This analysis was performed using the python package GSEAPY, with cancer hallmark pathways from serving as the gene sets for enrichment analysis. The computed normalized enrichment scores (NES) for each pathway, per sample, were subsequently averaged across each type of cancer, allowing us to explore patterns and variations in a broader biological context.
[0172] A comprehensive analysis was conducted to evaluate the model's proficiency in predicting tumor types. The class with the highest probability score was identified as the model's prediction. Using the Scikit-leam library, the model was subjected to an extensive suite of metrics to ensure an exhaustive examination of its performance. These metrics comprised the area under the precision-recall curve (AUPRC) and accuracy, all computed using micro-averaging. Additionally, these metrics were individually calculated for each distinct tumor type. To further evaluate the model's accuracy, top-k accuracy measures were employed, with k set to 1 , 2. This metricassessed the frequency at which the true label emerged within the model's highest confidence predictions.Example 6: Clinically applicable deep learning model for robust classification of central nervous system tumors from histopathology images
[0173] REengineered / REcursive DEep lEarning from histoPathoLOgy and methYlation (REDEPLOY) was developed. First, a state-of-the-art WSI-pretrained vision encoder was implemented to generate feature representations optimized for anatomic pathology tasks. Next, these feature representations were used to train three distinct deep learning models. First, a CNS tumor image classifier was trained to predict CNS tumor types from image representations paired with tumor type labels. Second, a DNA methylation model was trained to predict beta values from image representations paired with DNA methylation profiles. Third, a gene expression model was trained to predict mRNA expression levels from image representations paired with transcriptome profiles. Predicted molecular features were used to classify tumors using DNA methylationbased and gene expression-based CNS tumor classifiers. Finally, image classifications and molecular classifications were integrated with CNS tumor classifications based on patient demographic and tumor location information. The output of the integrated model was designed to trigger one or more recursive passes through the CNS tumor classifiers when multiple tumor subtypes could be distinguished within the higher-level tumor type predicted with highest confidence. The model was tested on a large, multi- institutional external dataset and demonstrated that REDEPLOY is a clinically applicable deep learning model for robust classification of CNS tumors from WSI. Methods
[0174] A multilayered, recursive machine learning framework (REDEPLOY) was built and trained to predict nine broad tumor types and 55 specific tumor subtypes from digital whole slide images (WSI) of CNS tumor histopathology slides by integrating predictions of molecular features along with classifications based on image representations and patient demographics. The model was trained on a diverse cohort of 5,592 samples with paired WSI and DNA methylation-based tumor classification data and tested it on a large external cohort of 3,000 samples.Findings
[0175] The model identified nine tumor types and 55 subtypes. The model outperformed human neuropathologists given an equivalent classification task. Model performance was proportional to the training set available. Errors occurred within a narrow spectrum of histologically similar CNS tumor entities. The model can provide human-interpretable summaries of WSIs to highlight diagnostically relevant areas for human review.
[0176] REDEPLOY provides the basis for a clinically applicable deep learning assistant to improve human efficiency and diagnostic accuracy of CNS tumors. The model can immediately be implemented to assist human pathologists in clinical workflows and will become more useful as additional training data is accrued for less frequently encountered tumor types.Example 7: MethodsData Collection
[0177] The data used in this study were whole slide images (WSIs) of hematoxylin and eosin (H&E)-stained sections of formalin fixed paraffin embedded (FFPE) tissues obtained from patients with CNS tumors, along with molecular assays performed on nucleic acids extracted from the same tissues.
[0178] Data used to train the molecular prediction and tumor classification models were collected from three cohorts (i.e. , NCI, CBTN, DBTA), along with additional samples from the NCI archives. Tumor samples from the NCI training cohort (‘NCI- train’) and CBTN cohort had concomitant WSI and methylation profiling data to permit development of a model to predict methylation beta values from WSIs. Tumor samples from select NCI-train cases had concomitant WSI and RNA-sequencing data to permit development of a model to predict transcripts per million reads from WSIs.
[0179] Data used to test the tumor classification models were collected from three cohorts independent of those used for training. The first test cohort was collected from samples accrued through the clinical consultation practice at the NCI after the data freeze for the NCI-train cohort. A second test cohort was collected from samples accrued through the clinical consultation practice at the University College London(UCL), part of the BRAIN UK collaborative virtual brain tumor archive. A third test cohort was collected from select samples accrued through the clinical consultation practice at the University of Pittsburgh Medical Center (UPMC). Samples were selected to enrich for uncommon tumor types expected to present a diagnostic challenge for a human neuropathologist or an algorithm.
[0180] Sample selection was based on availability of WSIs of relevant CNS tumor types and did not consider patient race, sex, or gender. Patient sex was self-reported or inferred from DNA methylation microarrays when available.Model Architecture
[0181] After preparing WSIs for processing, including dividing each one into tiles, the REDEPLOY framework starts with a recently described vision transformer, UNI, pretrained on the Mass-100K histopathology dataset. UNI extracted tile-level feature representations that were subsequently used as input vectors for deep neural networks (DNN) designed to make predictions of tumor types and molecular features. On the first pass through the algorithm, feature representations were passed to three independent DNNs. A tumor type DNN (‘direct model’) learned to predict CNS tumor types and subtypes directly from WSI feature representations by comparing feature representations to matched DNA methylation-based tumor classifications (assigned at the slide level by v12 of the DKFZ CNS tumor classifier). A DNA methylation DNN (‘methylation-indirect model’) learns to predict DNA methylation levels (i.e. , beta values) at 65,000 CpG sites across the genome by comparing WSI feature representations to matched bulk DNA methylation profiles. An mRNA expression DNN (‘expression- indirect model’) learns to predict relative abundance of mRNA expression (i.e., transcripts per million reads, TPM) across 18,000 genes by comparing WSI feature representations to matched bulk RNA sequencing profiles. Tile-level predictions of tumor types and molecular features are aggregated to produce slide-level predictions by calculating the mean of confidence scores, beta values, or TPM across all tiles for each tumor type, CpG site, or gene, respectively.
[0182] Molecular feature predictions were passed to two ensemble models comprising four independent classical machine learning algorithms - random forest (RF), support vector machine (SVM), K nearest neighbors (KNN), logistic regression(LR) - that learn to classify CNS tumors by DNA methylation (‘methylation-indirect model’) or gene expression (‘expression-indirect model’) levels, respectively. A third ensemble model (‘demographic model’) comprising the same four classical algorithms used for molecular feature-based classification learned to classify CNS tumors by patient age and sex and tumor location. Molecular and demographic featured-based classifications from the four submodels were aggregated by calculating the mean of confidence scores across all submodels within each of the three ensemble classifiers for each tumor type so that there is a single set of tumor type predictions with associated confidence scores for each ensemble model.
[0183] In the final step of the first pass through the REDEPLOY algorithm, a fourth ensemble model (‘integration model’) comprising the same four algorithms used in the previously described ensemble models (i.e. , RF, SVM, KNN, LR) learned to predict CNS tumor types from the results of the four previously described tumor type classifiers (i.e., direct model, methylation-indirect model, expression-indirect model, demographic model). The integration model generates a single vector of tumor types with associated confidence scores. The tumor type with the highest confidence score is considered the integrated prediction. If the confidence score of the integrated prediction reaches a tumor-type specific threshold (set to maximize the Youden index for a particular tumor type), and that tumor type has subtypes (true of all tumor types besides hemangioblastoma and central neurocytoma), the REDEPLOY framework executes one or more recursive steps.
[0184] On subsequent passes through the algorithm, tile-level WSI feature representations, slide-level molecular feature predictions, and sample-level demographic features are re-classified by their respective tumor type classifiers, this time trained to predict tumor subtypes within the tumor type predicted with highest confidence on the prior pass through the algorithm. The results of these four classifiers (i.e., direct model, methylation-indirect model, expression-indirect model, demographic model) are passed to the integration model, which produces tumor subtype classifications with associated confidence scores. The subtype with the highest score is considered the integrated prediction. The REDEPLOY framework executes additionalrecursive steps until a tumor subtype without additional subtypes is reached or the confidence score for the integrated prediction does not reach the relevant threshold.Model Evaluation
[0185] Output from the REDEPLOY algorithm was evaluated at the family and at the class level. The integrated prediction associated with the highest confidence score was considered the model output. Simple accuracy, balanced accuracy, F1 score, micro-averaged area under the precision-recall curve (AUPRC), precision, and recall were calculated using the sci-kit learn library. Top- accuracy was also measured with k set to 1 , 2, and 3. This metric calculates the frequency with which the correct classification result was included in the algorithm’s k predictions with highest confidence score. Confidence intervals and p-values were calculated using a bootstrap procedure with 1 ,000 resamples.Algorithm versus neuropathologist experiment
[0186] To compare the reliability and validity of REDEPLOY to human neuropathologists, REDEPLOY CNS tumor classifications for a representative set (selected by KDA) of 96 tumor samples from the test cohorts were compared with ground truth labels (v12 DKFZ classes) and the diagnoses of four human neuropathologists (CGL, CHD, MPN, PJC). Both REDEPLOY and human neuropathologists assigned each sample a family-level and class-level diagnosis from lists of nine possible tumor families and 53 possible tumor classes. Reliability was measured with Cohen’s kappa statistics calculated from contingency tables constructed for each pair of raters (FIG. 13A-13B). Validity was measured by calculating overall accuracy (i.e., percentage correct relative to ground truth) and sum correct (i.e. , total number correct relative to ground truth) (FIGs. 15A-16C).Whole slide image acquisition and preprocessing
[0187] Slides digitized at NCI were scanned on Hamamatsu S60, S210, or S360 digital slide scanners using a 20x objective lens with a numerical aperture of 0-75. Slides were scanned using the 20x or 40x scanning modes resulting in a scanningresolution per pixel of 0 46 x 10'6meter (20x mode) or 0 23 x 10'6meter (40x) mode. WSIs were deidentified prior to analysis using ‘anonymize-slide.py’.
[0188] WSIs were divided into non-overlapping square tiles of 512 x 512 pixels with a pixel size of 0 5 x 10'6meter per pixel (corresponding to 20x scanning magnification) for all images. Tiles in which at least half of pixels represented tissue (as opposed to background) were included in training and testing; tiles representing mostly background were excluded from analysis. Color values were normalized across slides using OpenSlide 1.1.2 with OpenCV 4.5.4 and Pillow (Python Imaging Library).Feature extraction and compression
[0189] To generate WSI feature representations, color-normalized tiles were passed to the UNI vision encoder, which produced a 1 ,024-dimensional feature embedding for each tile. WSIs were encoded as n_tiles x 1 ,024 features matrices.Image feature-based tumor classification (direct model)
[0190] To classify tumors based on WSI features, tile embeddings for each WSI were passed through a multi-layer perceptron (MLP) designed to learn relationships between the compressed WSI feature representations and CNS tumor classifications. CNS tumor classifications were determined by DNA methylation-based tumor classification performed by the DKFZ CNS tumor classifier (v12b6) on DNA methylation levels measured using the Illumina EPIC array. The MLP included three layers: a 1 ,024 node input layer, where each node corresponded to a compressed image feature; a 1 ,024 node hidden layer; and an n-types node output layer, where each node corresponded to a tumor classification and varied depending on which tumor type module was being trained or predicted.DNA methylation prediction (methylation-indirect model)
[0191] To predict DNA methylation levels based on WSI features, tile embeddings for each WSI were passed through an (MLP) designed to learn relationships between the compressed image feature representations and a subset of CpG site beta values selected from the genome-wide DNA methylation profile of the sample represented in the WSI. A set of 65,591 CpG sites was selected for MLPtraining to maximize coverage and beta value variance across samples in the training cohort. Site selection was performed as described below. The MLP included three layers: a 1 ,024-node input layer, where each node corresponded to a compressed image feature; a 1 ,024-node hidden layer; and a 2,000-node output layer, where each node corresponded to a CpG site. Relationships between image features and CpG sites were learned in sets of 2,000 CpG sites until all 65,591 selected sites were covered. The model was trained on 2,595 samples from 47 methylation-based CNS tumor classes.DNA methylation preprocessing and CpG site selection
[0192] Raw intensity values from Illumina Human Methylation EPIC 850K DNA methylation arrays were read into data objects using the minfi R library. Probe intensity values were normalized using the preprocesslllumina function. Probes were filtered to exclude those with ambiguous mappings to the reference genome, those representing non-CpG sites, those representing single nucleotide polymorphisms, those located on the sex chromosomes, and those shown to produce misleading results. Adjustment for the possible effect of FFPE or frozen tissue was performed as implemented in the mnp.v12b6 R package. Beta values were extracted using the getBeta function.
[0193] Of the 865,859 remaining probes, 51 ,052 were excluded due to missing measurements. Probes were then selected for high variance by filtering out those with a standard deviation across samples of the training cohort of less than 0.2. Of the remaining 130,285 probes, 65,591 were selected based on their relatively balanced distribution of beta values across training samples. Probes were considered balanced if the proportions of samples with beta values falling below 0.5 and above 0.5 for each probe were both beneath an imbalance threshold of 0.7. mRNA expression prediction (expression-indirect model)
[0194] To predict gene expression levels based on WSI features, tile embeddings for each WSI were passed through an MLP designed to learn relationships between the compressed image feature representations and a subset of gene transcripts per million reads (TPM) values selected from the transcriptom ic profile of the sample represented in the WSI. A set of 20,000 genes was selected for MLP training to maximize coverageand TPM variance across samples in the training cohort. Gene selection was performed as described below. The MLP included three layers: a 1 ,024-node input layer, where each node corresponded to a compressed image feature; a 1 ,024-node hidden layer; and a 2,000-node output layer, where each node corresponded to a gene. Relationships between image features and genes were learned in sets of 2,000 genes until all 20,000 selected genes were covered.Molecular feature-based tumor classification (indirect models)
[0195] To classify CNS tumors based on predicted DNA methylation and gene expression levels, we developed ensemble CNS tumor classifiers for each modality by integrating the results of four distinct traditional machine learning algorithms. The algorithms included -nearest neighbor (KNN), logistic regression (LR), random forest (RF), and support vector machine (SVM). Inputs to the algorithms were a subset of the predicted molecular features, scaled to a range of 0 to 1 using the MinMaxScaler function from scikit-learn. Features were selected by identifying the 1 ,000 CpG sites or genes with highest ANOVA F-values relative to tumor class using the f_classif and SelectKBest functions from scikit-learn. Outputs from the algorithms, which are confidence scores associated with tumor classifications, were integrated by taking the mean of the confidence scores within each tumor classification across the four algorithms.Patient demographics and tumor location classification (demographic model)
[0196] To classify CNS tumors based on patient age, sex, and tumor location, we developed an ensemble CNS tumor classifier by integrating the results of the four traditional machine learning algorithms used for molecular feature-based classification. Inputs to the algorithms were patient age, sex, and tumor location. Patient ages were scaled to a range of 0 to 1 using the MinMaxScaler function from scikit-learn. Tumor locations were standardized to include the following seven sites: cerebral hemisphere, dural based, lumbar spinal cord, pineal region, posterior fossa, spinal cord (non- lumbar), and ventricle. Outputs from the algorithms, which are confidence scores associated with tumor classifications, were integrated by taking the mean of the confidence scores within each tumor classification across the four algorithms.Reference sets
[0197] For each CNS tumor classifier (direct model, indirect models, demographic model), a reference set was selected to train the classifier. Unsupervised analyses were performed on relevant features to identify coherent groups using a tiered approach. First, the training cohort was divided into tumor families, which divided the CNS tumors into nine categories. Dimensionality reduction, using Uniform Manifold Approximation and Projection (UMAP) using the 4 data layers (methylation-indirect, expression-indirect, direct and demographic) as input demonstrated nine distinct groups of samples in two-dimensional UMAP space. A separate reference set was used to train each of the four preliminary classifiers at the family level. Following construction of the family-level reference sets, reference sets within the nine families were established. Some tumor families included only a single WHO tumor type (e.g., central neurocytoma tumor family included only central neurocytoma tumor class). Most tumor families contained multiple WHO tumor types, and for these, family-specific reference sets were constructed.Model training
[0198] To train and evaluate molecular prediction and tumor classification MLP models, a 5x5 nested cross-validation methodology was employed. The training cohort for each MLP was initially partitioned into five equal outer folds. Partition assignments were made within each CNS tumor type separately to obtain a balanced distribution of CNS tumor types across folds. One of the outer folds (i.e. , 20% of the total training cohort) was excluded from training in each of five successive outer training iterations to help reduce the risk of overfitting. In each outer iteration, the four remaining outer folds (i.e., 80% of the total training data) were themselves divided into five inner folds with a balanced distribution of tumor types across folds. A different inner fold was excluded from training and instead used as a validation set in each of five successive nested inner training iterations. For each outer training iteration, five models were trained, and after five outer iterations, 25 models were generated. Final molecular predictions and tumor classifications were generated by calculating the mean of the predictions of the 25 models.
[0199] Parameter tuning was performed using stochastic gradient descent coupled with mini-batches. The Adam optimizer was used for optimization with an initial learning rate of 0.0001 . The loss function for the direct model was cross-entropy loss. The loss function for the indirect models was mean square error (L2 loss). All MLP models were trained for 500 epochs with early stopping to help reduce the risk of overfitting. Early stopping was implemented by first calculating prediction accuracy (direct model) or mean correlation between predicted and actual molecular feature values (indirect models) in the validation set (i.e., held-out inner fold). If no improvement in prediction accuracy was observed after 30 epochs (direct model) or in mean correlation after 50 epochs (indirect models), training was halted. A dropout rate of 20% was applied to the first layer of each MLP to help reduce the risk of overfitting.
[0200] To train ensemble CNS tumor classifiers based on traditional machine learning algorithms for molecular-based and patient demographic data, a 5-fold cross- validation methodology was employed for each of the four submodels (i.e., RF, SVM, KNN, LR) using the same outer folds as those used for MLP model training. One of the outer folds (i.e., 20% of the total training cohort) was excluded from training in each of five successive outer training iterations and used instead as a validation set; no inner folds (i.e., no nesting) were used. For each outer training iteration, four submodels were trained (i.e., one for each submodel type), and after five iterations, 20 submodels were generated per classifier. Tumor classifications within each of the four submodel types were generated by calculating the mean of the five submodel-specific results. Final tumor classifications for each classifier were generated by calculating the mean of the confidence scores across the four submodels.
[0201] Tumor classifications from the 20 submodels in each ensemble classifier were used as inputs into the integration model. For example, RF has 20 submodels. RF prediction is the average across the 20 submodels and the ensemble output is the output of 80 submodels. Parameters for each of the submodels were set by optimizing the micro-averaged F1 score in the validation set (i.e., held-out outer fold) using the GridSearchCV function of scikit-learn.Example 8: Results
[0202] The REDEPLOY framework started with a general-purpose vision transformer, UNI, pretrained on the Mass-100K dataset, one of the largest and most diverse histopathology image datasets available.
[0203] After preparing H&E images for processing, including dividing each slide image into tiles, UNI extracts tile-level feature representations that are subsequently used as input vectors for deep neural networks (DNN) designed to make tile-level predictions of molecular features and tumor types. On a sample’s first pass through the REDEPLOY classification algorithm, tile-level feature representations are passed to three independent DNNs. A tumor type DNN (“direct model”) learns to predict CNS tumor types and subtypes directly from histopathology feature representations by comparing tile-level image representations to DNA methylation-based tumor classifications assigned at the slide level.
[0204] The recursive design of the REDEPLOY algorithm simplifies the complex task of accurately predicting a single, granular, class-level tumor type from the 55 possible tumor classes by breaking the task into multiple simpler tasks. Predictions are initially made at the family level, among nine possible tumor families. Subsequent predictions are made among the limited classes within a given tumor family. With fewer tumor types from which to choose, the likelihood of errors is reduced, and the REDEPLOY CNS tumor classification algorithm can accurately predict tumor types from among 55 possible classes.
[0205] It was reasoned that if each tier limited the number of categories in the classification task, the classifier would be more accurate. In addition, this allowed to second- and third tiers of the classifiers to be individualized and tunes to the task and samples within the tier. In addition, it allows assessment of individual confidence levels at the family, class and subclass levels, adding stringency and definition to the classifier predictions.
[0206] A DNA methylation predictor from the H&E slides was included, where a deep learning model (FIG. 11 ) was trained using matched slide images and DNA methylation profiles from the training cohort. The inferred methylation was then usedtrain and predict tumor types using the inferred DNA methylation data with standard machine learning (“methylation-indirect model”).
[0207] A predictor of gene expression was additionally trained from the histology image using a method (DeepPT). The inferred gene expression profile was then used to train a classifier model (“expression-indirect model”).
[0208] A third, “direct model” was used as a deep learning classifier for training and prediction of tumor types based on histopathology images.
[0209] The fourth, was the “demographic model”, which includes patient age, sex, and tumor location, to classify tumor types. Finally, the 4 models were then combined into a fifth (Integrated) model which was a composite of the 4 data layers described above.
[0210] The new model includes gene expression predictions and recursive tumor subtype modules that permit accurate prediction of 55 clinically relevant CNS tumor types and subtypes from WSI.
[0211] To further characterize the predictions of DNA methylation and gene expression across the diverse set of CNS tumors, the performance within individual CNS tumor categories was evaluated. Using the CNS tumor categories with the largest sample sizes, the correlation of predicted versus actual values in these individual tumor categories was assessed (FIG. 12, FIGs. 14A-14B).
[0212] Having described several embodiments, it will be recognized by those skilled in the art that various modifications, alternative constructions, and equivalents may be used without departing from the spirit of the disclosure. Additionally, a number of well-known processes and elements have not been described in order to avoid unnecessarily obscuring the present disclosure. Accordingly, the above description should not be taken as limiting the scope of the disclosure.
[0213] Those skilled in the art will appreciate that the presently disclosed embodiments teach by way of example and not by limitation. Therefore, the matter contained in the above description or shown in the accompanying drawings should be interpreted as illustrative and not in a limiting sense. The following claims are intended to cover all generic and specific features described herein, as well as all statements ofthe scope of the present method and system, which, as a matter of language, might be said to fall therebetween.Exemplary Embodiments
[0214] The following is a list of non-limiting exemplary embodiments and may include combinations thereof.
[0215] Embodiment 1 : A system for classifying a tumor, the system comprising: a processor in communication with a memory, the memory including instructions executable by the processor to: receive a histopathology image of a sample of the tumor; generate a first classification of the tumor by: providing the histopathology image to a trained methylation model to predict a methylation profile for the tumor; providing the predicted methylation profile to a trained methylation classifier; and generating a first set of prediction scores based on the trained methylation classifier; generate a second classification of the tumor by: providing the histopathology image to a trained direct classification model; and generating a second set of prediction scores based on the trained direct classification model; generate a third classification of the tumor by: providing demographic data to a trained demographic model; and generating a third set of prediction scores based on the trained demographic model; and generate an integrated classification by averaging the first set of prediction scores, the second set of prediction scores, and the third set of prediction scores and selecting the classification with the highest averaged prediction score.
[0216] Embodiment 2: The system of embodiment 1 , the memory further including instructions executable by the processor to: generate a fourth classification of the tumor by: providing the histopathology image to a trained gene expression model to predict a gene expression profile for the tumor; providing the predicted gene expression profile to a trained gene expression classifier; and generating a fourth set of prediction scores based on the trained gene expression classifier; and generate the integrated classification by averaging the first, second, third, and fourth sets of prediction scores, wherein the integrated classification is selected from a set of tumor families, types, or sub-types.
[0217] Embodiment 3: The system of embodiment 1 , the memory further including instructions executable by the processor to repeat one or more of the generating steps.
[0218] Embodiment 4: The system of embodiment 1 , wherein the histopathology image is preprocessed by dividing the histopathology image into a plurality of nonoverlapping tiles, omitting any non-overlapping tiles which include more than half background, extracting features from the non-overlapping tiles, and compressing the histopathology image.
[0219] Embodiment 5: The system of embodiment 4, wherein the first set of prediction scores and the second set of prediction scores comprise a set of prediction scores for each of the non-overlapping tiles in the histopathology image, thereby facilitating a spatial classification of the histopathology image using the methylation model, the methylation classifier, the direct classification model, or a combination thereof.
[0220] Embodiment 6: The system of embodiment 1 , the memory further including instructions executable by the processor to: update the trained methylation model and the trained methylation classifier with the methylation profile and the integrated classification; update the trained direct classification model with the histopathology image and the integrated classification; and update the trained demographic model with the demographic data and the integrated classification.
[0221] Embodiment 7: The system of embodiment 1 , wherein the trained methylation model comprises a deep learning model trained using labeled data comprising matched slide images and DNA methylation profiles.
[0222] Embodiment 8: The system of embodiment 7, wherein the trained methylation classifier comprises at least two machine learning models trained using labeled data comprising DNA methylation data and tumor types.
[0223] Embodiment 9: The system of embodiment 8, wherein the at least two machine learning models are selected from the group consisting of logistic regression, support vector machine, k-nearest neighbors, random forest, and combinations thereof.
[0224] Embodiment 10: The system of embodiment 1 , wherein the trained direct classification model comprises a deep learning model trained using labeled data comprising histopathology images and tumor types.
[0225] Embodiment 11 : The system of embodiment 1 , wherein the trained demographic model comprises a machine learning model trained using labeled data comprising demographic data and tumor types.
[0226] Embodiment 12: The system of embodiment 1 , wherein the demographic data comprises age, sex, and location of the tumor.
[0227] Embodiment 13: The system of embodiment 1 , wherein the integrated classification comprises at least one tumor sub-type.
[0228] Embodiment 14: The system of embodiment 1 , wherein the integrated classification comprises more than one tumor sub-type.
[0229] Embodiment 15: The system of embodiment 13, wherein the tumor is a central nervous system tumor.
[0230] Embodiment 16: The system of embodiment 15, wherein the classification comprises a tumor sub-type selected from the group consisting of glioblastoma, medulloblastoma, ependymoma, pilocytic astrocytoma, meningioma, astrocytoma IDH- mutant, choroid plexus, subependymoma, myxopapillary ependymoma, oligodendroglioma, and combinations thereof.
[0231] Embodiment 17: The system of embodiment 1 , wherein the predicted methylation profile comprises DNA methylation beta values.
[0232] Embodiment 18: The system of embodiment 1 , the memory further including instructions executable by the processor to: train the methylation model, the methylation classifier, the direct model, and / or the demographic model prior to providing the histopathology image.
[0233] Embodiment 19: The system of embodiment 1 , the memory further including instructions executable by the processor to: diagnose the tumor using the generated integrated classification.
[0234] Embodiment 20: The system of embodiment 19, the memory further including instructions executable by the processor to: form a treatment plan specific to the diagnosis of the tumor.
[0235] Embodiment 21 : The system of embodiment 1 , wherein the histopathology image is an H&E image.
[0236] Embodiment 22: A system for classifying a tumor, the system comprising: a processor in communication with a memory, the memory including instructions executable by the processor to: receive a histopathology image of a sample of the tumor; and generate a first classification of the tumor by: providing the histopathology image to a trained methylation model to predict a methylation profile for the tumor; providing the predicted methylation profile to a trained methylation classifier; and generating a first set of prediction scores based on the trained methylation classifier, wherein the first classification is selected from a set of tumor families, types, or subtypes.
[0237] Embodiment 23: The system of embodiment 22, the memory further including instructions executable by the processor to: generate a second classification of the tumor by: providing the histopathology image to a trained direct classification model; and generating a second set of prediction scores based on the trained direct classification model, wherein the second classification is selected from the set of tumor families, types, or sub-types.
[0238] Embodiment 24: The system of embodiment 23, the memory further including instructions executable by the processor to: generate a third classification of the tumor by: providing demographic data to a trained demographic model; and generating a third set of prediction scores based on the trained demographic model, wherein the third classification is selected from the set of tumor families, types, or subtypes.
[0239] Embodiment 25: The system of embodiment 24, the memory further including instructions executable by the processor to: generate a fourth classification of the tumor by: providing the histopathology image to a trained gene expression model to predict a gene expression profile for the tumor; providing the predicted gene expression profile to a trained gene expression classifier; and generating a second set of prediction scores based on the trained gene expression classifier, wherein the third classification is selected from the set of tumor families, types, or sub-types.
[0240] Embodiment 26: The system of embodiment 25, the memory further including instructions executable by the processor to: generate an integrated classification by averaging the first set of prediction scores, the second set of prediction scores, the third set of prediction scores, and the fourth set of prediction scores and selecting the classification with the highest averaged prediction score, wherein the integrated classification is selected from a set of tumor families, types, or sub-types.
[0241] Embodiment 27: The system of embodiment 26, the memory further including instructions executable by the processor to repeat one or more of the generating steps.
Claims
CLAIMSWhat is claimed is:1 . A system for classifying a tumor, the system comprising: a processor in communication with a memory, the memory including instructions executable by the processor to: receive a histopathology image of a sample of the tumor; generate a first classification of the tumor by: providing the histopathology image to a trained methylation model to predict a methylation profile for the tumor; providing the predicted methylation profile to a trained methylation classifier; and generating a first set of prediction scores based on the trained methylation classifier; generate a second classification of the tumor by: providing the histopathology image to a trained direct classification model; and generating a second set of prediction scores based on the trained direct classification model; generate a third classification of the tumor by: providing demographic data to a trained demographic model; and generating a third set of prediction scores based on the trained demographic model; and generate an integrated classification by averaging the first set of prediction scores, the second set of prediction scores, and the third set of prediction scores and selecting the classification with the highest averaged prediction score, wherein the integrated classification is selected from a set of tumor families, types, or sub-types.
2. The system of claim 1 , the memory further including instructions executable by the processor to: generate a fourth classification of the tumor by: providing the histopathology image to a trained gene expression model to predict a gene expression profile for the tumor;providing the predicted gene expression profile to a trained gene expression classifier; and generating a fourth set of prediction scores based on the trained gene expression classifier; and generate the integrated classification by averaging the first, second, third, and fourth sets of prediction scores, wherein the integrated classification is selected from a set of tumor families, types, or sub-types.
3. The system of claim 1 , the memory further including instructions executable by the processor to repeat one or more of the generating steps.
4. The system of claim 1 , wherein the histopathology image is preprocessed by dividing the histopathology image into a plurality of non-overlapping tiles, omitting any non-overlapping tiles which include more than half background, extracting features from the non-overlapping tiles, and compressing the histopathology image.
5. The system of claim 4, wherein the first set of prediction scores and the second set of prediction scores comprise a set of prediction scores for each of the non-overlapping tiles in the histopathology image, thereby facilitating a spatial classification of the histopathology image using the methylation model, the methylation classifier, the direct classification model, or a combination thereof.
6. The system of claim 1 , the memory further including instructions executable by the processor to: update the trained methylation model and the trained methylation classifier with the methylation profile and the integrated classification; update the trained direct classification model with the histopathology image and the integrated classification; and update the trained demographic model with the demographic data and the integrated classification.
7. The system of claim 1 , wherein the trained methylation model comprises a deep learning model trained using labeled data comprising matched slide images and DNA methylation profiles.
8. The system of claim 7, wherein the trained methylation classifier comprises at least two machine learning models trained using labeled data comprising DNA methylation data and tumor types.
9. The system of claim 8, wherein the at least two machine learning models are selected from the group consisting of logistic regression, support vector machine, k-nearest neighbors, random forest, and combinations thereof.
10. The system of claim 1 , wherein the trained direct classification model comprises a deep learning model trained using labeled data comprising histopathology images and tumor types.11 . The system of claim 1 , wherein the trained demographic model comprises a machine learning model trained using labeled data comprising demographic data and tumor types.
12. The system of claim 1 , wherein the demographic data comprises age, sex, and location of the tumor.
13. The system of claim 1 , wherein the integrated classification comprises at least one tumor sub-type.
14. The system of claim 13, wherein the integrated classification comprises more than one tumor sub-type.
15. The system of claim 13, wherein the tumor is a central nervous system tumor.
16. The system of claim 15, wherein the classification comprises a tumor sub-type selected from the group consisting of glioblastoma, medulloblastoma, ependymoma, pilocytic astrocytoma, meningioma, astrocytoma IDH-mutant, choroidplexus, subependymoma, myxopapillary ependymoma, oligodendroglioma, and combinations thereof.
17. The system of claim 1 , wherein the predicted methylation profile comprises DNA methylation beta values.
18. The system of claim 1 , the memory further including instructions executable by the processor to: train the methylation model, the methylation classifier, the direct model, and / or the demographic model prior to providing the histopathology image.
19. The system of claim 1 , the memory further including instructions executable by the processor to: diagnose the tumor using the generated integrated classification.
20. The system of claim 17, the memory further including instructions executable by the processor to: form a treatment plan specific to the diagnosis of the tumor.21 . The system of claim 1 , wherein the histopathology image is an H&E image.
22. A system for classifying a tumor, the system comprising: a processor in communication with a memory, the memory including instructions executable by the processor to: receive a histopathology image of a sample of the tumor; and generate a first classification of the tumor by: providing the histopathology image to a trained methylation model to predict a methylation profile for the tumor; providing the predicted methylation profile to a trained methylation classifier; and generating a first set of prediction scores based on the trained methylation classifier,wherein the first classification is selected from a set of tumor families, types, or sub-types.
23. The system of claim 22, the memory further including instructions executable by the processor to: generate a second classification of the tumor by: providing the histopathology image to a trained direct classification model; and generating a second set of prediction scores based on the trained direct classification model, wherein the second classification is selected from the set of tumor families, types, or sub-types.
24. The system of claim 23, the memory further including instructions executable by the processor to: generate a third classification of the tumor by: providing demographic data to a trained demographic model; and generating a third set of prediction scores based on the trained demographic model, wherein the third classification is selected from the set of tumor families, types, or sub-types.
25. The system of claim 24, the memory further including instructions executable by the processor to: generate a fourth classification of the tumor by: providing the histopathology image to a trained gene expression model to predict a gene expression profile for the tumor; providing the predicted gene expression profile to a trained gene expression classifier; and generating a second set of prediction scores based on the trained gene expression classifier, wherein the third classification is selected from the set of tumor families, types, or sub-types.
26. The system of claim 25, the memory further including instructions executable by the processor to: generate an integrated classification by averaging thefirst set of prediction scores, the second set of prediction scores, the third set of prediction scores, and the fourth set of prediction scores and selecting the classification with the highest averaged prediction score, wherein the integrated classification is selected from a set of tumor families, types, or sub-types.
27. The system of claim 26, the memory further including instructions executable by the processor to repeat one or more of the generating steps.
Citation Information
Patent Citations
Determining biomarkers from histopathology slide images
CA3133826A1
Systems and methods for deep orthogonal fusion for multimodal prognostic biomarker discovery
EP4239647A1