System and method for determining IDH status in gliomas

A deep learning-based method combining EsViT and DeepMIL classifiers analyzes H&E images to accurately predict IDH mutation status in gliomas, addressing the limitations of traditional methods and improving diagnostic efficiency.

WO2025094178A1PCT designated stage expired Publication Date: 2025-05-08SHEBA IMPACT LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/IL2024/051044
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-30
Filing Date
2024-10-30
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Traditional methods for determining IDH mutation status in gliomas, such as immunohistochemistry and molecular testing, are limited by sensitivity, can produce inconclusive results, and require complex laboratory procedures and specialized equipment, leading to delays in diagnosis and treatment.

Method used

A deep learning-based method using a combination of a self-supervised Vision Transformer (EsViT) and an attention-based Deep Multiple Instance Learning (DeepMIL) classifier to analyze H&E histopathological images, extracting region-specific feature embeddings, and integrating patient-specific data to predict IDH mutation status.

Benefits of technology

This approach enables more accurate and efficient identification of IDH mutation status, reducing reliance on complex laboratory procedures and specialized equipment, and improving accessibility in various healthcare settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IL2024051044_08052025_PF_FP_ABST
    Figure IL2024051044_08052025_PF_FP_ABST
Patent Text Reader

Abstract

A system and method for determining IDH status in gliomas includes activating a trained classifier on region-specific feature embeddings of an input stained slide of a patient to produce a set of region-specific probabilities, and combining patient-specific data with the set of region-specific probabilities to produce an isocitrate dehydrogenase (IDH) mutation status prediction for the input stained slide. The method may include dividing the input slide into a multiplicity of tiles and extracting the region-specific feature embeddings from the tiles. The trained classifier may be an attention-based deep multiple instance learning (DeepMIL) classifier. The combining may include performing a logistic regression on the IDH mutation status prediction and the patient-specific data.
Need to check novelty before this filing date? Find Prior Art

Description

TITLE OF THE INVENTIONSYSTEM AND METHOD FOR DETERMINING IDH STATUS IN GLIOMASCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority from US provisional patent application 63 / 594,111, filed October 30, 2023, which is incorporated herein by reference.FIELD OF THE INVENTION

[0002] The present invention relates generally to methods for determining isocitrate dehydrogenase (IDH) mutation status in gliomas and to deep learning-based approaches for predicting IDH status from histopathological images in particular.BACKGROUND OF THE INVENTION

[0003] Gliomas are a group of common primary malignant brain tumors originating from glial cells or stem cells that develop glial characteristics. They are characterized by high morbidity, high recurrence, high mortality and low cure rate which is largely dependent on its subtype. Determining the molecular features of gliomas is crucial for accurate diagnosis, prognosis, and treatment planning.

[0004] Among the molecular markers, the isocitrate dehydrogenase (IDH) mutation status has emerged as a critical factor in glioma classification and patient stratification. IDH mutations occur in the IDH1 and IDH2 genes, leading to abnormal enzymatic activity in cellular metabolism. These mutations are frequently observed in lower-grade gliomas, such as diffuse astrocytomas and oligodendrogliomas, while they are less prevalent in high-grade gliomas, including glioblastomas.

[0005] The IDH mutation status has been recognized as an independent prognostic factor, with IDH-mutated gliomas generally associated with a more favorable clinical outcome compared to theirIDH wild-type counterparts. Accurate determination of the IDH mutation status is therefore useful for effective glioma diagnosis and treatment decision-making.

[0006] Traditional methods for assessing IDH status involve immunohistochemistry (IHC) stains and molecular testing. These techniques are widely used in pathology laboratories to analyze tissue samples and determine the presence of specific molecular markers. For example, hematoxylin and eosin (H&E) staining is a routine staining technique commonly employed to visualize cellular structures in tissue samples. While H&E staining provides valuable information regarding the morphology of tissues, it is insufficient for determining IDH mutation status on its own.

[0007] In recent years, artificial intelligence-based approaches have emerged as promising tools for analyzing medical imaging data, including histopathological images. One such approach is discussed in the online article, by W. Wang, et al., and entitled, "Neuropathologist-level integrated classification of adult-type diffuse gliomas using deep learning from whole-slide pathological images", Nature Communications, October 11, 2023.SUMMARY OF THE PRESENT INVENTION

[0008] There is therefore provided, in accordance with a preferred embodiment of the present invention, a method for determining IDH status in gliomas. The method includes activating a trained classifier on region-specific feature embeddings of an input stained slide of a patient to produce a set of region-specific probabilities, and combining patient-specific data with the set of region-specific probabilities to produce an isocitrate dehydrogenase (IDH) mutation status prediction for the input stained slide.

[0009] Moreover, in accordance with a preferred embodiment of the present invention, the method also includes dividing an H&E (hematoxylin and eosin) slide of a patient into a multiplicity of tiles, and where the activating includes extracting the region-specific feature embeddings from the multiplicity of tiles.

[0010] Further, in accordance with a preferred embodiment of the present invention, the extracting includes activating a vision transformer model on the multiplicity of tiles.

[0011] Still further, in accordance with a preferred embodiment of the present invention, the trained classifier is an attention-based deep multiple instance learning (DeepMIL) classifier.

[0012] Additionally, in accordance with a preferred embodiment of the present invention, the combining includes performing a logistic regression on the IDH mutation status prediction and the patient-specific data to produce the IDH mutation status prediction.

[0013] Moreover, in accordance with a preferred embodiment of the present invention, the patientspecific data includes at least one of age and sex of the patient.

[0014] Further, in accordance with a preferred embodiment of the present invention, the attentionbased DeepMIL classifier includes a self-attention module configured to generate attention scores for the region-specific feature embeddings.

[0015] Still further, in accordance with a preferred embodiment of the present invention, the method further includes training the trained classifier using a dataset of H&E stained slides with known IDH mutation statuses.

[0016] Additionally, in accordance with a preferred embodiment of the present invention, the training includes extracting a plurality of tiles from each H&E stained slide from the dataset, generating tile embeddings for the plurality of tiles using the pre-trained vision transformer model, and training the DeepMIL classifier using the tile embeddings and corresponding IDH mutation statuses.

[0017] Moreover, in accordance with a preferred embodiment of the present invention, the pretrained vision transformer model is an EsViT (self-supervised vision transformer) model.

[0018] Further, in accordance with a preferred embodiment of the present invention, the method further includes pre-training the EsViT model using a self-supervised learning approach on the dataset.

[0019] Still further, in accordance with a preferred embodiment of the present invention, the selfsupervised learning approach includes creating multiple augmented views of each tile, matching corresponding regions across the augmented views, and minimizing a distance between matched regions.

[0020] Additionally, in accordance with a preferred embodiment of the present invention, training the DeepMIL classifier includes assigning attention scores to tiles indicating their contribution to an overall IDH mutation status prediction of their corresponding H&E stained slide, combining the attention scores with the tile embeddings to generate weighted aggregations, and generating the set of region-specific probabilities from the weighted aggregations.

[0021] Moreover, in accordance with a preferred embodiment of the present invention, the method further includes using a weakly supervised loss function to compute a per slide loss value from a predicted IDH mutation status probability and an actual IDH mutation status per slide.

[0022] Further, in accordance with a preferred embodiment of the present invention, the weakly supervised loss function is a Binary Cross Entropy loss.

[0023] Still further, in accordance with a preferred embodiment of the present invention, the method further includes updating weights of the attention-based DeepMIL classifier using an AdamW optimizer.

[0024] Additionally, in accordance with a preferred embodiment of the present invention, the method further includes generating an attention map for the input stained slide, where the attention map indicates regions of high attention corresponding to areas indicative of IDH mutation status.

[0025] Moreover, in accordance with a preferred embodiment of the present invention, the method is implemented as part of a decision support tool for glioma cases that test negative in immunohistochemistry (IHC) staining and includes recommending molecular tests for confirming IDH mutation status when the IDH mutation status prediction indicates a positive result.

[0026] There is therefore provided, in accordance with a preferred embodiment of the present invention, a system for determining IDH status in gliomas. The system is implemented on a computing device having a processor and a memory. The system includes a trained classifier and a clinical data integrator. The trained classifier processes region-specific feature embeddings of an input stained slide of a patient to produce a set of region-specific probabilities. The clinical data integrator combines patient-specific data with the set of region-specific probabilities to produce an IDH mutation status prediction for the input stained slide.

[0027] Moreover, in accordance with a preferred embodiment of the present invention, the system also includes a tile extractor to divide an H&E slide of a patient into a multiplicity of tiles, and where the trained classifier is configured to extract the region-specific feature embeddings from the multiplicity of tiles.

[0028] Further, in accordance with a preferred embodiment of the present invention, the trained classifier includes a pre-trained vision transformer to extract the region-specific feature embeddings from the multiplicity of tiles.

[0029] Still further, in accordance with a preferred embodiment of the present invention, the trained classifier also includes an attention-based deep multiple instance learning classifier.

[0030] Additionally, in accordance with a preferred embodiment of the present invention, the clinical data integrator includes a logistic regression module to perform a logistic regression on the IDH mutation status prediction and the patient-specific data to produce the IDH mutation status prediction.

[0031] Further, in accordance with a preferred embodiment of the present invention, the attentionbased DeepMIF classifier includes a self-attention module to generate attention scores for the regionspecific feature embeddings.

[0032] Still further, in accordance with a preferred embodiment of the present invention, the system further includes a trainer to train the trained classifier using a dataset of H&E stained slides with known IDH mutation statuses.

[0033] Additionally, in accordance with a preferred embodiment of the present invention, the trainer includes the tile extractor to extract a plurality of tiles from each H&E stained slide of the dataset, the pre-trained vision transformer to generate tile embeddings for the plurality of tiles usinga pre-trained vision transformer model, and a loss function to train the DeepMIL classifier using the tile embeddings and corresponding IDH mutation statuses.

[0034] Moreover, in accordance with a preferred embodiment of the present invention, the pretrained vision transformer is an EsViT model.

[0035] Further, in accordance with a preferred embodiment of the present invention, the trainer is further configured to pre-train the EsViT model using a self-supervised learning approach on the dataset.

[0036] Still further, in accordance with a preferred embodiment of the present invention, the selfsupervised learning approach includes creating multiple augmented views of each tile, matching corresponding regions across the augmented views, and minimizing a distance between matched regions.

[0037] Additionally, in accordance with a preferred embodiment of the present invention, the region-specific feature embeddings indicate their contribution to an overall IDH mutation status prediction of their corresponding H&E stained slide.

[0038] Moreover, in accordance with a preferred embodiment of the present invention, the loss function is a weakly supervised loss function to compute a per slide loss value from a predicted IDH mutation status probability and an actual IDH mutation status per slide.

[0039] Further, in accordance with a preferred embodiment of the present invention, the weakly supervised loss function is a Binary Cross Entropy loss.

[0040] Still further, in accordance with a preferred embodiment of the present invention, the trainer is further configured to update weights of the attention-based DeepMIL classifier using anAdamW optimizer.

[0041] Additionally, in accordance with a preferred embodiment of the present invention, the system further includes an attention map generator configured to generate an attention map for the input stained slide, where the attention map indicates regions of high attention corresponding to areas indicative of IDH mutation status.

[0042] Moreover, in accordance with a preferred embodiment of the present invention, the system is implemented as part of a decision support tool for glioma cases that test negative in immunohistochemistry staining and includes a recommendation module configured to recommend molecular tests for confirming IDH mutation status when the IDH mutation status prediction indicates a positive result.BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The subject matter regarded as the invention is particularly pointed out and distinctly claimed in the concluding portion of the specification. The invention, however, both as to organization and method of operation, together with objects, features, and advantages thereof, may best be understood by reference to the following detailed description when read with the accompanying drawings in which:

[0044] Fig. 1 is a block diagram illustration of an IDH mutation identifying system, constructed and operative according to an embodiment of the present invention;

[0045] Figs. 2A and 2B together form a block diagram illustration of components of a trainer for training an IDH mutation identifier, forming part of the system of Fig. 1 ;

[0046] Fig. 3 is a block diagram illustration of the elements of a DeepMIL classifier, forming part of the trainer of Fig. 2B;

[0047] Fig. 4 is a block diagram illustration of the elements of a clinical data integrator, forming part of the system of Fig. 1 ;

[0048] Fig. 5 is a block diagram illustration of a trained IDH mutation identifier, forming part of the system of Fig. 1 ;

[0049] Fig. 6 is a graphical illustration of IDH mutant samples with positive staining regions and high attention areas, showing the results of operating the system of Fig. 1 ; and

[0050] Fig. 7 is a block diagram illustration of the IDH mutation identifier with a decision support tool, constructed and operative according to an embodiment of the present invention.

[0051] It will be appreciated that for simplicity and clarity of illustration, elements shown in the Figures have not necessarily been drawn to scale. For example, the dimensions of some of the elements may be exaggerated relative to other elements for clarity. Further, where consideredappropriate, reference numerals may be repeated among the Figures to indicate corresponding or analogous elements.DETAILED DESCRIPTION OF THE PRESENT INVENTION

[0052] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the invention. However, it will be understood by those skilled in the art that the present invention may be practiced without these specific details. In other instances, well-known methods, procedures, and components have not been described in detail so as not to obscure the present invention.

[0053] Applicant has realized that the prior art systems for determining IDH mutation status in gliomas face significant challenges. Traditional methods such as immunohistochemistry (IHC) stains and molecular testing are often limited by their sensitivity, can produce inconclusive results, and may be time-consuming and labor-intensive. Additionally, these methods often require complex laboratory procedures and specialized equipment, which can be costly and may not be readily available in all healthcare settings. These limitations can lead to delays in diagnosis and treatment decisions, potentially impacting patient outcomes.

[0054] Applicant has realized that deep learning techniques, particularly when applied to histopathological images, offer a promising avenue for addressing these challenges. Such an approach could potentially reduce the reliance on complex laboratory procedures and specialized equipment, making IDH identification more accessible in various healthcare settings, including those with limited resources.

[0055] Applicant has realized that a machine learning system which may capture and analyze specific regions within the tissue may be particularly effective for analyzing glioma samples, since IDH mutation-related features may be localized to specific regions within the tissue. In particular, Applicant has realized that a combination of a self-supervised Vision Transformer (EsViT) and a Deep Multiple Instance Learning (DeepMIL) classifier may be particularly useful for such a system.The EsViT component may capture fine-grained details and patterns in these localized regions, while the DeepMIL component may focus attention on the most relevant areas of the slide. This approach may allow for more accurate identification of IDH mutation status from H&E histopathological images by leveraging the strengths of both models to detect and analyze the specific tissue regions that are most indicative of IDH mutation status.

[0056] Moreover, Applicant has realized that the identification of the IDH mutation status may be improved by adding patient-specific clinical data to the image analysis.

[0057] Reference is now made to Fig. 1 , which is a block diagram of an IDH mutation identifying system 10. IDH mutation identifying system 10 may comprise a tile extractor 12, an IDH mutation identifier 14, an IDH mutation identifier trainer 16, and a clinical data integrator 18.

[0058] IDH mutation identifying system 10 may process an input-stained slide, such as an H&E stained glioma slide, from a patient to generate a final IDH mutation status probability. Tile extractor 12 may first break down the slide image into tiles after which IDH mutation identifier 14 may analyze these tiles. In accordance with a preferred embodiment of the present invention, clinical data integrator 18 may combine this analysis with data about the patient whose slide it is, to produce the final IDH mutation status probability.

[0059] Tile extractor 12 may receive the stained slide as input and may extract tiles of the stained slide, normalizing the stained slide in the process. These extracted tiles may then be passed to IDH mutation identifier 14. As described in more detail hereinbelow, IDH mutation identifier 14 may process the tiles to determine region-specific probabilities related to IDH mutation status. As described in more detail hereinbelow, IDH mutation identifier 14 may be a machine learning system which may be trained by IDH mutation identifier trainer 16. Clinical data integrator 18 may integrateprobabilities from IDH mutation identifier 14 with patient data of the input-stained slide to produce an IDH mutation status probability as output.

[0060] Tile extractor 12 may segment tissue regions using, for example, binary thresholding in LAB color-space. Subsequently, tile extractor 12 may tile the segmented stained slide into patches, such as of size 256 x 256 pixels at 20x magnification with a resolution of 0.5 microns per pixel. To reduce the impact of the staining color, tile extractor 12 may apply a stain normalization and augmentation, such as the RandStainNA normalization, on all of the extracted tiles.

[0061] Reference is now made to Figs. 2A and 2B, which, together, detail IDH mutation identifier trainer 16. Trainer 16 has two portions, a feature representation portion, shown in Fig. 2A, and a probability determining portion, shown in Fig. 2B.

[0062] The feature representation portion may use an EsViT (Efficient Self-supervised Vision Transformer) model, such as that described in the paper, “Li C, Yang J, Zhang P, Gao M, Xiao B, Dai X, Yuan L, Gao J. Efficient self-supervised vision transformers for representation learning. arXiv preprint arXiv:2106.09785. 2021 Jun 17," whose disclosure is incorporated herein by reference.

[0063] Specifically, this part of trainer 16 may comprise tile extractor 12, and an EsViT model in training 20, which may be trained using a self-supervised loss function 22. In this feature representation portion, tile extractor 12 may receive a multiplicity of stained glioma slides as input and may extract a plurality of tiles from each of these slides. For example, there may be up to 5,000 tiles per slide and there may be up to 1600 slides.

[0064] At each iteration of the training, EsViT model 20 may attempt to produce a set of embeddings (i.e. region-specific feature embeddings) per slide. Self-supervised loss function 22 may compare the tile embeddings with the embeddings of corresponding regions in augmented views of the same image, leveraging a region-matching approach for self-supervision. Specifically, thismethod involves creating multiple augmented views (known as “teacher” model views) of each slide tile (known as a “student” model’s view) and matching corresponding regions across these views. For each tile region in the "student" model’s view, EsViT model 20 may find the most similar region in the "teacher" model’s view by using cosine similarity. Self-supervised loss function 22 may then minimize the distance between matched regions across views, encouraging the embeddings of similar regions to align closely. This region-matching task enables EsViT model 20 to learn meaningful representations by capturing fine-grained spatial relationships without requiring explicit labels.

[0065] This self-supervised approach may allow the model to learn representations from the input data without requiring explicit labels. The output of the EsViT trainer 20 may be a trained EsViT model 30 (shown in Fig. 2B), which may serve as a feature extractor, capable of generating representations from glioma slide images.

[0066] In accordance with a preferred embodiment of the present invention, to accelerate convergence, EsViT model 20 may be initialized with published weights, such as those which were pretrained on a variety of different types of images (e.g. images from the Imagenet database), and then trained on tiles from H&E slides until convergence.

[0067] The probability determining portion of IDH mutation identifier trainer 16, shown in Fig. 2B, may comprise tile extractor 12, trained EsViT 30, which is the output of the feature representation portion of Fig. 2A, an attention-based DeepMIL classifier 32, a weakly supervised loss function 34, and clinical data integrator 18.

[0068] Tile extractor 12 may receive mutant H&E slides and wild-type H&E slides as input and may extract a plurality of tiles of the input-stained slides. Trained EsViT 30 may generate tile embeddings for the plurality of tiles of each stained slide. The per slide, tile embeddings may then be input into DeepMIL classifier 32, which may be in a training phase. DeepMIL classifier 32, with itsintegrated self-attention layer, may assign exemplary attention scores to tiles, indicating their contribution to the overall IDH mutation status prediction of their stained slide, and may produce region-specific probabilities from these attention scores. This approach facilitates weakly-supervised learning, allowing trainer 16 to focus on pertinent regions of the stained slides without explicit tilelevel annotations.

[0069] For each iteration of the training phase, the exemplary region-specific probabilities may be then fed into clinical data integrator 18. Clinical data integrator 18 may integrate the region-specific probabilities of a slide with patient data related to that slide to produce an exemplary IDH mutation status probability for each slide. Weakly supervised loss function 34 may compute a per slide loss value from the exemplary IDH mutation status probability and the actual IDH mutation status of the slide, which may then be used to train DeepMIL classifier 32. After multiple training iterations, trainer 16 may output a trained DeepMIL, labeled 62 in Fig. 5.

[0070] For example, for training DeepMIL classifier 32, 80% of the tiles extracted by tile extractor 12 may be used, for a 5-fold cross-validation training. In this embodiment, weakly supervised loss function 34 may be a Binary Cross Entropy loss which may update the weights of DeepMIL classifier 32 using an AdamW optimizer. For example, the weight decay factor may be 0.001 with a learning rate of 0.001 and an exponential scheduling of 0.99 gamma. The weights generated from loss function 34 may be normalized by the proportion of IDH mutant H&E slides in the training set.

[0071] Reference is now made to Fig. 3, which illustrates the elements of DeepMIL classifier 32 in training. DeepMIL classifier 32 may comprise a self-attention module 40, a summer 42, and a fully connected (FC) classifier module 44.

[0072] Self-attention module 40 may be any neural network which generates attention scores based on the input tile embeddings. For example, self-attention module 40 may comprise a two-layerneural network to compute attention weights directly from the tile embeddings. The first layer may apply a hyperbolic tangent (tanh) activation, which may support gradient flow and may enable module 40 to learn complex relationships. The second layer may be a gating mechanism that may combine tanh with a sigmoid function to refine the attention weights.

[0073] Self-attention module 40 may be trained to compute attention weights for each tile embedding of a slide, indicating its importance to the classification task. Self-attention module 40 may normalize the weights for a slide to sum to one. This approach allows the model to effectively aggregate the tile embeddings while enhancing the classifier's interpretability and performance.

[0074] The attention scores may then be passed to summer 42, which may perform a weighted aggregation of the scores with the tile embeddings. FC classifier module 44 may process this aggregated data and may output region-specific probabilities as the final result of the classification process.

[0075] FC classifier 44 may comprise one or more neural network layers, where each layer may perform a linear transformation followed by an activation function (such as a ReLU function). The final layer may use a Softmax activation function to convert the processed features into class probabilities. This allows FC classifier 44 to predict the likelihood of each tile's embedding belonging to a specific class (e.g., IDH mutant or wild-type) based on the relevant features identified by selfattention module 40.

[0076] Reference is now made to Fig. 4, which details clinical data integrator 36. Integrator 36 may comprise a Gaussian discriminant analysis module 50 and a logistic regression module 52. Gaussian discriminant analysis module 50 may use Gaussian discriminant analysis to transform patient data, such as age and sex data for each patient, into a single patient data score. Logistic regression module 52 may receive the single patient data score and the per tile probabilities, producedby DeepMIL classifier 32 for the tiles of the patient's slide, and may use logistic regression analysis to produce therefrom an exemplary IDH mutation status probability for the patient.

[0077] Once the training process has been completed, IDH mutation identifier trainer 16 may provide trained EsViT 30 and trained DeepMIL 62 to IDH mutation identifier 14.

[0078] Reference is now made to Fig. 5, which illustrates the elements of IDH mutation identifier 14 in operation. IDH mutation identifier 14 may comprise tile extractor 12, trained EsViT 30, trained DeepMIL 62, and clinical data integrator 18 and may receive an input-stained slide along with its associated patient data.

[0079] Tile extractor 12 may receive the input H&E slide as input and may extract its tiles. Trained EsViT 30 may process the tiles and may generate tile embeddings therefrom. Trained DeepMIL 62 may analyze the tile embeddings and may output the region-specific probabilities. Clinical data integrator 18 may receive patient data related to the H&E slide and may combine the region-specific probabilities from trained DeepMIL 62 with the associated patient data to produce an IDH mutation status probability as output. Thus, IDH mutation identifier 14 may integrate patient-specific clinical data with image-derived probabilities to generate more accurate and personalized IDH mutation status predictions.

[0080] It will be appreciated that trained EsViT 30, which is trained specifically on domainspecific pathology H&E slides, may recognize subtle patterns inherent in the pathology H&E slides, leading to improved performance metrics. In particular, this training may enable trained EsViT 30 to discern between the positive tile embeddings (associated with IDH mutant statuses) and negative tile embeddings (associated with IDH wild-type statuses), even though trained EsViT 30 was not explicitly trained for this specific task.

[0081] Reference is now made to Fig. 6, which shows the results for two slides 60 of IDH mutant glioma samples in the left column and for two slides 61 of IDH wild type glioma samples in the right column. A top row shows the slides 60 and 61, a middle row shows their associated immunohistochemistry, and a bottom row shows their attention maps as generated by trained DeepMIL 62.

[0082] In the immunohistochemistry row, for the IDH mutant samples, there are distinct dark- stained regions, labeled 64A and 64B, indicating positive staining for IDH mutation. However, in the immunohistochemistry row for the IDH wild type samples, there is only minimal staining, appearing mostly gray and faint.

[0083] In the attention map row, the attention for each tile is indicated by a dot whose color indicates the strength of the attention. Thus, the darker dots indicate areas of high attention. For the IDH mutant samples, areas of high attention are labeled 66 A and 66B and are outlined. It will be appreciated that the outlines of areas of high attention 66A and 66B correspond closely with the positively stained regions 64A and 64B in the immunohistochemistry images. The high attention tiles for IDH wild type samples, on the other hand, are dispersed throughout the tissue map, aligning with the IHC staining that didn't emphasize any specific tissue region.

[0084] Importantly, Fig. 6 may demonstrate spatial correlations between regions that tested positive in IHC staining with high-attention score regions in attention maps for IDH mutant samples. This correspondence may suggest that IDH mutation identifier 14 may successfully identify areas most indicative of IDH mutation status.

[0085] It will be noted that, in experiments, IDH mutation identifier 14 identified a set of slides as being IDH mutants, even though the pathologies initially classified the slides as being IDH wild-type, based on their IHC staining. These slides were later identified as IDH mutants through PCR tests.This finding suggests the potential of integrating IDH mutation identifier 14 as a decision support tool, especially for glioma cases that test negative in IHC staining.

[0086] Reference is now made to Fig. 7, which illustrates such a decision support tool, here labeled 70. Tool 70 may comprise IDH mutation identifier 14 along with the tests 72 and 74 to confirm the decision provided by IDH mutation identifier 14. Identifier 14 may process an input H&E slide which is classified by the pathologist as having IDH wild-type (i.e. a negative IDH result) and may produce either a negative or a positive IDH result. If IDH mutation identifier 14 classifies the input slide as positive, then the patient may be referred for molecular tests to confirm the IDH mutation status. Otherwise, the patient may be referred for further tests according to the guidelines.

[0087] Unless specifically stated otherwise, as apparent from the preceding discussions, it is appreciated that, throughout the specification, discussions utilizing terms such as "analyzing", "generating", "processing," "computing," "calculating," "determining," or the like, refer to the action and / or processes of a general purpose computer of any type, such as a client / server system, mobile computing devices, smart appliances, cloud computing units or similar electronic computing devices that manipulate and / or transform data within the computing system’s registers and / or memories into other data within the computing system’s memories, registers or other such information storage, transmission or display devices.

[0088] Tile extractor 12, IDH mutation identifier 14, IDH mutation identifier trainer 16, clinical data integrator 18, and tool 70 may be implemented on a suitable apparatus. This apparatus may be specially constructed for the desired purposes, or it may comprise a computing device or system typically having at least one processor and at least one memory, selectively activated or reconfigured by a computer program stored in the computer. The resultant apparatus when instructed by software may turn the general-purpose computer into inventive elements as discussed herein. The instructionsmay define the inventive device in operation with the computer platform for which it is desired. Such a computer program may be stored in a computer readable storage medium, such as, but not limited to, any type of disk, including optical disks, magnetic-optical disks, read-only memories (ROMs), volatile and non-volatile memories, random access memories (RAMs), electrically programmable read-only memories (EPROMs), electrically erasable and programmable read only memories (EEPROMs), magnetic or optical cards, Flash memory, disk-on-key or any other type of media suitable for storing electronic instructions and capable of being coupled to a computer system bus. The computer readable storage medium may also be implemented in cloud storage.

[0089] Some general-purpose computers may comprise at least one communication element to enable communication with a data network and / or a mobile communications network.

[0090] The processes and displays presented herein are not inherently related to any particular computer or other apparatus. Various general-purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct a more specialized apparatus to perform the desired method. The desired structure for a variety of these systems will appear from the description below. In addition, embodiments of the present invention are not described with reference to any particular programming language. It will be appreciated that a variety of programming languages may be used to implement the teachings of the invention as described herein.

[0091] While certain features of the invention have been illustrated and described herein, many modifications, substitutions, changes, and equivalents will now occur to those of ordinary skill in the art. It is, therefore, to be understood that the appended claims are intended to cover all such modifications and changes as fall within the true spirit of the invention.

Claims

CLAIMSWhat is claimed is:

1. A method for determining IDH status in gliomas, the method comprising: activating a trained classifier on region-specific feature embeddings of an input-stained slide of a patient to produce a set of region-specific probabilities; and combining patient-specific data with said set of region-specific probabilities to produce an isocitrate dehydrogenase (IDH) mutation status prediction for said input- stained slide.

2. The method according to claim 1 and also comprising dividing an H&E (hematoxylin and eosin) slide of a patient into a multiplicity of tiles, and wherein said activating comprises extracting said region-specific feature embeddings from said multiplicity of tiles.

3. The method according to claim 2 and wherein said extracting comprises activating a pre-trained vision transformer model on said multiplicity of tiles.

4. The method according to claim 1 and wherein said trained classifier is an attention-based deep multiple instance learning (DeepMIL) classifier.

5. The method according to claim 1 wherein said combining comprises performing a logistic regression on said IDH mutation status prediction and said patient-specific data to produce said IDH mutation status prediction.

6. The method of claim 5, wherein said patient- specific data comprises at least one of age and sex of said patient.

7. The method of claim 4, wherein said attention-based DeepMIL classifier comprises a selfattention module configured to generate attention scores for said region-specific feature embeddings.

8. The method according to claim 4, further comprising training said trained classifier using a dataset of H&E stained slides with known IDH mutation statuses.

9. The method according to claim 8, wherein said training comprises: extracting a plurality of tiles from each H&E stained slide from said dataset; generating tile embeddings for said plurality of tiles using a pre-trained vision transformer model; and training said DeepMIL classifier using said tile embeddings and corresponding IDH mutation statuses.

10. The method according to claim 9, wherein said pre-trained vision transformer model is an EsViT (self-supervised vision transformer) model.

11. The method according to claim 10, further comprising pre-training said EsViT model using a self-supervised learning approach on said dataset.

12. The method according to claim 11, wherein said self-supervised learning approach comprises: creating multiple augmented views of each tile; matching corresponding regions across said augmented views; and minimizing a distance between matched regions.

13. The method according to claim 9, wherein training said DeepMIL classifier comprises: assigning attention scores to tiles indicating their contribution to an overall IDH mutation status prediction of their corresponding H&E stained slide; combining said attention scores with said tile embeddings to generate weighted aggregations; and generating said set of region-specific probabilities from said weighted aggregations.

14. The method according to claim 13, further comprising using a weakly supervised loss function to compute a per slide loss value from a predicted IDH mutation status probability and an actual IDH mutation status per slide.

15. The method according to claim 14, wherein said weakly supervised loss function is a Binary Cross Entropy loss.

16. The method according to claim 15, further comprising updating weights of said attention-based DeepMIL classifier using an AdamW optimizer.

17. The method according to claim 1, further comprising generating an attention map for said input stained slide, wherein said attention map indicates regions of high attention corresponding to areas indicative of IDH mutation status.

18. The method according to claim 1, wherein said method is implemented as part of a decision support tool for glioma cases that test negative in immunohistochemistry (IHC) staining and comprises: recommending molecular tests for confirming IDH mutation status when said IDH mutation status prediction indicates a positive result.

19. A system for determining IDH status in gliomas, the system implemented on a computing device having a processor and a memory, the system comprising: a trained classifier to process region-specific feature embeddings of an input stained slide of a patient to produce a set of region-specific probabilities; and a clinical data integrator to combine patient-specific data with said set of regionspecific probabilities to produce an isocitrate dehydrogenase (IDH) mutation status prediction for said input stained slide.

20. The system according to claim 19 and also comprising a tile extractor to divide an H&E (hematoxylin and eosin) slide of a patient into a multiplicity of tiles, and wherein said trained classifier is configured to extract said region-specific feature embeddings from said multiplicity of tiles.

21. The system according to claim 20 and wherein said trained classifier comprises a pre-trained vision transformer to extract said region-specific feature embeddings from said multiplicity of tiles.

22. The system according to claim 21 and wherein said trained classifier also comprises an attention-based deep multiple instance learning (DeepMIL) classifier.

23. The system according to claim 19 wherein said clinical data integrator comprises a logistic regression module to perform a logistic regression on said IDH mutation status prediction and said patient-specific data to produce said IDH mutation status prediction.

24. The system of claim 23, wherein said patient-specific data comprises at least one of age and sex of said patient.

25. The system of claim 22, wherein said attention-based DeepMIL classifier comprises a selfattention module to generate attention scores for said region-specific feature embeddings.

26. The system according to claim 22, further comprising a trainer to train said trained classifier using a dataset of H&E stained slides with known IDH mutation statuses.

27. The system according to claim 26, wherein said trainer comprises: said tile extractor to extract a plurality of tiles from each H&E stained slide of said dataset; said pre-trained vision transformer to generate tile embeddings for said plurality of tiles using a pre-trained vision transformer model; anda loss function to train said DeepMIL classifier using said tile embeddings and corresponding IDH mutation statuses.

28. The system according to claim 27, wherein said pre-trained vision transformer is an EsViT model.

29. The system according to claim 28, wherein said trainer is further configured to pre-train said EsViT model using a self-supervised learning approach on said dataset.

30. The system according to claim 29, wherein said self-supervised learning approach comprises: creating multiple augmented views of each tile; matching corresponding regions across said augmented views; and minimizing a distance between matched regions.

31. The system according to claim 25, wherein said region-specific feature embeddings indicate their contribution to an overall IDH mutation status prediction of their corresponding H&E stained slide.

32. The system according to claim 27, wherein said loss function is a weakly supervised loss function to compute a per slide loss value from a predicted IDH mutation status probability and an actual IDH mutation status per slide.

33. The system according to claim 32, wherein said weakly supervised loss function is a Binary Cross Entropy loss.

34. The system according to claim 33, wherein said trainer is further configured to update weights of said attention-based DeepMIL classifier using an AdamW optimizer.

35. The system according to claim 19, further comprising an attention map generator configured to generate an attention map for said input stained slide, wherein said attention map indicates regions of high attention corresponding to areas indicative of IDH mutation status.

36. The system according to claim 19, wherein said system is implemented as part of a decision support tool for glioma cases that test negative in immunohistochemistry (IHC) staining and comprises: a recommendation module configured to recommend molecular tests for confirming IDH mutation status when said IDH mutation status prediction indicates a positive result.

Citation Information

Patent Citations

  • Platform, device and process for annotation and classification of tissue specimens using convolutional neural network

    US20200272864A1