Predicting immunotherapy outcomes using deep learning

Deep learning models enhance immunotherapy outcome prediction by analyzing tissue samples with multiple machine learning techniques, providing accurate therapy success probabilities and reducing resource consumption, while identifying novel biomarkers for precision medicine and various therapeutic applications.

WO2025170862A1PCT designated stage Publication Date: 2025-08-14VERILY LIFE SCIENCES LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/014305
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-09
Filing Date
2025-02-03
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Existing techniques for predicting immunotherapy outcomes, such as checkpoint blockade immunotherapy for cancer, are not accurate enough, with biomarkers like PD-1/PD-L1 status providing reliable indications only about half the time, leading to inefficiencies and resource waste.

Method used

A deep learning approach using multiple machine learning models, including a classification model, a self-supervised pathology foundation model, and a DeepMIL model with a gated attention mechanism, analyzes tissue samples to predict therapy outcomes by encoding pathological structures into embedded representations and determining therapy success probabilities.

Benefits of technology

The deep learning models provide more accurate predictions of immunotherapy outcomes, reducing computational resource needs and identifying novel biomarkers, improving precision medicine and therapeutic applications beyond cancer, such as tumor detection and diagnostics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025014305_14082025_PF_FP_ABST
    Figure US2025014305_14082025_PF_FP_ABST
Patent Text Reader

Abstract

Techniques for predicting immunotherapy outcomes using deep learning are disclosed. In an example method, a computing device receives an image of a tissue sample. The computing device classifies, using a first machine learning model, one or more pathological structures in the tissue sample. The computing device generates, using a second machine learning model, an embedded representation for a first pathological structure from the one or more pathological structures. The computing device determines, using a third machine learning model, a prediction relating to an outcome of applying a therapy. Responsive to the prediction exceeding a predetermined threshold, the computing device outputs an indication corresponding to a likelihood of success of the therapy.
Need to check novelty before this filing date? Find Prior Art

Description

PREDICTING IMMUNOTHERAPY OUTCOMES USING DEEP LEARNINGCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to provisional application U.S. Ser. No. 63 / 551,744 entitled “Predicting Immunotherapy Outcomes Using Deep Learning’ and filed on February 9, 2024. the entire disclosure of which is incorporated herein by reference for any purpose.FIELD

[0002] The present application generally relates to image analysis using deep learning and more particularly relates to predicting immunotherapy outcomes using deep learning.BACKGROUND

[0003] The immune system of the human body has a natural tendency to attack cancer cells. However, the immune system response to a dangerous cancer cell can be muted by an adaptation that exploits the immune system’s own tendency to modulate itself to minimize collateral tissue damage.

[0004] This modulation can occur by way of immune checkpoints, which are regulatory pathways in immune cells, such as T-cells. Proteins such as PD-1, PD-L1, and CTLA-4, which are found on the surface of T-cells, may biochemically recognize and bind to partner proteins on other cells, such as some cancer cells. When these proteins bind together, the immune system response is inhibited, preventing the bound T-cells from attacking the cancer cells.

[0005] In checkpoint blockade immunotherapy, drugs known as checkpoint inhibitors are used to block certain proteins from binding to enable the T-cells to instead attack and destroy the cancer cells. This form of immunotherapy has proven successful in treating certain types of cancers such as melanoma, non-small cell lung cancer, and kidney cancer.SUMMARY

[0006] Various examples are described for predicting immunotherapy outcomes using deep learning. One example method includes receiving an image of a tissue sample; classifying, using a first machine learning model, one or more pathological structures in the tissue sample; generating, using a second machine learning model, an embeddedrepresentation for a first pathological structure from the one or more pathological structures; determining, using a third machine learning model, a prediction relating to an outcome of applying a therapy; and responsive to the prediction exceeding a predetermined threshold, outputting an indication corresponding to a likelihood of success of the therapy.

[0007] One example system for predicting immunotherapy outcomes using deep learning includes a non-transitory computer-readable medium, one or more processors in communication with the non-transitory computer-readable medium, the one or more processors configured to execute processor-executable instructions stored in the non- transitory computer-readable medium configured to cause the one or more processors to receive an image of a tissue sample; classify, using a first machine learning model, one or more pathological structures in the tissue sample; generate, using a second machine learning model, an embedded representation for a first pathological structure from the one or more pathological structures; determine, using a third machine learning model, a prediction relating to an outcome of applying a therapy; and responsive to the prediction exceeding a predetermined threshold, output an indication corresponding to a likelihood of success of the therapy.

[0008] One example non-transitory computer-readable medium for predicting immunotherapy outcomes using deep learning comprising processor-executable instructions configured to cause one or more processors to receive an image of a tissue sample; classify, using a first machine learning model, one or more pathological structures in the tissue sample; generate, using a second machine learning model, an embedded representation for a first pathological structure from the one or more pathological structures; determine, using a third machine learning model, a prediction relating to an outcome of applying a therapy; and responsive to the prediction exceeding a predetermined threshold, output an indication corresponding to a likelihood of success of the therapy.

[0009] In some embodiments, an apparatus is provided, which includes means for implementing part or all of the operations and / or methods disclosed herein.

[0010] In some embodiments, a computer program product is provided, which includes computer instructions that, when executed by a processor, implement part or all of the operations and / or methods disclosed herein.

[0011] These illustrative examples are mentioned not to limit or define the scope of this disclosure, but rather to provide examples to aid understanding thereof. Illustrative examples are discussed in the Detailed Description, which provides further description. Advantages offered by various examples may be further understood by examining this specification.BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings, which are incorporated into and constitute a part of this specification, illustrate one or more certain examples and, together with the description of the example, serve to explain the principles and implementations of the certain examples.

[0013] Figure 1 shows an example system for predicting immunotherapy outcomes using deep learning, according to some examples of the present disclosure.

[0014] Figure 2 depicts a detail view of a system including an example implementation of the image analysis software shown in Figure 1, according to some examples of the present disclosure.

[0015] Figure 3 shows a system for predicting immunotherapy outcomes using deep learning, including a detail view of an example implementation of a deep multiple instance learning model, according to some examples of the present disclosure.

[0016] Figures 4A-4C show graphs illustrating the effectiveness of the techniques of the present application compared with certain existing techniques, according to some examples of the present disclosure.

[0017] Figure 5 shows an example method for predicting immunotherapy outcomes using deep learning training as shown in detail in Figure 2, according to some examples of the present disclosure.

[0018] Figure 6 shows an example computing device suitable for use in example systems or methods for predicting immunotherapy outcomes using deep learning according to this disclosure, according to some examples of the present disclosure.DETAILED DESCRIPTION

[0019] Examples are described herein in the context of predicting immunotherapy outcomes using deep learning. Those of ordinary skill in the art will realize that the following description is illustrative only and is not intended to be in any way limiting. Reference will now be made in detail to implementations of examples as illustrated in the accompanying drawings. The same reference indicators will be used throughout the drawings and the following description to refer to the same or like items.

[0020] In the interest of clarity, not all of the routine features of the examples described herein are shown and described. It will, of course, be appreciated that in the development of any such actual implementation, numerous implementation-specific decisions must be made in order to achieve the developer’s specific goals, such as compliance with application- and business-related constraints, and that these specific goals will vary from one implementation to another and from one developer to another.

[0021] Checkpoint blockade immunotherapy is an increasingly effective method for treatment of cancer, and in particular, lung cancer. However, patients’ response to checkpoint blockade immunotherapy can vary. To avoid waste, prevent needless discomfort, and to control costs, identification of patients who will respond favorably to checkpoint blockade immunotherapy is an important aspect of developing a treatment plan for each individual patient.

[0022] Existing techniques rely on biomarkers to identify patients who will respond favorably to checkpoint blockade immunotherapy. For example, some existing techniques use a biomarker that includes a patient’s PD-1 / PD-L1 status. PD-1 / PD-L1 status refers to the expression or presence of the proteins PD-1 (Programmed Death-1) on T-cells and the protein PD-L1 (Programmed Death-Ligand 1) on cancer cells or other types of cells in the tumor environment. However, experience with existing techniques has demonstrated that use of this and other biomarkers is an accurate indication of the effectiveness of checkpoint blockade immunotherapy only about half of the time or less.

[0023] Techniques for predicting immunotherapy outcomes using deep learning are provided that can more accurately predict immunotherapy outcomes than existing techniques. In an example method, a computing device first receives an image of a tissue sample. The tissue sample may be obtained, for example, from a patient that is a candidate for a therapy. The patient may be a candidate for checkpoint blockade immunotherapy to treat non-small cell lung cancer. The tissue sample can be obtained from the patient using a technique such as by biopsy or resection. The tissue sample is then stained using dyes such as hematoxylin and eosin (H&E).

[0024] The computing device classifies, using a first machine learning (ML) model, one or more pathological structures in the tissue sample. For example, the first ML model may be a classification model based on a neural network that is trained to identify pathological structures such as tumor regions within the stained tissue sample.

[0025] The computing device generates, using a second ML model, an embedded representation for a particular pathological structure from the one or more pathological structures. For example, the second ML model may be a self-supervised pathology foundation model. The self-supervised pathology foundation model can be used to obtain a compressed visual representation of each pathological structure, also known as an embedding or embedded representation. For instance, the particular pathological structure (e.g., a tumor region) can be encoded by the self- supervised pathology foundation model into an embedded vector representation of a predetermined, lower dimension.

[0026] The computing device then determines, using a third ML model, a prediction relating to an outcome of applying a therapy. For example, the third ML model may be a trained deep multiple instance learning (DeepMIL) model with a gated attention mechanism trained to predict the binary 1-year overall survival (OS) status for each patient. The binary 1-year OS, in this context, refers to whether the patient is living or not living one year from the time treatment is commenced. In contrast to traditional ML models trained using supervised learning, a DeepMIL ML model operates on sets or “bags” in which a label is provided for the entire “bag.” For example, the entire embedded representation is associated with the outcome via a label, but only a portion of the embedded representation actually contains the information related to the outcome. The DeepMIL ML model is thus trained to predict outcomes for the therapy based on the classified content of the region containing the pathological structure, in which the exact location or characteristics of the pathology (e.g.. the precise biomarkers indicative of response to the therapy) in each image are not annotated and may not even be known.

[0027] The computing device, responsive to the prediction exceeding a predetermined threshold, outputs an indication corresponding to a likelihood of success of the therapy. For example, the computing device can update a graphical user interface (GUI) to include an indication that the prediction has exceeded the predetermined threshold, which may cause a healthcare provider to apply the therapy (e.g., an immunotherapy regime) to a patient.

[0028] The techniques for predicting immunotherapy outcomes using deep learning described herein result in an advantageous technical effect by improving the technical field of image analysis using deep learning for therapeutic applications. The predictions output by, for example, the DeepMIL ML model exhibit a stronger association with outcomes compared to using biomarkers such as PD-1 / PD-L1 status. The techniques demonstrate the potential of deep learning using pathology foundation models to improve immunotherapy outcomes prediction. Use of these techniques may additionally enable the discovery of novel biomarkers from stained tissue samples that can contribute to advances in precision medicine.

[0029] Furthermore, the techniques are not limited to predictions for immunotherapy effectiveness or cancer. The ability to use deep learning ML models such as a DeepMIL ML model can be applied to image analysis for a variety of therapeutic applications. For example, with respect to cancer and in addition to immunotherapy, deep learning ML models such as the attention-based multiple instance learning approach disclosed herein can be used for tumor detection and localization, assessing tumor size and i)growth rate, planning radiation therapy, monitoring metastasis spread, evaluating tumor heterogeneity, stratifying patients for clinical trials, identifying genetic mutations through imaging phenotypes, and so on. Beyond cancer, the techniques can be used for assays such as early detection of Alzheimer's disease, evaluation of liver fibrosis stages, detection of coronary artery disease, and so on. The techniques can additionally be combined with explanatory analysis of ML model output to determine which aspects and to what extent of the stained samples the deep learning ML models are relying on for generation of predictions to further improve the field of diagnostics and efficacy determination using image analysis.

[0030] The techniques disclosed herein can also improve the functioning of a computer. Manual image analysis by radiation oncologists may be error-prone and require frequent repetition and rework, consuming computational resources while images are stored, processed, and transmitted to third-parties for analysis. Additionally, the techniques as presented herein may obviate the need for biomarker analysis, such as determination of PD-1 / PD-L1 status, which may require further manual or automated image analysis, saving additional computational resources. Finally, once the ML models, such as the DeepMIL, are trained, subsequent image analysis, particular on reduceddimensional embedded representations may consume significantly fewer computational resources. Likewise, the use of pre-trained models for, for example, classification means that the overall training burden and concomitant need for computational resources may be reduced.

[0031] This illustrative example is given to introduce the reader to the general subject matter discussed herein and the disclosure is not limited to this example. The following sections describe various additional non-limiting examples and examples of predicting immunotherapy outcomes using deep learning.

[0032] Referring now to Figure 1. Figure 1 shows an example system 100 for predicting immunotherapy outcomes using deep learning. The system 100 includes an imaging system 150 that is connected to a computing device 110. The computing device 110 has image analysis software 116. which includes multiple trained ML models 120-126 for classification, encoding, and analyzing images of tissue sample. The computing device 110 is connected to an imaging system 150, a display device 114. a local data store 112, and to a remote server 140 via one or more communication networks 130. The remote server 140 is, in turn, connected to its own data store 142.

[0033] The multiple ML models 120-126 in the image analysis software 116 can be trained and provided by the remote server 140. The ML models 120, 122, 124, and 126 arejust examples of trained ML models. There can be less than four trained ML models or more than four trained ML models in the image analysis software 116. For instance, in the example method described above, three ML models are used. In some examples, a trained ML model can be located on the remote server 140 and accessed remotely via, for example, a web-based application programming interface (API) by the image analysis software 116.

[0034] The ML models 120, 122, 124, and 126 may include, for example, a classification model for identifying or classifying pathological structures in a tissue sample, an ML model for generating an embedded representation of the pathological structures, or an ML model to determine, based on the embedded representations, a prediction relating to an outcome of applying a therapy, such as immunotherapy.

[0035] The remote server 140 can train an ML model and provide one or more trained ML models for predicting immunotherapy outcomes using deep learning. In some examples, the remote server 140 can use training data including sets of images of stained tissue samples of a particular tissue type. Some tissue samples may be stained using dyes such as hematoxylin and eosin (H&E). Other example dyes can include the Hoechst stain. DAPI stain, Masson's trichrome stain. Papanicolaou (Pap) stain, silver stain. Giemsa stain, chromogenic or fluorescent immunostains, and so on. Certain images may include tissue edges showing edge effects, for example tearing, deformation, or damage at the edges, leading to local discrepancies between adjacent sections. The image analysis software 116 can generate and apply masks to exclude edges or other portions of stained images from being used in model training.

[0036] While the process of training an ML model occurs on the remote server 140. in some examples, a third-party provider (not shown) trains ML models for predicting immunotherapy outcomes using deep learning. The third-party provider can train ML models for generating predictions for, for example, different types of tissue samples, cancers, sampling methods, etc., and provide trained ML models to the remote server 140, which can then provide the trained ML models to the computing device 110.

[0037] The imaging system 150 includes a microscope and camera to capture images of pathology samples. Imaging system 150 in this example is a conventional pathology imaging system that can capture digital images of tissue samples, stained or unstained, using broad-spectrum visible light. The imaging system 150 can include (for example) a microscope (e.g., a light microscope) and / or a camera. In some instances, the camera is integrated within the microscope and the microscope can include a stage on which the portion of the sample (e.g., a slice mounted onto a slide) is placed, one or more lenses (e.g., one or more objective lenses and / or an eyepiece lens), one or more focuses, and / or a lightsource. The camera may be positioned such that a lens of the camera is adjacent to the eyepiece lens. In some instances, a lens of the camera is included within image collection system 104 in lieu of an eyepiece lens of a microscope. The camera can include one or more lenses, one or more focuses, one or more shutters, and / or a light source (e.g., a flash). The digital images from the imaging system 150 can be autofluorescence images. Alternatively, the imaging system 150 can implement other suitable imaging techniques.

[0038] The tissue samples can include, but are not limited to, a sample collected via a biopsy (such as a core-needle biopsy), fine needle aspirate, surgical resection, or the like. In one scenario, a tissue sample can be prepared for imaging within the conventional imaging system 150. such as by obtaining one or more thin slices of tissue taken from a patient, applying a suitable stain, and positioning them on corresponding slides, which are then inserted in sequence into the imaging system 150. The imaging system 150 then captures images of stained samples and provides them to the computing device 110. A set of stained images may be then generated by the imaging system 150 and each image of the set of images may correspond to different portions of the biological sample or samples.

[0039] The computing device 110 receives digital images from the imaging system 150 corresponding to a particular tissue sample and provides them to one of the ML models 120-126 to generate a likelihood of success of a particular therapy based on the tissue sample. After receiving the captured stained image or multiple captured stained images, the computing device 110 may store the image(s) in the local data store 112. It then executes the image analysis software 116 on an image for a particular biological sample. The tissue samples can be displayed via a display device 114.

[0040] While in this example, the entire process occurs on the local computing device 110 and imaging system 150, such an arrangement is not needed. For example, an example system may omit the imaging system 150. Instead, the computing device 110 could obtain whole slide images from the local data store 112 or from the remote server 140.Alternatively, while image analysis software 116 is executed at the computing device 110, in some examples, the whole slide images may be provided to the remote server 140, which may execute image analysis software 116, including suitable ML models, e.g., ML models 120-126. Thus, the system shown in Figure 1 may, according to different examples, provide predictions relating to therapy outcomes in contexts in which tissue samples are available locally or by receiving images of pathology tissue from a third party for processing, including in a cloud environment provided by a remote server 140.

[0041] Turning next to Figure 2, Figure 2 depicts a detail view of a system 200 including an example implementation of the image analysis software 116 shown in Figure1. In example system 200, the imaging system 150 receives a stained tissue sample. The tissue sample may. for example, be obtained from a patient (not shown). The patient may be a candidate for a therapy, such as an immunotherapy treatment for a particular type of cancer. The imaging system 150 generates an image of the stained tissue sample 205 that is suitable for image analysis by the ML models 120-126 or the ML models described below. Figure 2 depicts an example of a stained tissue sample image 205 of a lung biopsy in support of a determination of an effectiveness of an immunotherapy treatment for nonsmall cell lung cancer.

[0042] The stained tissue sample image 205 is received by image analysis software 116. In the example of system 200, the stained tissue sample image 205 is first processed by classification model 210. The classification model 210 can be used for identifying or classifying pathological structures 215A...N in a tissue sample. The classification model 210 may be, for example, a neural network that is trained to identify regions of interest in stained tissue sample images including certain pathological structures 215A...N.

[0043] The neural network may be or may include, for example, a convolutional neural network (CNN). Other types of neural networks or ML models may be used alone or in combination, such as recurrent neural networks (RNN), long short-term memories (LSTM), generative adversarial networks (GAN), support vector machines (SVM), random forests, gradient boosting machines (GBM), a k-nearest neighbors algorithm, among many others.

[0044] In some examples, the classification model 210 can be developed and trained by external research organizations such as the Cancer Genome Atlas of the Center for Cancer Genomics at the National Cancer Institute. In that case, the classification model 210 may be accessed through an API provided by the external research organization. In some examples, the external research organizations can export the trained classification model 210 for local use on, for example, the computing device 110, as shown in Figure 1.

[0045] The remote server 140 can train an ML model and provide one or more trained ML models for predicting immunotherapy outcomes using deep learning. For example, the remote server 140 can be used for training an ML model for generating an embedded representation for a pathological structure identified by the classification ML model. In some examples, the ML model for generating the embedded representation can be a self-supervised foundation model. In this context, “foundation model” refers to an ML model trained on diverse datasets self-supervised learning techniques. The foundation model can be trained to learn embedded representations 225A...N of the underlying data, such as the embedded representations 225A... N of the identified pathological structures215A... N, useful for capturing general features applicable across a wide range of tasks. The foundation model is then generally adaptable to specific tasks through further fine-tuning. For example, an ML models such as a DeepMIL model 230 with a gated attention mechanism can be trained to make predictions using the embedded representation generated by the foundation model 220.

[0046] In some examples, the foundation model 220 is an encoder trained to generate the embedded representation. For example, the encoder may be a CNN trained to generate embedded representations of the imaged tissue samples. The encoder may also be part of an autoencoder, which can include both an encoder and a decoder, trained to minimize the difference between the input image and its reconstructed version. As with the classification ML model 210, the foundation model 220 can be or be a part of a combination including any number of ML model types, including, but not limited to the example types listed above.

[0047] In some examples, the encoder is trained using self- supervised training methods. For instance, the encoder can be trained to generate embedded representations 225A... N by deriving its own supervisory signals from the input data. Self- supervised training techniques used for generation of embedded representations include contrastive learning, predictive coding, data augmentation, “jigsaw puzzle” solving, and so on.

[0048] The foundation model 220, the ML model for generating embedded representations 225A... N, may be trained to generate an embedded representation that is a vector or other numerical representation of the image of the tissue sample. The vector, for example, may be a certain-dimensional (e.g., MxN) array in which each element represents a specific characteristic or feature extracted from the image, such as texture, color, or shape information. The training method may be selected to be robust to variations in imaging conditions, ensuring consistent representation across different samples. For example, an MxN pixel image of a tissue sample image, or portion thereof, may be encoded to a fixed- length, 1- or 2-D vector by a CNN. The vector may be of significantly lower dimension than the original MxN pixels.

[0049] In another example, the remote server 140 can be used for training an ML model to determine a prediction relating to an outcome of applying a therapy, such as immunotherapy. In particular, the prediction ML model can be trained based on the embedded representations 225A... N generated by the foundation model 220, such as the ML model for generating the embedded representation, as described above.

[0050] The prediction ML model can be, for example, a deep neural network using a multiple instance learning (DeepMIL) technique. In a DeepMIL model 230, an ML model istrained using training data that includes “sets” or “bags." each set or bag containing multiple instances of interest, rather than individual labeled instances.

[0051] There are various techniques to label the bags based on available information. In one example technique, the bag label is known. For example, the bag label could be the binary 1-year OS status of a patient, in which it is known whether the patient is living or not living one year from the time treatment is commenced. In another example technique, each bag is labeled positive if it contains at least one positive instance; otherwise, it is labeled negative. For instance, if an embedded representation of a tissue sample in a training data set includes both a pathological structure and non-pathological structure, then the embedded representation would be labeled as including the pathological structure.

[0052] The DeepMIL 230 model thus trained using MIL techniques can predict the labels of bags not included in the training data. In some examples, the DeepMIL model 230 can include neural networks and other related components. For example, some ML models utilizing MIL trained for image analysis may include convolutional neural networks (CNNs), fully-connected networks (FCNs). transformers, or other neural networks. Some DeepMIL models 230 may include attention mechanisms that can. for example, adjust ML model parameters or weights during training according to a particular bag’s contribution to the bag’s label. An “attention mechanism” in the context of image analysis refers generally to a component in a neural network that selectively focuses on certain parts of an image while processing information, which can have the effect of apparently prioritizing and weighing different regions or features of an image based on their relevance to the prediction task.

[0053] Example system 200 depicts the ML models 210, 220, and 230 as separate components. However, in some examples, some or all of the functions of one or more of the ML models 210, 220, and 230 may be combined. For example, the functions of ML models 210, 220, and 230 may be combined into a single ML model. In another example, the functions of two ML models 210, 220 could be combined into a single ML model whose output is input to DeepMIL model 230. Additionally, some examples may omit some or all functions of the ML models 210, 220. and 230. For example, the classification performed by classification model 210 may be omitted. In this case, the tissue sample image 205 may be input to the foundation model 220, given suitable pre-processing, and all regions of the input image may be used for generation of embedded representations.

[0054] In some examples, the attention mechanism is a gated attention mechanism. In a gated attention mechanism trained for image analysis, the attention mechanism caninclude “gates”' that control the relative impact of information from different parts of the image during training and inference. The gates can thus determine the extent to which different parts of the image contribute to the output prediction. For example, some gating mechanism implementations may utilize a sigmoid activation function to modulate contributions from different parts of the image.

[0055] The DeepMIL model 230 is trained to output likelihood 235. Likelihood 235 may be a probability score, representing the likelihood of a certain outcome or class such as a positive outcome for the therapy, which can be used along with a threshold value to make a classification. For example, a threshold value of 60% probability may be used to make a classification for a particular example implementation. Alternatively, likelihood 235 could be a quantification of risk level, a confidence level, a binary decision or a set of values for a number of classes, a continuous score that quantifies the extent or magnitude of a certain condition or attribute, and so on, according to the particular scenario.

[0056] Turning next to Figure 3, Figure 3 shows a system 300 for predicting immunotherapy outcomes using deep learning, including a detail view of an example implementation of a DeepMIL model 230 as may be used in certain embodiments. In the example system 300. DeepMIL model 230 includes a first convolutional neural network layer (CNN) 305. For example, the first convolutional neural network layer 305 may apply filters, sometimes called kernels, to the input tissue sample image 205 to create feature maps. The filters may be trained to detect low-level features such as edges, colors, or textures. The application of the filters can include convolution of successive portions of the input tissue sample image 205 with the filters followed by application of an activation function, such as the Rectified Linear Unit (ReLU) nonlinear function.

[0057] DeepMIL model 230 includes a second CNN layer 310. The second CNN layer 310 can be trained to apply a similar convolution operation to the reduced-dimensional feature maps output by the first CNN layer 305. In the second CNN layer 310, however, the filters can be trained to recognize more complex patterns within the image such as tissues or anomalies.

[0058] DeepMIL model 230 includes a first fully-connected network (FCN) 315, sometimes referred to as a fully-connected layer. The first FCN 315 following the CNN layers can receive input from each element of the output feature map and thus combine the learned features to integrate and interpret the high-level features extracted by the CNN layers 305, 310. In some FCN implementations, neurons perform a weighted sum of their inputs, followed by a nonlinear activation function, to learn non-linear combinations of the extracted features.

[0059] DeepMIL model 230 includes a gated attention mechanism 320. Some DeepMIL model implementations may include an assumption of non-ordering or no dependency of instances within each feature bag or set. sometimes referred to as permutation-invariance. Following transformation of the input tissue sample images as described above, the permutation invariance can be implemented using an MIL pooling function.

[0060] In some implementations, the MIL pooling function is an attention mechanism. For instance, the attention mechanism can include a neural network with weights used to determine a weighted average of feature sets or bags (or low-dimensional embeddings thereof). The weighted average may include nonlinear functions such as the hyperbolic tangent (tanh) so that similarities and dissimilarities can be learned during training.

[0061] In some examples, the nonlinear functions may not efficiently learn complex relation within or between sets or bags when, for example, the selected nonlinear functions are linear over domain ranges of interest, which may limit the final expressiveness of learned relations among instances. Training of the attention mechanism may thus include a gating mechanism such as a sigmoid non-linearity that can introduce a non-linearity over the linear ranges of concern.

[0062] DeepMIL model 230 includes a second FCN 325. As with the first FCN 315, the second FCN 325 can receive input from each output neuron of the gated attention mechanism 320 and combine the learned features. The second FCN 325 can be trained to output a prediction or likelihood associated with the probable success of a therapy, such as immunotherapy.

[0063] The example DeepMIL implementation described above can additionally be used to determine the weights assigned, during training and inference, to particular instances within a set or bag because the gated attention mechanism 320 operates at the instance-level of the constituents of a set or bag. The weights thus determined can be used to find key instances within a set or bag. The weights of the gated attention mechanism 320 can additionally be used to interpret the output predictions or likelihoods in terms of instance -level labels. Thus, the example DeepMIL model implementation above as well as the gated attention mechanism 320 can provide additional insight to the interpretability of the output prediction or likelihood, including information about particular regions of interest within the tissue sample image.

[0064] Referring now to Figures 4A-4C, Figures 4A-4C show graphs illustrating the effectiveness of the techniques of the present application compared with certain existingtechniques. The effectiveness is illustrated using Kaplan-Meier (K-M) survival curves. K-M survival curves can show a graphical representation of the probability of an individual or a group of individuals surviving from the time of their initial diagnosis or treatment until a specified future time point. In Figures 4A-4C, the graphs show the probability of survival (y-axis) over a given period of time (x-axis. about 1 year in units of days). The hazard ratios (HR) and p-values of Figures 4A-4C were calculated using a univariate Cox regression, but other comparable mathematical methods may yield similar results.

[0065] The K-M curves depicted in Figures 4A-4C illustrate the predicted overall survival rates for non-small cell lung cancer patients with a known positive or negative status. The status can be, for example, a known status of a clinical biomarker such as PD- L1 status or a predicted status such as the output from an ML model, and so on. The known positive PD-L1 status corresponds to a tumor proportion score (TPS) of at least 1%. The TPS quantifies the percentage of tumor cells showing a specific characteristic, such as in this case, the presence of a particular protein marker. For instance, a TPS of at least 1% means that out of all the tumor cells observed in the sample, at least 1% of them exhibit a positive PD-L1 status.

[0066] In Figure 4A, the K-M survival curves 405 are shown using only PD-L1 status as a predictor of survival, following application of anti-PD-l / PD-Ll immunotherapy. The survival curve 410 corresponds to patients with positive PD-L1 status. Survival curve 410 shows some improvement over the survival curve 412, corresponding to patients with negative PD-L1 status. This improvement is quantified by the hazard ratio 414 of 0.65, which suggests a modestly reduced risk of the event in the subset with positive PD-L1 status. The hazard ratio is a measure of effectiveness of a predictor or model in modeling the predicted risk of an event (e.g., death) between two groups. In this case, the two groups are the group of patients with positive PD-L1 status and the group of patients with negative PD-L1 status. Use of PD-L1 status as a predictor of survival is accompanied by a relatively high p-value 416 of 0.32, indicating a low level of statistical confidence in the use of PD-L1 status as a predictor of survival following the application of immunotherapy.

[0067] In Figure 4B. the K-M survival curves 425 are shown using a linear-probe logistic regression model as a predictor of survival, following application of anti-PD- 1 / PD-L1 immunotherapy for comparison with the DeepMIL approach of the present application, as described below. The linear-probe logistic regression model is selected as a baseline for comparison. The survival curve 430 corresponds to patients with positive status predicted by the linear-probe logistic regression model. Survival curve 430 again shows some improvement over the survival curve 432. Survival curve 432 corresponds to patients withnegative status predicted by the linear-probe logistic regression model. However, use of the linear-probe logistic regression model results in a lower hazard ratio 434 of 0.46, indicating an improved usefulness of using a linear-probe logistic regression model as a predictor of survival following application of anti-PD-l / PD-Ll immunotherapy. This model is accompanied by a significantly lower p-value 436 of 0.09, potentially indicating a greater statistical significance, although still short of 0.05, a typical threshold for statistical significance.

[0068] In Figure 4C. the K-M survival curves 445 are shown using the techniques of the present disclosure, including the DeepMIL model 230 as a predictor of survival, following application of anti-PD-l / PD-Ll immunotherapy. The survival curve 450 corresponds to patients with positive status as predicted by the DeepMIL model. Survival curve 450 again shows some improvement over the non-survival curve 452, which corresponds to patients with negative status predicted by the DeepMIL model. The use of the DeepMIL model 230 also results in a lower hazard ratio 454 of 0.40, indicating an improved accuracy of using the DeepMIL model 230 as a predictor of survival following application of anti-PD-l / PD-Ll immunotherapy. This model is accompanied by a lower p- value 456 of 0.04. that is less than 0.05, corresponding to statistical significance and the superiority of the techniques of the present disclosure over existing techniques. The improved statistical significance is further illustrated in that the confidence interval regions 458, 460 have minimal overlap.

[0069] Similar analyses illustrating the effectiveness of the techniques of the present disclosure over existing techniques can be performed on different patient groups and subsets thereof. For example, similar outcomes may be obtained for predictions applied to patient populations with or without a known PD-L1 status. Likewise, a multivariable Cox regression can be used to adjust the predicted outcomes for age group and smoking status to obtain similar results.

[0070] Referring now to Figure 5. Figure 5 shows an example method 500 for predicting immunotherapy outcomes using deep learning training as shown in detail in Figure 2. The method 500 will be described with respect to the example system 200 shown in Figure 2; however, any suitable system according to this disclosure may be used.

[0071] At block 510. an image analysis software 116 receives an image of a tissue sample 205 as described above with respect to Figure 2. The tissue sample 205 may be, for example, obtained from a patient who is a candidate for a therapy. For example, the method 500 can be used for predicting a likelihood of the effectiveness of an immunotherapy treatment, such as anti-FD-l / rD-Ll immunotherapy, for combating non-small cell lung cancer. The image can be generated from an imaging system 150 based on a tissue sample. The tissue samples may be stained using dyes such as hematoxylin and eosin (H&E), or suitable other dyes according to the particular application or treatment.

[0072] At block 520, the image analysis software 116 classifies, using a first machine learning (ML) model, one or more pathological structures 215A... N in the tissue sample 205. For example, the one or more pathological structures 215A... N in the tissue sample 205 may be tumor patches, such as a localized area within a tissue sample where tumor cells are present, corresponding to regions where cancerous cells have formed a cluster or group. For instance, the tumor patches may be caused by non-small cell lung cancer.

[0073] The classification may be performed by, for example, classification model 210 including a neural network trained for identifying or classifying pathological structures 215A...N in a tissue sample image 205. The neural network may be or include, for example, a convolutional neural network (CNN). For instance, features may first be extracted from the tissue sample images through convolution and pooling layers, followed by fully connected layers trained to interpret the extracted features for classification.

[0074] Other types of neural networks or ML models may be used for classification instead of or in combination with a CNN such as support vector machines (SVM). random forest classifiers, gradient boosting machines (GBM), decision trees, naive Bayes classifiers, recurrent neural networks (RNN), long short-term memory networks (LSTM), generative adversarial networks (GAN), multi-layer perceptrons (MLP), XGBoost classifiers, LightGBM classifiers, among other possibilities. ML models may be combined by using models in sequence or through ensemble models such as by stacking models or through selection of champion models. In some examples, the classification model 210 can be developed and trained by external research organizations and provided via a suitable API independently or in combination with another ML model.

[0075] At block 530, the image analysis software 116 generates, using a second ML model, an embedded representation for a first pathological structure from the one or more pathological structures 215A... N. For example, the second machine learning model may be foundation model 220 trained to learn representations of the underlying data, such as the embedded representation of the identified pathological structures. For example, the foundation model 220 may be an encoder or autoencoder including one or more CNN layers. The foundation model 220 can thus provide feature extraction for use by downstream ML models. For example, for an implementation of the foundation model 220 including one or more CNN layers, the CNN layers can be fine-tuned to capture particularpatterns in the identified pathological structures that may be associated with particular outcomes of interest.

[0076] At block 540, the image analysis software 116 determines, using a third ML model, a prediction relating to an outcome of applying a therapy. For example, the third ML model may be a DeepMIL model 230 trained to predict the labels of “bags” or encoded subsets of the pathological structures as containing indications that the therapy may be effective. Some ML models utilizing MIL may include attention mechanisms or gated attention mechanisms that can, for example, adjust ML model parameters or weights during training according to a particular bag’s contribution to the bag's label. An example of such an implementation is described with respect to Figure 3 above.

[0077] At block 550, responsive to the prediction exceeding a predetermined threshold, the image analysis software 116 outputs an indication corresponding to a likelihood 235 of success of the therapy. The DeepMIL model 230 may, as an intermediate step during processing, generate predictions for each instance within a bag, which can later be used for interpreting the prediction. The intermediate results of DeepMIL model 230 can be aggregated using an aggregation function, sometimes referred to as a pooling operation, to combine the instance-level features into a single representation. The aggregated representation can then be processed using, for example, fully connected layers to output a final prediction or likelihood for the entire bag. This prediction can be used to assign a class label corresponding to a classification. For instance, one example label may indicate that a particular patient is or is not a good candidate for immunotherapy.

[0078] In some examples, given a classification or the predicted likelihood of success exceeding a predetermined threshold, the therapy may then be applied to a patient by a healthcare provider. The classification or the predicted likelihood may be used along with additional indicators of the effectiveness of the outcome of applying the therapy to the patient. For example, an oncologist may apply immunotherapy when, among other considerations, the likelihood of success is output as greater than 70%. The other considerations may include, for example, the presence of PD-L1 proteins, tumor proportion score, type and stage of cancer, genetic mutations, overall health and comorbidities, previous cancer treatments, other cancer biomarkers, immune system function, potential side effects, patient preferences, and so on.

[0079] Referring now to Figure 6. Figure 6 shows an example computing device 600 suitable for use in example systems or methods for predicting immunotherapy outcomes using deep learning according to this disclosure. The example computing device 600 includes a processor 610 which is in communication with the memory 620 and othercomponents of the computing device 600 using one or more communications buses 602. The processor 610 is configured to execute processor-executable instructions stored in the memory 620 to perform one or more methods for training or executing ML models or predicting immunotherapy outcomes according to different examples, such as part or all of the example method 500 described above with respect to Figure 5. In this example, the memory 620 may include the image analysis software 116, such as is depicted in the example system 100 shown in Figure 1. In addition, the computing device 600 also includes one or more user input devices 650, such as a keyboard, mouse, touchscreen, microphone, etc., to accept user input; however, in some examples, the computing device 600 may lack such user input devices, such as remote servers or cloud servers. The computing device 600 also includes a display 640 to provide visual output to a user.

[0080] The computing device 600 also includes a communications interface 630. In some examples, the communications interface 630 may enable communications using one or more networks, including a local area network (“LAN”); wide area network (“WAN”), such as the Internet; metropolitan area network (“MAN”); point-to-point or peer-to-peer connection; etc. Communication with other devices may be accomplished using any suitable networking protocol. For example, one suitable networking protocol may include the Internet Protocol (“IP”), Transmission Control Protocol (“TCP”), User Datagram Protocol (“UDP”), or combinations thereof, such as TCP / IP or UDP / IP.

[0081] While some examples of methods and systems herein are described in terms of software executing on various machines, the methods and systems may also be implemented as specifically configured hardware, such as field-programmable gate array (FPGA) specifically to execute the various methods according to this disclosure. For example, examples can be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or in a combination thereof. In one example, a device may include a processor or processors. The processor comprises a computer-readable medium, such as a random-access memory (RAM) coupled to the processor. The processor executes computer-executable program instructions stored in memory, such as executing one or more computer programs. Such processors may comprise a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), field programmable gate arrays (FPGAs). and state machines. Such processors may further comprise programmable electronic devices such as PLCs, programmable interrupt controllers (PICs), programmable logic devices (PLDs). programmable read-only memories (PROMs), electronically programmable read-only memories (EPROMs or EEPROMs), or other similar devices.

[0082] Such processors may comprise, or may be in communication with, media, for example one or more non -transitory computer-readable media, that may store processorexecutable instructions that, when executed by the processor, can cause the processor to perform methods according to this disclosure as carried out, or assisted, by a processor. Examples of non-transitory computer-readable medium may include, but are not limited to, an electronic, optical, magnetic, or other storage device capable of providing a processor, such as the processor in a web server, with processor-executable instructions. Other examples of non-transitory computer-readable media include, but are not limited to. a floppy disk, CD-ROM. magnetic disk, memory chip. ROM, RAM. ASIC, configured processor, all optical media, all magnetic tape or other magnetic media, or any other medium from which a computer processor can read. The processor, and the processing, described may be in one or more structures, and may be dispersed through one or more structures. The processor may comprise code to carry out methods (or parts of methods) according to this disclosure.

[0083] Embodiments may be implemented by using a computer program product, comprising computer program / instructions which, when executed by a processor, cause the processor to perform any of the methods described in the disclosure.

[0084] The foregoing description of some examples has been presented only for the purpose of illustration and description and is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. Numerous modifications and adaptations thereof will be apparent to those skilled in the art without departing from the spirit and scope of the disclosure.

[0085] Reference herein to an example or implementation means that a particular feature, structure, operation, or other characteristic described in connection with the example may be included in at least one implementation of the disclosure. The disclosure is not restricted to the particular examples or implementations described as such. The appearance of the phrases “in one example,’ “in an example,” “in one implementation,” or “in an implementation,” or variations of the same in various places in the specification does not necessarily refer to the same example or implementation. Any particular feature, structure, operation, or other characteristic described in this specification in relation to one example or implementation may be combined with other features, structures, operations, or other characteristics described in respect of any other example or implementation.

[0086] Use herein of the word “or” is intended to cover inclusive and exclusive OR conditions. In other words, A or B or C includes any or all of the following alternativecombinations as appropriate for a particular usage: A alone: B alone; C alone; A and B only: A and C only: B and C only; and A and B and C.

Claims

CLAIMSThat which is claimed is:

1. A method, comprising: receiving an image of a tissue sample; classifying, using a first machine learning model, one or more pathological structures in the tissue sample: generating, using a second machine learning model, an embedded representation for a first pathological structure from the one or more pathological structures; determining, using a third machine learning model, a prediction relating to an outcome of applying a therapy; and responsive to the prediction exceeding a predetermined threshold, outputting an indication corresponding to a likelihood of success of the therapy.

2. The method of claim 1, wherein the tissue sample is stained.

3. The method of any one of claims 1 or 2, wherein the tissue sample is stained using hematoxylin and eosin.

4. The method of any one of claims 1-3. wherein the first machine learning model is a neural network trained to identify regions of images comprising at least one pathological structure.

5. The method of claim 4, wherein the first machine learning model comprises a convolutional neural network.

6. The method of any one of claims 1-5, wherein the second machine learning model is an encoder, wherein the encoder is trained to generate the embedded representation.

7. The method of claim 6, wherein the encoder is trained using self- supervised training methods.

8. The method of claim 6, wherein the encoder is part of an autoencoder.

9. The method of claim 6, wherein the embedded representation is a vector, the vector comprising a numerical representation of the image of the tissue sample.

10. The method of any one of claims 1-9, wherein: the third machine learning model is deep neural network comprising an attention mechanism; and the third machine learning model is trained to predict the outcome of applying the therapy to a patient, comprising using a multiple instance learning technique.

11. The method of any one of claims 1-10, wherein the first machine learning model, the second machine learning model, and the third machine learning model are the same machine learning model.

12. The method of claim 10, wherein: the embedded representation comprises at least one pathological structure; and using the multiple instance learning technique comprises labeling the embedded representation based on a known outcome.

13. The method of any one of claims 1-12. wherein the one or more pathological structures in the tissue sample include tumor patches.

14. The method of claim 13, wherein the tumor patches are caused by non-small cell lung cancer.

15. The method of any one of claims 1-14, wherein the therapy is immunotherapy.

16. The method of claim 15, wherein the immunotherapy is anti-PD-l / PD-Ll immunotherapy.

17. The method of any one of claims 1-16, further comprising: responsive to the prediction exceeding the predetermined threshold, applying the therapy to a patient.

18. The method of claim 17, wherein applying the therapy to the patient is further responsive to an additional indicator of an effectiveness of the outcome of applying the therapy to the patient.

19. A system comprising: a non-transitory computer-readable medium: one or more processors in communication with the non-transitory computer-readable medium, the one or more processors configured to execute processor-executable instructions stored in the non-transitory computer-readable medium configured to cause the one or more processors to: receive an image of a tissue sample; classify, using a first machine learning model, one or more pathological structures in the tissue sample; generate, using a second machine learning model, an embedded representation for a first pathological structure from the one or more pathological structures; determine, using a third machine learning model, a prediction relating to an outcome of applying a therapy; and responsive to the prediction exceeding a predetermined threshold, output an indication corresponding to a likelihood of success of the therapy.

20. A non-transitory computer-readable medium comprising processor-executable instructions configured to cause one or more processors to: receive an image of a tissue sample; classify, using a first machine learning model, one or more pathological structures in the tissue sample; generate, using a second machine learning model, an embedded representation for a first pathological structure from the one or more pathological structures; determine, using a third machine learning model, a prediction relating to an outcome of applying a therapy; and responsive to the prediction exceeding a predetermined threshold, output an indication corresponding to a likelihood of success of the therapy.

21. An apparatus, comprising: means for implementing the operations of the method of any of claims 1-17.

22. A computer program product comprising computer instructions that, when executed by a processor, implement the operations of the method of any of claims 1-17.

Citation Information

Patent Citations

  • A Deep Learning Method For Predicting Patient Response To A Therapy

    US20200184641A1

  • Multiple instance learner for tissue image classification

    US20220237788A1

  • Systems and methods for generating histology image training datasets for machine learning models

    US20230343074A1