Systems and methods for multimodal prediction of patient outcomes

A multimodal deep learning system integrates various patient data types to enhance NSCLC treatment prediction accuracy by fusing histological, genetic, clinical, and imaging data, addressing the inaccuracy of current biomarkers and improving treatment selection.

JP2025533816APending Publication Date: 2025-10-09VENTANA MEDICAL SYSTEMS INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025519114
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-08-18
Filing Date
2023-10-02
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Current biomarkers for selecting patients for targeted therapy or immunotherapy in non-small cell lung cancer (NSCLC) are inaccurate, leading to potential misidentification of patients who can benefit from these treatments.

Method used

A multimodal prediction system using a deep learning framework that integrates histological H&E image data, gene sequencing data, clinical data, MRI data, and IHC image data to generate a survivability score for NSCLC patients, leveraging convolutional and feedforward neural networks to fuse these modalities and produce a personalized prediction.

Benefits of technology

The system provides more accurate patient survivability predictions than existing methods, enabling better treatment selection by considering diverse patient data modalities and improving treatment efficacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025533816000001_ABST
    Figure 2025533816000001_ABST
Patent Text Reader

Abstract

A method for predicting a patient's overall survivability by a prediction system based on machine learning includes receiving, by a prediction system having a processor and a memory, a plurality of input modalities corresponding to the patient, the input modalities being of different types from each other; generating, by the prediction system, a plurality of intermediate features based on the plurality of input modalities, each input modality of the plurality of input modalities corresponding to one or more features of the plurality of intermediate features; and determining, by the prediction system, a survivability score corresponding to the patient's overall survivability based on a fusion of the plurality of intermediate features.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation of U.S. Provisional Application No. 63 / 378,164, filed on October 3, 2022, entitled "SYSTEM AND METHOD FOR MULTIMODAL PREDICTION OF PATIENT OUTCOMES," U.S. Provisional Application No. 63 / 533,572, filed on August 18, 2023, entitled "ATTENTION-BASED MULTIMODAL-FUSION FOR NON-SMALL CELL LUNG CANCER (NSCLC) PATIENT SURVIVAL PREDICTION," and Indian Application No. 202311030916, filed on April 30, 2023, entitled "INTEGRATING MULTIMODAL DATA FOR NON-SMALL CELL LUNG CANCER (NSCLC) PATIENT SURVIVAL PREDICTION." "PREDICTION" and claims priority to and benefit of Indian Application No. 202311030917, filed April 30, 2023, entitled "FEATURE GENERATION AND SELECTION FOR AN EFFICIENT MULTIMODAL ANALYSIS ON BIOLOGICAL DATA" and Indian Application No. 202311044011, filed June 30, 2023, entitled "INTERPRETABLE FEATURE BASED NETWORK FOR CLASSIFYING CELL-OF-ORIGIN FROM WHOLE SLIDE IMAGES IN DIFFUSE LARGE B-CELL LYMPHOMA PATIENTS."

[0002] One or more aspects of some embodiments according to the present disclosure relate to systems and methods for predicting patient outcomes. [Background technology]

[0003] Various forms of cancer have become a leading cause of death worldwide. Lung cancer, in particular, is one of the most common malignancies, accounting for approximately 25% of all cancer-related deaths. Approximately 84% of lung cancers are non-small cell lung cancer (NSCLC), a group of similarly behaving lung cancers. Immunotherapy with checkpoint inhibitors, such as anti-PD1 and anti-PD-L1 drugs, has shown promising clinical outcomes for patients with locally advanced (ad) or metastatic (m) NSCLC. However, biomarkers currently used to select patients who can benefit from targeted therapy or immunotherapy are inaccurate and have significant potential for improvement.

[0004] The above information disclosed in this Background section is intended only to enhance understanding of the background art, and therefore, the information described in this Background section does not necessarily constitute prior art. Summary of the Invention

[0005] Aspects of embodiments of the present disclosure are directed to a multimodal prediction system that utilizes a deep learning framework to predict overall survival of patients (e.g., NSCLC patients) from diverse multimodal data. In some embodiments, the deep learning framework leverages digital pathology, genetic information, clinical data, patient demographic data, as well as many other modalities to generate more accurate predictions than other prediction methods in the related art.

[0006] According to some embodiments of the present disclosure, there is provided a method of predicting a patient's overall survivability by a machine learning based prediction system, the method including: receiving, by a prediction system including a processor and a memory, a plurality of input modalities corresponding to the patient, the input modalities being of different types from each other; generating, by the prediction system, a plurality of intermediate features based on the plurality of input modalities, each input modality of the plurality of input modalities corresponding to one or more features of the plurality of intermediate features; and determining, by the prediction system, a survivability score corresponding to the patient's overall survivability based on a fusion of the plurality of intermediate features.

[0007] In some embodiments, the multiple input modalities include a first modality comprising histological hematoxylin and eosin (H&E) image data, a second modality comprising gene sequencing data, a third modality comprising clinical data, a fourth modality comprising radiological magnetic resonance imaging (MRI) data, and a fifth modality comprising immunohistochemistry (IHC) image data.

[0008] In some embodiments, the histological H&E image data comprises digitized images of patient tissue samples stained with hematoxylin and eosin stain.

[0009] In some embodiments, the digitized image comprises multiple images of different tumor regions of the tissue sample.

[0010] In some embodiments, the gene sequencing data comprises mRNA gene expression extracted from the patient's tumor tissue.

[0011] In some embodiments, the clinical data includes the patient's age, the patient's sex, the patient's tumor stage, and the patient's performance status.

[0012] In some embodiments, the IHC image data comprises digitized images of patient tissue samples stained with the PD-L1 biomarker.

[0013] In some embodiments, the prediction system includes a first convolutional neural network configured to receive a first modality of the plurality of input modalities and generate one or more first intermediate features of the plurality of intermediate features; a first feedforward neural network configured to receive a second modality of the plurality of input modalities and generate one or more second intermediate features of the plurality of intermediate features; a second feedforward network configured to receive a third modality of the plurality of input modalities and generate one or more third intermediate features of the plurality of intermediate features; a second convolutional neural network configured to receive a fourth modality of the plurality of input modalities and generate one or more fourth intermediate features of the plurality of intermediate features; and a third convolutional neural network configured to receive a fifth modality of the plurality of input modalities and generate one or more fifth intermediate features of the plurality of intermediate features.

[0014] In some embodiments, the method further includes combining the first intermediate feature, the second intermediate feature, the third intermediate feature, and the fourth intermediate feature to form a fusion of the multiple intermediate features.

[0015] In some embodiments, the prediction system further includes a fusion layer neural network configured to receive a fusion of the plurality of intermediate features and generate a survivability score.

[0016] In some embodiments, the method further includes receiving, by the classifier, an input image from the patient tissue sample, and extracting, by the classifier, cell spatial graph data from the input image, the cell spatial graph data including a cell type and a location of each cell in the input image.

[0017] In some embodiments, the input image comprises one of a histological H&E image and an IHC image.

[0018] In some embodiments, extracting the cell spatial graph data comprises detecting cells within the input image, generating cell classification data for the cells detected within the input image, the cell classification data comprising a cell type and a location of each cell in the input image, and constructing the cell spatial graph data based on the cell classification data.

[0019] In some embodiments, the prediction system further includes a graph convolutional network configured to receive a sixth modality of the plurality of input modalities and generate one or more sixth intermediate features of the plurality of intermediate features, wherein the sixth modality includes a cell spatial graph corresponding to a histological hematoxylin and eosin (H&E) image or an immunohistochemistry (IHC) image from a tissue sample of the patient.

[0020] In some embodiments, the prediction system includes a multimodal fusion model configured to correlate multiple input modalities into a survivability score.

[0021] In some embodiments, the method further comprises transmitting the survivability score to a display device for display to a user.

[0022] According to some embodiments of the present disclosure, there is provided a method for predicting a patient's overall survivability by a machine learning-based prediction system, the method including: receiving, by the prediction system, a plurality of input modalities corresponding to the patient, the input modalities including a first modality including histological hematoxylin and eosin (H&E) image data, a second modality including gene sequencing data, a third modality including clinical data, a fourth modality including radiological magnetic resonance imaging (MRI) data, and a fifth modality including immunohistochemistry (IHC) image data; generating, by the prediction system, a plurality of intermediate features based on the plurality of input modalities, each input modality of the plurality of input modalities corresponding to one or more features of the plurality of intermediate features; and determining, by the prediction system, a survivability score corresponding to the patient's overall survivability based on a fusion of the plurality of intermediate features.

[0023] In some embodiments, the prediction system includes a first convolutional neural network configured to receive a first modality and generate one or more first intermediate features of the plurality of intermediate features, a first feedforward neural network configured to receive a second modality and generate one or more second intermediate features of the plurality of intermediate features, a second feedforward neural network configured to receive a third modality and generate one or more third intermediate features of the plurality of intermediate features, a second convolutional neural network configured to receive a fourth modality and generate one or more fourth intermediate features of the plurality of intermediate features, and a third convolutional neural network configured to receive a fifth modality and generate one or more fifth intermediate features of the plurality of intermediate features.

[0024] In some embodiments, the method further includes receiving, by a classifier of the prediction system, an input image from a tissue sample of a patient, and extracting, by the classifier, cell spatial graph data from the input image, the cell spatial graph data including a cell type and location of each cell in the input image, wherein the input image includes one of a histological H&E image and an IHC image.

[0025] According to some embodiments of the present disclosure, there is provided a prediction system including: a first convolutional neural network configured to receive histological hematoxylin and eosin (H&E) images and generate one or more first intermediate features; a first feed-forward neural network configured to receive gene sequencing data and generate one or more second intermediate features; a second feed-forward neural network configured to receive clinical data and generate one or more third intermediate features; a second convolutional neural network configured to receive radiological magnetic resonance imaging (MRI) data and generate one or more fourth intermediate features; a fusion circuit configured to combine the first, second, third, and fourth intermediate features to form a fusion of intermediate features via attention-gated tensor fusion; and a fusion layer neural network configured to receive the fusion of intermediate features and generate a survivability score corresponding to an overall survivability of the patient. [Brief explanation of the drawings]

[0026] Non-limiting and non-exhaustive embodiments according to the present disclosure are described with reference to the following figures, in which like reference numerals refer to like parts throughout the various views unless otherwise specified.

[0027] [Figure 1] FIG. 1 is a flow diagram illustrating various actions that may occur in a pathological situation or environment, according to some embodiments.

[0028] [Figure 2] FIG. 1 is a block diagram illustrating a forecasting system, according to some embodiments of the present disclosure.

[0029] [Figure 3] FIG. 1 is a block diagram illustrating a prediction system utilizing derivative modalities, according to some embodiments of the present disclosure.

[0030] [Figure 4] FIG. 1 is a block diagram illustrating the internal architecture of a prediction system, according to some embodiments of the present disclosure.

[0031] [Figure 5] FIG. 1 is a flow diagram illustrating a process for predicting a patient's overall chance of survival by a machine learning-based prediction system, according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0032] Aspects of some exemplary embodiments will now be described in more detail with reference to the accompanying drawings, in which like reference numerals refer to like elements throughout. However, the present invention may be embodied in a variety of different forms and should not be construed as limited to only the embodiments illustrated herein. Rather, these embodiments are provided as examples so that this disclosure will be thorough and complete, and will fully convey the aspects and features of the present invention to those skilled in the art. Accordingly, processes, elements, and techniques that are not necessary for those skilled in the art to fully understand the aspects and features of the present invention may not be described. Unless otherwise noted, like reference numerals indicate like elements throughout the accompanying drawings and written specification, and therefore, descriptions thereof will not be repeated. In the drawings, relative sizes of elements, layers, and regions may be exaggerated for clarity.

[0033] Pathology is a medical field that attempts to facilitate the diagnosis and treatment of disease by studying samples of a patient's tissues, cells, and bodily fluids. In many applications, tissue samples are collected from a patient and processed into a form that can be analyzed by a physician (e.g., a pathologist), often under magnification, to diagnose and characterize associated disease states based on the tissue sample.

[0034] 1 is a flow diagram illustrating various operations that may occur in a pathology environment or system 100. For example, if a treating physician or healthcare provider identifies a patient for whom analysis of a tissue or bodily fluid sample may be beneficial in diagnosing or treating a medical condition, the tissue or bodily fluid sample may be collected in operation 102. Patient identification information may be collected and matched to the patient's sample, and the sample may be placed in a sterile container and / or collection medium for further processing.

[0035] The sample may then be transported to a pathology receiving laboratory in operation 104 where it may be received, sorted, organized, and labeled along with other samples from other patients for further processing.

[0036] In operation 106, the sample may be further processed as part of a grossing operation. For example, individual tissue samples or specimens may be sectioned for embedding and subsequent cutting into a collection onto slides.

[0037] The sample or specimen may then be mounted or deposited on one or more glass slides in operation 108. Preparing the slides may include applying one or more reagents or stains to the sample, for example, to improve the visibility of or contrast between different portions of the sample.

[0038] In some cases, during or after the reagent or stain treatment, multiple slides may be organized or collected into a case or folio in operation 110. The case may be carefully labeled with, for example, individual patient identification information.

[0039] Between each of operations 102 and 110, in operation 112, the sample, specimen, or slide may be transported within or between medical facilities (e.g., between a doctor's office and a laboratory) or may be stored between processing operations.

[0040] Once sample and slide processing is complete and the pathologist is ready to review the sample, the slide and / or a case holding multiple slides corresponding to a patient may be transported back to the pathologist in operation 112. In operation 114, the pathologist may review the slides, for example, under magnification using a microscope. Individual slides may be placed under the microscope objective so that the microscope and slides can be manipulated and adjusted as the pathologist reviews the tissue or bodily fluid.

[0041] Once the pathologist has completed review of the slide, the pathologist may attempt to form a medical opinion or make a diagnosis in operation 116. Meanwhile, the sample or slide may be transported again to a storage facility for a longer period of time in operation 112. In some cases, the sample or slide may be transported again to another physician before or after a storage period for further analysis, a second opinion, etc.

[0042] An example of the above operations may be performed in a pathology setting, where a pathologist analyzes certain information and arrives at a score (e.g., a viability score) that indicates the success of the treatment, thereby identifying patients (e.g., breast cancer patients) who are likely to respond to a particular immunotherapy. Some related scoring solutions are single-modality systems that use a single biomarker (e.g., PD-L1) or oligobiomarkers. However, in many cases, a single biomarker is not an accurate predictor of the effectiveness of an immunotherapy. For example, in the case of the PD-L1 biomarker, some negative patients may respond to the treatment, while some positive patients may not. Furthermore, some treatments do not yet have a discovered biomarker that can serve as a predictor of the treatment's effectiveness.

[0043] Accordingly, some aspects of the present disclosure are directed to deep fusion, multi-modality predictive systems that can consider a plethora of heterogeneous data about a patient, such as demographic information, the patient's tumor histology, biomarker-stained slides, disease stage (e.g., cancer stage), patient performance data, etc., to derive an individualized / personalized patient score based on the patient's profile. The score may determine how well the patient will respond to a particular treatment.

[0044] FIG. 2 is a block diagram illustrating a prediction system 200 according to some embodiments of the present disclosure.

[0045] According to some embodiments, the prediction system 200 is a multi-modality deep learning framework that integrates various modality data to predict patient outcomes in a more robust manner than related art qualitative clinical assessments or unimodal strategies. In some embodiments, the prediction system 200 is configured to receive multiple input modalities 202 associated with a patient and determine a patient survivability score based on the modalities 202. The various modalities utilized by the prediction system 200 may be of different types.

[0046] For example, the first modality 202a may include histological hematoxylin and eosin (H&E) image data. The H&E data 202a may include one or more digitized images of a patient's tissue sample (e.g., a tumor tissue sample) stained with hematoxylin and eosin dye. The H&E dye stains cell nuclei, extracellular matrix, and cytoplasm, as well as other cellular structures, with different colors, thereby allowing the pathologist and the prediction system 200 to distinguish between different cellular structures. The overall pattern of coloration from the staining also indicates the general layout and distribution of cells, providing a view of the structure of the tissue sample. In some examples, the H&E image data 202a may include multiple image patches (e.g., three image patches) extracted (e.g., randomly selected and extracted) from a viable tumor region of the stained tissue sample.

[0047] The second modality 202b may include genetic sequencing data, such as DNA information and / or mRNA gene expression of tumor mutations extracted from the patient's tumor tissue. Each tumor cell may have hundreds or thousands of tumor mutation genes. The second modality 202b may include some or all of the genetic mutations discovered in the tissue sample. In some examples, only those expressions most relevant to the patient's chance of survival may be included in the second modality 202b.

[0048] The third modality 202c may include clinical data associated with the patient, such as the patient's age, gender, tumor stage, and performance status. Performance status may be measured by an ECOG score, which describes the patient's level of function in terms of the patient's ability to care for themselves, daily activities, and physical abilities (e.g., walking, working, etc.). Performance status may also be the Karnofsky Performance Score, which measures the cancer patient's ability to perform usual tasks. Scores may range from 0 to 100, with higher scores indicating the patient is better able to perform daily activities.

[0049] The fourth modality 202d may include radiological magnetic resonance imaging (MRI) data, which may include one or more digitized MRI images of the patient's tumor region (such as Gd-T1w and T2w-FLAIR scans). The MRI image data may be useful for assessing tumor volume.

[0050] The fifth modality 202e may include immunohistochemistry (IHC) image data. The IHC data 202e may include one or more digitized images of a patient's tissue sample (e.g., a tumor tissue sample) stained with the PD-L1 biomarker. The PD-L1 biomarker may produce a brown stain when antibodies are able to attach to those tumor cells that have PD-L1 expression. The IHC image may correspond to a slice of the tissue sample adjacent to the slice on which the H&E image 202a is based. In some examples, the cellular structure captured in the H&E and IHC images may be the same or substantially the same, but this is not necessarily the case. For example, an H&E image may provide information about the pattern, shape, and structure of cells in a tissue sample, while an IHC image showing the distribution and localization of specific proteins in the sample may not clearly reveal the cellular structure of the tissue sample.

[0051] Once the prediction system 200 estimates the survivability score 204, the score may be transmitted to a server (e.g., a remote server or a cloud server) 206 for further processing and / or to a display device 208 for display to a user.

[0052] Although the above description describes five modalities as examples of input modalities to the prediction system 200, embodiments of the present disclosure are not so limited, and any suitable type and / or number of modalities may be employed by the prediction system 200 to determine the viability score 204.

[0053] For example, the prediction system 200 may utilize one or more derivatives of multiple modalities 202 .

[0054] FIG. 3 is a block diagram illustrating a prediction system 200 utilizing derivative modalities, according to some embodiments of the present disclosure.

[0055] According to some embodiments, the cellular structure captured by the H&E image data 202a or IHC image data 202e is used to generate cell graph data that identifies the location and cell type of each cell within the image. Generally, the graph may represent the spatial arrangement and neighborhood relationships of different tissue components, which may be visually recognized by a pathologist during specimen examination. Depending on the approach employed (e.g., Voronoi diagram, Delaunay triangulation, nearest neighbor graph, etc.), the nodes and edges of the graph may represent different elements or characteristics. For example, each node of the graph may identify a cell (e.g., the center of a cell's nucleus), and each edge may identify the Euclidean distance between adjacent cells or represent the similarity between adjacent cells. The prediction system 200 may utilize the cell graph data to extract features that can be used for survival outcome prediction. Cell classification may be performed manually by a pathologist or by a trained classifier.

[0056] In some embodiments, the classifier (e.g., machine learning classifier) ​​230a / b receives an input image that is an image of a stained tissue sample (e.g., an H&E image 202a or an IHC image 202e), detects cells within the input image, and generates cell classification data corresponding to the detected cells. The classification data, including the type and location of each cell in the input image, can be used to generate cell graph data 203a / e.

[0057] In some embodiments, the classifier 230a / b includes a neural network (e.g., a convolutional neural network) capable of cell detection and cell classification. The neural network may include several layers, each of which performs a convolution operation on an input feature map (IFM) via the application of a kernel / filter to generate an output feature map that serves as the input feature map for subsequent layers. In the first layer of the neural network, the input feature map may be an input image (e.g., an H&E image 202a or an IHC image 202e). The neural network may be a convolutional neural network (ConvNet / CNN), which takes in an input image and assigns importance to various aspects / objects of the image (e.g., via learnable weights and biases) to distinguish one from another. However, embodiments of the present disclosure are not limited thereto. For example, the neural network may be a recurrent neural network (RNN) with convolution operations, a random forest network, or the like. In some embodiments, the neural network further generates a spatial cell graph 203a / e based on the classification data.

[0058] As shown in FIG. 3 , in some embodiments, a first classifier 230a generates first cell graph data 203a based on H&E image data 202a, and a second classifier 230b generates second cell graph data 203b based on IHC image data 202e. However, embodiments of the present disclosure are not limited thereto. For example, the first and second cell graph data 203a and 203b may be generated by the same classifier 230a / b, as opposed to two separate classifiers. Furthermore, in some embodiments, only one of the first and second cell graph data 203a and 203b may be generated and utilized by the prediction system 200. While the classifiers 230a and 230b are shown as external to the prediction system 200, embodiments of the present disclosure are not limited thereto, and one or more of the first and second classifiers 230a and 230b may be included in (e.g., be part of) the prediction system 200.

[0059] The prediction system may include a multimodal fusion model (e.g., a multimodal deep orthogonal fusion (DOF) model) configured to correlate multiple input modalities to a viability score. In some embodiments, the prediction system 200 includes multiple neural networks (e.g., multiple unimodal neural networks) for processing the multiple input modalities 202. Each modality may have a corresponding neural network of an appropriate type.

[0060] FIG. 4 is a block diagram illustrating the internal architecture of a prediction system 200, according to some embodiments of the present disclosure.

[0061] According to some embodiments, the prediction system 200 includes multiple neural networks (also referred to as unimodal networks / sub-models or feature extraction models) corresponding to multiple modalities (e.g., having a one-to-one correspondence). Each neural network generates one or more intermediate features (IFs; also referred to as data modality feature representations or modality feature vectors) corresponding to an input modality. The intermediate features from the various neural networks are fused and utilized by the prediction system 200 to generate the viability score 204. In some examples, the intermediate features are self-identified by the prediction system 200 during training. However, embodiments of the present disclosure are not limited thereto, and one or more intermediate features may be manually selected by a user of the prediction system 200 during training of the prediction system 200.

[0062] In some embodiments, the prediction system 200 includes a first convolutional neural network (e.g., U-Net) 210 for processing the H&E image data of a first modality 202a to generate one or more corresponding first intermediate data (e.g., histological features; IF1) 220; a first feedforward network 212 for processing the genetic data of a second modality 202b to generate one or more corresponding second intermediate data (e.g., genetic features; IF2) 222; and a first feedforward network 213 for processing the genetic data of a second modality 202b to generate one or more corresponding third intermediate data (e.g., clinical features; IF3) 224. The input modalities 202 may include a second feedforward network (or fully connected (FC) network, etc.) 214 for processing the clinical data of the third modality 202c to generate one or more corresponding fourth intermediate data (e.g., MRI features; IF4) 226; a second convolutional neural network 216 for processing the MRI image data of the fourth modality 202d to generate one or more corresponding fifth intermediate data (e.g., IHC features; IF5) 228; and a third convolutional neural network 218 for processing the IHC image data of the fifth modality 202e to generate one or more corresponding fifth intermediate data (e.g., IHC features; IF5) 228. In some examples, one or more of the input modalities 202 may be combined and fed into the same network. For example, the genetic data 202b and the clinical data 202c may be combined before being fed into a single feedforward network, FC network, etc.

[0063] In some embodiments, prediction system 200 includes a graph convolutional neural network 219 for processing spatial cell graph data of a sixth modality 203 (e.g., 203a / b), which may correspond to (e.g., based on or extracted from) H&E image data or IHC image data, to generate corresponding one or more sixth intermediate data (e.g., graph features; IF6) 229. In embodiments in which cell graph data is extracted from two or more images (e.g., both H&E and IHC images), one or more additional graph convolutional networks may be used by prediction system 200 to generate additional intermediate features.

[0064] Thus, each input data modality 202 is processed by machine learning (e.g., a dedicated deep learning submodel / unimodel) 210 / ... / 219, which is trained to generate intermediate data (e.g., modality-specific feature representations or feature representation vectors) 220 / ... / 229.

[0065] In some embodiments, the prediction system 200 also includes a fusion network 240 that combines the generated feature representation vectors into a single fused representation (e.g., a single fused vector) and outputs a survivability score (also referred to as a fused prognostic risk score) using the fused representation vector. The fusion network 240 may include a data fusion block (e.g., a fusion circuit) 242 and a fusion layer neural network 244. The fusion block 242 is configured to combine (e.g., fuse) various intermediate features (e.g., IF1-IF6) generated by multiple unimodal neural networks (e.g., 210, 212, 214, 216, 218, and 219), for example, via attention-gated tensor fusion, which partially controls the representability of each modality 202. The fusion layer neural network 244 is configured to generate a survivability score based on the fusion of multiple intermediate features, for example, by utilizing a Cox partial likelihood loss function. The fusion layer neural network 244 may include an integrated model of pairwise feature interactions across modalities. In some examples, the fusion layer neural network 244 may be an FC network (composed of several fully connected layers), or the like.

[0066] The neural networks that make up the prediction system 200 can be trained in one shot through a single end-to-end training session using a large training dataset, or different neural networks corresponding to modalities can be trained separately and then combined together to form one large system, i.e., the prediction system 200. For example, the unimodal network 202 can generally be first trained separately for survival prediction. However, one or more of the image modality feature extraction models 202a, 202d, and 202e can be pre-trained models trained on a separate public dataset (e.g., ImageNet) for a different task. This can be done to leverage non-medical data to learn richer feature representations. Then, during the multimodal network training phase, the fusion network (including the fusion block 242 and the fusion layer neural network 244) is trained using the features extracted by the trained unimodal network 202 in the first step as input. During multimodal network training, the unimodal network 202 can either be frozen, unfrozen, or frozen for the first few epochs and then unfrozen to allow for further optimization of unimodal feature extraction.

[0067] The holistic approach of the prediction system 200, which utilizes multiple modalities associated with a patient, allows the system 200 to generate a more accurate prediction score for each individual patient.

[0068] Table 1 compares the performance of prediction system 200 with that of other approaches based on the concordance index (CI), a standard performance measure for model evaluation in survival analysis. In Table 1, the baseline represents a statistical method used to predict overall survival, while the other models use deep learning predictions. As shown in the example in Table 1, despite the fact that prediction system 200's fusion method uses only three modalities, its predictive ability exceeds that of a unimodality method (histologic H&E) and a two-modality method (omics) (clinical and mRNA).

[0069] [Table 1]

[0070] FIG. 5 is a flow diagram illustrating a process 500 for predicting a patient's overall chance of survival with a machine learning-based prediction system, according to some embodiments of the present disclosure.

[0071] According to some embodiments, the prediction system 200 receives multiple input modalities 202 corresponding to a patient, the input modalities being of different types from one another (S502). The multiple input modalities may include one or more of a first modality 202a including histological hematoxylin and eosin (H&E) image data, a second modality 202b including gene sequencing data, a third modality 202c including clinical data, a fourth modality 202d including radiological magnetic resonance imaging (MRI) data, and a fifth modality 202e including immunohistochemistry (IHC) image data.

[0072] In some embodiments, the prediction system 200 then generates a plurality of intermediate features based on the plurality of input modalities, each input modality of the plurality of input modalities corresponding to one or more features of the plurality of intermediate features (S504). The prediction system 200 combines the intermediate features to form a fusion of the intermediate features via attention-gated tensor fusion.

[0073] The prediction system 200 determines a survivability score corresponding to the overall survivability of the patient based on the fusion of the intermediate features (S506).

[0074] As described above, according to some embodiments, the prediction system fuses heterogeneous multimodal data using a single deep learning framework for patient survival prediction. The prediction system enables better characterization of the interactions between different modalities, allowing clinicians to gain more insights from a wealth of clinical and diagnostic information and build integrated, personalized healthcare solutions. Furthermore, the prediction system produces a more accurate predictive score for an individual patient than single or oligoparameters such as demographics, histology, PD-L1 status, and omics. The score can be used to select the optimal treatment regimen for that patient.

[0075] According to various embodiments of the present disclosure, prediction system 200 is implemented using one or more processing or electronic circuits configured to perform various operations as described above. Types of electronic circuits may include central processing units (CPUs), graphics processing units (GPUs), artificial intelligence (AI) accelerators (e.g., vector processors that may include vector arithmetic logic units configured to efficiently perform operations common to neural networks, such as dot products and softmaxes), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), digital signal processors (DSPs), etc. For example, in some situations, aspects of embodiments of the present disclosure are implemented in program instructions stored in non-volatile computer-readable memory that, when executed by an electronic circuit (e.g., a CPU, a GPU, an AI accelerator, or a combination thereof), perform the described operations. The operations performed by prediction system 200 may be performed by a single electronic circuit (e.g., a single CPU, a single GPU, etc.) or may be allocated among multiple electronic circuits (e.g., multiple GPUs, or a CPU in conjunction with a GPU). The multiple electronic circuits may be local to each other (e.g., located on the same die, located in the same package, or located in the same embedded device or computer system) and / or remote from each other (e.g., communicating over a network such as a local personal area network such as Bluetooth®, communicating over a local area network such as a local wired and / or wireless network, and / or communicating over a wide area network such as the Internet, where some operations are performed locally and other operations are performed on a server hosted by a cloud computing service). One or more electronic circuits operating to implement prediction system 200 may be referred to herein as a computer or computer system, which may include memory that stores instructions that, when executed by the one or more electronic circuits, implement the systems and methods described herein.

[0076] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the invention. As used herein, the singular forms "a" and "an" are intended to include the plural forms unless the context clearly suggests otherwise. It will be further understood that the terms "comprises," "comprising," "includes," and "including," when used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. The term "and / or," as used herein, includes any and all combinations of one or more of the associated listed items. Expressions such as "at least one of," when preceding a list of elements, modify the entire list of elements and not the individual elements of the list.

[0077] Terms such as "first," "second," and "third" may be used herein to describe various elements, components, regions, layers, and / or sections, but it is understood that these elements, components, regions, layers, and / or sections should not be limited by these terms. These terms are used to distinguish one element, component, region, layer, or section from another element, component, region, layer, or section. Thus, a first element, component, region, layer, or section described below may be referred to as a second element, component, region, layer, or section without departing from the spirit and scope of the present invention.

[0078] As used herein, the terms "substantially," "about," and similar terms are used as terms of approximation, not as terms of degree, and are intended to account for inherent variations in measurements or calculations recognized by those of ordinary skill in the art. Furthermore, the use of "may" when describing embodiments of the present invention refers to "one or more embodiments of the present invention." As used herein, the terms "use," "using," and "used" may be considered synonymous with the terms "utilize," "utilizing," and "utilized," respectively.

[0079] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art to which this invention belongs. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and / or this specification, and should not be interpreted in an idealized or overly formal sense unless explicitly defined herein.

[0080] While aspects of certain exemplary embodiments of the system and method for quantifying pathology slides using a cell-based scoring system have been described and illustrated herein, various modifications and variations may be implemented as understood by those skilled in the art without departing from the spirit and scope of the embodiments according to the present disclosure. Accordingly, it should be understood that the pathology slide manufacturing system and method according to the principles of the present disclosure may be embodied in other ways than those specifically described herein. The present disclosure is also defined in the following claims and their equivalents.

Claims

1. 1. A method for predicting overall survival chances of a patient by a machine learning based prediction system, said method comprising: receiving, by the prediction system comprising a processor and a memory, a plurality of input modalities corresponding to the patient, the input modalities being of different types from one another; generating, by the prediction system, a plurality of intermediate features based on the plurality of input modalities, each input modality of the plurality of input modalities corresponding to one or more features of the plurality of intermediate features; and determining, by the prediction system, a survivability score corresponding to an overall survivability probability of the patient based on a fusion of the plurality of intermediate features; A method comprising:

2. The plurality of input modalities include: a first modality comprising histological hematoxylin and eosin (H&E) image data; a second modality comprising genetic sequencing data; a third modality including clinical data; a fourth modality including magnetic resonance imaging (MRI) data; a fifth modality comprising immunohistochemistry (IHC) image data; The method of claim 1 , comprising:

3. The method of claim 2 , wherein the histological H&E image data comprises a digitized image of the patient's tissue sample stained with hematoxylin and eosin dye.

4. The method of claim 3 , wherein the digitized image comprises multiple images of different tumor regions of the tissue sample.

5. 3. The method of claim 2, wherein the gene sequencing data comprises mRNA gene expression extracted from the patient's tumor tissue.

6. the clinical data the age of the patient; the patient's sex, the stage of the patient's tumor, and the patient's performance status; The method of claim 2 , comprising:

7. 3. The method of claim 2, wherein the IHC image data comprises a digitized image of the patient's tissue sample stained with a PD-L1 biomarker.

8. The prediction system comprises: a first convolutional neural network configured to receive a first modality of the plurality of input modalities and to generate one or more first intermediate features of the plurality of intermediate features; a first feedforward neural network configured to receive a second of the plurality of input modalities and to generate one or more second intermediate features of the plurality of intermediate features; a second feedforward network configured to receive a third modality of the plurality of input modalities and to generate one or more third intermediate features of the plurality of intermediate features; a second convolutional neural network configured to receive a fourth modality of the plurality of input modalities and to generate one or more fourth intermediate features of the plurality of intermediate features; and a third convolutional neural network configured to receive a fifth modality of the plurality of input modalities and to generate one or more fifth intermediate features of the plurality of intermediate features; The method of claim 1 , comprising:

9. 9. The method of claim 8, further comprising combining the first intermediate feature, the second intermediate feature, the third intermediate feature, and the fourth intermediate feature to form the blend of the plurality of intermediate features.

10. The prediction system comprises: The method of claim 8 , further comprising a fusion layer neural network configured to receive the fusion of the plurality of intermediate features and generate the survivability score.

11. receiving, by a classifier, an input image from a tissue sample of the patient; and extracting, by the classifier, cell spatial graph data from the input image, the cell spatial graph data including a cell type and a location of each cell in the input image; The method of claim 1 further comprising:

12. The method of claim 11 , wherein the input image comprises one of a histological hematoxylin and eosin (H&E) image and an immunohistochemistry (IHC) image.

13. The extracting of the cell spatial graph data includes: detecting cells within the input image; generating cell classification data for the cells detected within the input image, the cell classification data including the cell type and the location of each cell in the input image; and constructing the cell spatial graph data based on the cell classification data; The method of claim 11 , comprising:

14. The prediction system comprises: a graph convolutional network configured to receive a sixth modality of the plurality of input modalities and to generate one or more sixth intermediate features of the plurality of intermediate features; 10. The method of claim 1, wherein the sixth modality comprises a cell spatial graph corresponding to a histological hematoxylin and eosin (H&E) image or an immunohistochemistry (IHC) image from the patient tissue sample.

15. The method of claim 1 , wherein the prediction system includes a multimodal fusion model configured to correlate the multiple input modalities to the survivability score.

16. The method of claim 1 , further comprising transmitting the survivability score to a display device for display to a user.

17. 1. A method for predicting overall survival chances of a patient by a machine learning based prediction system, said method comprising: receiving, by the prediction system, a plurality of input modalities corresponding to the patient, the input modalities comprising: a first modality comprising histological hematoxylin and eosin (H&E) image data; a second modality comprising genetic sequencing data; a third modality including clinical data; a fourth modality including magnetic resonance imaging (MRI) data; a fifth modality comprising immunohistochemistry (IHC) image data; receiving a plurality of input modalities, including generating, by the prediction system, a plurality of intermediate features based on the plurality of input modalities, each input modality of the plurality of input modalities corresponding to one or more features of the plurality of intermediate features; and determining, by the prediction system, a survivability score corresponding to an overall survivability probability of the patient based on a fusion of the plurality of intermediate features; A method comprising:

18. The prediction system comprises: a first convolutional neural network configured to receive the first modality and generate one or more first intermediate features of the plurality of intermediate features; a first feedforward neural network configured to receive the second modality and generate one or more second intermediate features of the plurality of intermediate features; a second feedforward neural network configured to receive the third modality and generate one or more third intermediate features of the plurality of intermediate features; a second convolutional neural network configured to receive the fourth modality and generate one or more fourth intermediate features of the plurality of intermediate features; and a third convolutional neural network configured to receive the fifth modality and generate one or more fifth intermediate features of the plurality of intermediate features; 20. The method of claim 17, comprising:

19. receiving an input image from a tissue sample of the patient by a classifier of the prediction system; and extracting, by the classifier, cell spatial graph data from the input image, the cell spatial graph data including a cell type and a location of each cell in the input image; further comprising The method of claim 17 , wherein the input image comprises one of the histological H&E image and the IHC image.

20. 1. A prediction system, comprising: a first convolutional neural network configured to receive a histological hematoxylin and eosin (H&E) image and generate one or more first intermediate features; a first feedforward neural network configured to receive the genetic sequencing data and generate one or more second intermediate features; a second feedforward neural network configured to receive the clinical data and generate one or more third intermediate features; a second convolutional neural network configured to receive the radiological magnetic resonance imaging (MRI) data and generate one or more fourth intermediate features; a fusion circuit configured to combine the first intermediate feature, the second intermediate feature, the third intermediate feature, and the fourth intermediate feature to form a fusion of intermediate features via attention-gated tensor fusion; and a fusion layer neural network configured to receive the fusion of intermediate features and generate a survivability score corresponding to an overall survivability of the patient; A prediction system comprising: