System and method for multi-modal prediction of patient outcome

Through the multimodal deep learning framework, the problem of inaccurate prediction of treatment choices in patients with NSCLC in the prior art is solved, and higher prediction accuracy and accuracy of treatment plan selection are achieved.

CN120092302APending Publication Date: 2025-06-03VENTANA MEDICAL SYSTEMS INC +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202380070520.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-08-18
Filing Date
2023-10-02
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Existing biomarkers for selecting NSCLC patients who can benefit from targeted or immunotherapy are inaccurate and lack effective prediction methods.

Method used

The multimodal deep learning framework is adopted to extract intermediate features from various modal data such as digital pathology, genetic information, clinical data, demographic data, and other modal data, and the overall survival ability score of patients is generated through the fusion layer neural network.

Benefits of technology

Improves the prediction accuracy of overall survival in patients with NSCLC, providing more accurate prediction scores than single or oligoparameter methods to help select the best treatment options.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120092302A_ABST
    Figure CN120092302A_ABST
Patent Text Reader

Abstract

A method of predicting overall viability of a patient by a machine learning based prediction system includes receiving, by the prediction system, a plurality of input modalities corresponding to the patient, the prediction system including a processor and a memory, the input modalities belonging to different types from each other; generating, by the prediction system, a plurality of intermediate features based on the plurality of input modalities, each input modality of the plurality of input modalities corresponding to one or more features of the plurality of intermediate features; and determining, by the prediction system, a survivability score corresponding to an overall survivability of the patient based on a fusion of the plurality of intermediate features.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims priority and the benefit of U.S. Provisional Application No. 63 / 378,164, filed on October 3, 2022, entitled "SYSTEM AND METHOD FOR MULTIMODAL PREDICTION OF PATIENT OUTCOMES"; U.S. Provisional Application No. 63 / 533,572, filed on August 18, 2023, entitled "ATTENTION - BASED MULTIMODAL - FUSION FOR NON - SMALL CELL LUNG CANCER(NSCLC)PATIENT SURVIVAL PREDICTION"; Indian Application No. 202311030916, filed on April 30, 2023, entitled "INTEGRATING MULTIMODAL DATA FOR NON - SMALL CELL LUNG CANCER(NSCLC)PATIENT SURVIVAL PREDICTION"; Indian Application No. 202311030917, filed on April 30, 2023, entitled "FEATURE GENERATION AND SELECTION FOR AN EFFICIENT MULTIMODAL ANALYSIS ON BIOLOGICAL DATA" and Indian Application No. 202311044011, filed on June 30, 2023, entitled "INTERPRETABLE FEATURE BASED NETWORK FOR CLASSIFYING CELL - OF - ORIGIN FROM WHOLE SLIDE IMAGES IN DIFFUSE LARGE B - CELL LYMPHOMA PATIENTS". Technical Field

[0003] One or more aspects in accordance with some embodiments of the present disclosure relate to systems and methods for predicting patient outcomes. Background Art

[0004] Various forms of cancer have become one of the leading causes of death globally. In particular, lung cancer is one of the most common malignancies and is responsible for approximately 25% of all cancer-related deaths. Approximately 84% of lung cancers are non-small cell lung cancers (NSCLCs), which are a group of lung cancers with similar presentations. Immunotherapy using checkpoint inhibitors, such as anti-PD1 and anti-PD-L1 drugs, has brought favorable clinical outcomes for patients with locally advanced (ad) or metastatic (m) NSCLC. However, the current biomarkers used to select patients who can benefit from targeted or immunotherapy are inaccurate and have great potential for improvement.

[0005] The above information disclosed in this background art section is only for enhancing the understanding of the background art, and thus the information discussed in this background art section does not necessarily constitute the prior art. Summary of the Invention

[0006] Aspects of embodiments of the present disclosure relate to a multimodal prediction system that utilizes a deep learning framework to predict the overall survival of a patient (e.g., an NSCLC patient) from various multimodal data. In some embodiments, the deep learning framework utilizes digital pathology, genetic information, clinical data, patient demographic data, and many other modalities to generate a more accurate prediction than other prediction methods in the relevant field.

[0007] According to some embodiments of the present disclosure, a method for predicting the overall survival ability of a patient by a machine learning-based prediction system is provided. The method includes: receiving, by the prediction system, a plurality of input modalities corresponding to the patient, the prediction system including a processor and a memory, the input modalities belonging to different types from each other; generating, by the prediction system, a plurality of intermediate features based on the plurality of input modalities, each input modality of the plurality of input modalities corresponding to one or more features of the plurality of intermediate features; and determining, by the prediction system, a survival ability score corresponding to the overall survival ability of the patient based on the fusion of the plurality of intermediate features.

[0008] In some embodiments, the plurality of input modalities include: a first modality including histological hematoxylin and eosin (H&E) image data; a second modality including gene sequencing data; a third modality including clinical data; a fourth modality including radiological magnetic resonance imaging (MRI) data; and a fifth modality including immunohistochemistry (IHC) image data.

[0009] In some embodiments, the histological H&E image data includes digital images of tissue samples of the patient stained with hematoxylin and eosin dyes.

[0010] In some embodiments, the digital images include a plurality of images of different tumor regions of the tissue sample.

[0011] In some embodiments, the gene sequencing data includes mRNA gene expression extracted from the tumorous tissue of the patient.

[0012] In some embodiments, the clinical data includes: the age of the patient; the gender of the patient; the tumor stage of the patient; and the performance status of the patient.

[0013] In some embodiments, the IHC image data includes digital images of tissue samples of the patient stained with the PD-L1 biomarker.

[0014] In some embodiments, the prediction system includes: a first convolutional neural network configured to receive a first modality of the plurality of input modalities and generate one or more first intermediate features of the plurality of intermediate features; a first feed-forward neural network configured to receive a second modality of the plurality of input modalities and generate one or more second intermediate features of the plurality of intermediate features; a second feed-forward network configured to receive a third modality of the plurality of input modalities and generate one or more third intermediate features of the plurality of intermediate features; a second convolutional neural network configured to receive a fourth modality of the plurality of input modalities and generate one or more fourth intermediate features of the plurality of intermediate features; and a third convolutional neural network configured to receive a fifth modality of the plurality of input modalities and generate one or more fifth intermediate features of the plurality of intermediate features.

[0015] In some embodiments, the method further includes: combining the first intermediate feature, the second intermediate feature, the third intermediate feature, and the fourth intermediate feature to form the fusion of the plurality of intermediate features.

[0016] In some embodiments, the prediction system further includes: a fusion layer neural network configured to receive the fusion of the plurality of intermediate features and generate the viability score.

[0017] In some embodiments, the method further includes: receiving, by a classifier, an input image of a tissue sample from the patient; and extracting, by the classifier, cell spatial map data from the input image, the cell spatial map data including the cell type and location of each cell in the input image.

[0018] In some embodiments, the input image includes one of a histological H&E image and an IHC image.

[0019] In some embodiments, extracting the cellular spatial map data includes: detecting cells within the input image; generating cellular classification data for the cells detected within the input image, the cellular classification data including the cell type and the location of each cell in the input image; and constructing the cellular spatial map data based on the cellular classification data.

[0020] In some embodiments, the prediction system further includes: a graph convolutional network configured to receive a sixth modality among the plurality of input modalities and generate one or more sixth intermediate features among the plurality of intermediate features, where the sixth modality includes a cellular spatial map corresponding to a histological hematoxylin and eosin (H&E) image or an immunohistochemistry (IHC) image of a tissue sample from the patient.

[0021] In some embodiments, the prediction system includes a multimodal fusion model configured to correlate the plurality of input modalities with the viability score.

[0022] In some embodiments, the method further includes: transmitting the viability score to a display device for display to a user.

[0023] According to some embodiments of the present disclosure, a method for predicting the overall viability of a patient by a machine learning-based prediction system is provided. The method includes: receiving, by the prediction system, a plurality of input modalities corresponding to the patient, the input modalities including: a first modality including histological hematoxylin and eosin (H&E) image data; a second modality including gene sequencing data; a third modality including clinical data; a fourth modality including radiological magnetic resonance imaging (MRI) data; and a fifth modality including immunohistochemistry (IHC) image data; generating, by the prediction system, a plurality of intermediate features based on the plurality of input modalities, each input modality among the plurality of input modalities corresponding to one or more features among the plurality of intermediate features; and determining, by the prediction system, a viability score corresponding to the overall viability of the patient based on the fusion of the plurality of intermediate features.

[0024] In some embodiments, the prediction system includes: a first convolutional neural network configured to receive the first modality and generate one or more first intermediate features of the plurality of intermediate features; a first feed-forward neural network configured to receive the second modality and generate one or more second intermediate features of the plurality of intermediate features; a second feed-forward neural network configured to receive the third modality and generate one or more third intermediate features of the plurality of intermediate features; a second convolutional neural network configured to receive the fourth modality and generate one or more fourth intermediate features of the plurality of intermediate features; and a third convolutional neural network configured to receive the fifth modality and generate one or more fifth intermediate features of the plurality of intermediate features.

[0025] In some embodiments, the method further includes: receiving, by a classifier of the prediction system, an input image of a tissue sample from the patient; and extracting, by the classifier, cell spatial map data from the input image, the cell spatial map data including the cell type and location of each cell in the input image, wherein the input image includes one of the histological H&E image and the IHC image.

[0026] According to some embodiments of the present disclosure, there is provided a prediction system including: a first convolutional neural network configured to receive a histological hematoxylin and eosin (H&E) image and generate one or more first intermediate features; a first feed-forward neural network configured to receive gene sequencing data and generate one or more second intermediate features; a second feed-forward neural network configured to receive clinical data and generate one or more third intermediate features; a second convolutional neural network configured to receive radiological magnetic resonance imaging (MRI) data and generate one or more fourth intermediate features; a fusion circuit configured to combine the first intermediate features, the second intermediate features, the third intermediate features, and the fourth intermediate features to form a fusion of intermediate features via attention gating tensor fusion; and a fusion layer neural network configured to receive the fusion of intermediate features and generate a survival ability score corresponding to the overall survival ability of the patient. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Non-limiting and non-exhaustive embodiments according to the present disclosure are described with reference to the following drawings, in which like reference numerals refer to like components throughout the various views, unless otherwise specified.

[0028] Figure 1 is a flowchart showing various operations that can occur in a pathology context or pathology environment according to some embodiments;

[0029] Figure 2 is a block diagram showing a prediction system according to some embodiments of the present disclosure.

[0030] Figure 3 is a block diagram showing a prediction system according to some embodiments of the present disclosure, the prediction system utilizing derived modalities.

[0031] Figure 4 is a block diagram showing the internal architecture of a prediction system according to some embodiments of the present disclosure.

[0032] Figure 5 is a flowchart showing a process of predicting a patient's overall survival ability by a machine learning-based prediction system according to some embodiments of the present disclosure. Detailed Description

[0033] Hereinafter, aspects of some exemplary embodiments will be described in more detail with reference to the accompanying drawings, in which like reference numerals always refer to like elements. However, the present invention may be embodied in various different forms and should not be construed as limited to only the embodiments shown herein. Instead, these embodiments are provided as examples so that the present disclosure will be thorough and complete and will convey the aspects and features of the present invention fully to those skilled in the art. Accordingly, processes, elements, and techniques that are not necessary for those of ordinary skill in the art to fully understand the aspects and features of the present invention may not be described. Unless otherwise noted, like reference numerals in the entire drawings and the written description represent like elements and thus their description will not be repeated. In the drawings, for clarity, the relative sizes of elements, layers, and regions may be exaggerated.

[0034] Pathology is a medical discipline that attempts to contribute to the diagnosis and treatment of diseases by studying a patient's tissue, cell, and fluid samples. In many applications, tissue samples can be collected from a patient and processed into a form that can typically be analyzed by a physician (e.g., a pathologist), usually under magnification, to diagnose and characterize a related medical condition based on the tissue sample.

[0035] Figure 1 is a flowchart showing various operations that can occur in a pathology environment or pathology system 100. For example, when a treating physician or healthcare provider identifies a patient for whom the analysis of a tissue or fluid sample can be beneficial for diagnosing or treating a medical condition, a tissue or fluid sample can be collected at operation 102. The identity of the patient can be collected and matched with the patient's sample, and the sample can be placed in a sterile container and / or collection medium for further processing.

[0036] Then, the sample can be transported to the pathology registration laboratory at operation 104, where the sample can be received, sorted, organized, and labeled with other samples from other patients for further processing.

[0037] At operation 106, the sample can be further processed as part of a grossing operation. For example, individual tissue samples or specimens can be cut into smaller sections for embedding and subsequent cutting to assemble onto slides.

[0038] Then, at operation 108, the sample or specimen can be mounted or deposited on one or more glass slides. Preparation of the slides can involve applying one or more reagents or stains to the sample, for example, to improve the visibility of different parts of the sample or the contrast between different parts of the sample.

[0039] In some cases, at operation 110, several slides can be assembled or collected in a box or folio during or after the reagent or staining treatment. For example, the box can be carefully labeled with identification information of the individual patient.

[0040] Between each of operations 102 and 110, at operation 112, the sample, specimen, slides can be transported within or between medical facilities (e.g., between a physician's office and a laboratory), or can be stored between processing operations.

[0041] Once the processing of the sample and slides is complete and the pathologist is ready to examine the sample, the slides corresponding to the patient and / or the box holding multiple slides can be transported to the pathologist again at operation 112. At operation 114, the pathologist can examine the slides, for example, using magnification with a microscope. Individual slides can be placed under the objective lens of the microscope, and the microscope and slides can be manipulated and adjusted while the pathologist examines the tissue or fluid.

[0042] Once the pathologist has completed the examination of the slides, the pathologist can attempt to form a medical opinion or diagnosis at operation 116. Meanwhile, at operation 112, the sample or slides can be transported to a long-term storage facility again. In some cases, the sample or slides can be transported to other physicians for further analysis, a second opinion, etc., again before or after a certain storage period.

[0043] An example of the above operation can be performed in a pathology setting, where a pathologist identifies patients (e.g., breast cancer patients) who may respond to a particular immunotherapy by analyzing certain information and deriving a score (e.g., a viability score) that indicates treatment success. Some scoring solutions in the relevant art are single-modal systems that use a single biomarker (e.g., PD-L1) or oligobiomarkers. However, generally, a single biomarker does not accurately predict the effectiveness of immunotherapy. For example, in the case of the PD-L1 biomarker, some negative patients may respond to treatment, while some positive patients may not. Additionally, some treatments have not been found to have a biomarker that can be used as a predictor of treatment efficacy.

[0044] Accordingly, some aspects of the present disclosure relate to a deeply fused multi-modal prediction system that can consider a large amount of different data about a patient, such as demographic information, the histology of the patient's tumor, biomarker-stained slides, the stage of the disease (e.g., cancer staging), patient activity data, etc., to obtain a personalized / personalized patient score based on the patient profile. The score can determine the degree of response of the patient to a particular treatment.

[0045] Figure 2 is a block diagram showing a prediction system 200 according to some embodiments of the present disclosure.

[0046] According to some embodiments, the prediction system 200 is a multi-modal deep learning framework that integrates various modal data to predict patient outcomes and does so in a more robust manner than qualitative clinical assessments or single-modal strategies in the relevant art. In some embodiments, the prediction system 200 is configured to receive a plurality of input modalities 202 associated with a patient and determine a viability score of the patient based on the modalities 202. The individual modalities utilized by the prediction system 200 can be of different types.

[0047] For example, the first modality 202a can include histological hematoxylin and eosin (H&E) image data. The H&E data 202a can include one or more digital images of a tissue sample (e.g., a neoplastic tissue sample) of the patient stained with hematoxylin and eosin dyes. The H&E dyes stain the cell nucleus, extracellular matrix, cytoplasm, and other cellular structures in different colors, enabling a pathologist and the prediction system 200 to distinguish different cellular structures. Additionally, the overall staining pattern of the stain shows the overall layout and distribution of the cells and provides a view of the structure of the tissue sample. In some examples, the H&E image data 202a can include a plurality of image patches (e.g., three image patches) extracted (e.g., randomly selected and extracted) from the viable tumor region of the stained tissue sample.

[0048] The second modality 202b may include genetic sequencing data, such as DNA information of tumor mutations and / or mRNA gene expression extracted from a patient's tumorous tissue. Each tumor cell may have hundreds or thousands of tumor mutant genes. The second modality 202b may include some or all of the gene mutations found in the tissue sample. In some examples, only those expressions most relevant to the patient's survival ability may be included in the second modality 202b.

[0049] The third modality 202c may include clinical data associated with the patient, such as the patient's age, gender, tumor stage, and performance status. The performance status can be measured by the ECOG score, which describes the functional level of the patient in terms of self-care ability, daily activities, and physical ability (e.g., walking, working, etc.). The performance status can also be the Karnofsky performance status, which measures the ability of cancer patients to perform normal tasks. The score can range from 0 to 100, with a higher score indicating that the patient can perform daily activities better.

[0050] The fourth modality 202d may include radiological magnetic resonance imaging (MRI) data, which may include one or more digital MRI images of the patient's tumor region (such as Gd-T1w and T2w-FLAIR scans). The MRI image data can help evaluate the volume of the tumor.

[0051] The fifth modality 202e may include immunohistochemistry (IHC) image data. The IHC data 202e may include one or more digital images of a tissue sample (e.g., a tumorous tissue sample) of the patient stained with the PD-L1 biomarker. The PD-L1 biomarker may produce brown spots when the antibody can attach to tumor cells with PD-L1 expression. The IHC image may correspond to a section of the tissue sample adjacent to the section on which the H&E image 202a is based. In some examples, the cell structures captured in the H&E and IHC images may be the same or substantially the same; however, this may not necessarily be the case. For example, while the H&E image can provide information about the pattern, shape, and structure of the cells in the tissue sample, the IHC image showing the distribution and localization of specific proteins in the sample may not clearly reveal the cell structure of the tissue sample.

[0052] Once the prediction system 200 estimates the survival ability score 204, the score can be transmitted to a server (e.g., a remote server or a cloud server) 206 for further processing and / or transmitted to a display device 208 for display to the user.

[0053] While the above description describes five modalities as examples of input modalities to the prediction system 200, embodiments of the present disclosure are not limited thereto, and the prediction system 200 may employ any suitable type of modality and / or any suitable number of modalities to determine the viability score 204.

[0054] For example, the prediction system 200 may utilize one or more derivatives of the plurality of modalities 202.

[0055] Figure 3 is a block diagram showing the prediction system 200 according to some embodiments of the present disclosure, which utilizes derivative modalities.

[0056] According to some embodiments, the cell structures captured by the H&E image data 202a or the IHC image data 202e are used to generate cell map data that identifies the location and cell type of each cell within the image. Generally, a map can represent the spatial arrangement and neighborhood relationships of different tissue components, which can be characteristics visually observed by a pathologist during specimen examination. Depending on the method employed (e.g., Voronoi diagram, Delaunay triangulation, nearest neighbor graph, etc.), the nodes and edges of the map can represent different elements or characteristics. For example, each node in the map can identify a cell (e.g., the center of the cell's nucleus), and each edge can identify the Euclidean distance between adjacent cells or represent the similarity between adjacent cells. The prediction system 200 can utilize the cell map data to extract features that can be used for viability outcome prediction. Cell classification can be performed manually by a pathologist or by a trained classifier.

[0057] In some embodiments, a classifier (e.g., a machine learning classifier) 230a / b receives an input image that is an image of a stained tissue sample (e.g., an H&E image 202a or an IHC image 202e); detects cells within the input image; and generates cell classification data corresponding to the detected cells. The classification data, including the type and location of each cell in the input image, can be used to generate the cell map data 203a / e.

[0058] In some embodiments, classifier 230a / b includes a neural network (e.g., a convolutional neural network) capable of cell detection and cell classification. The neural network may include multiple layers, and each of the multiple layers performs a convolution operation on an input feature map (IFM) via an applied kernel / filter to generate an output feature map, which serves as the input feature map for subsequent layers. In the first layer of the neural network, the input feature map may be an input image (e.g., H&E image 202a or IHC image 202e). The neural network may be a convolutional neural network (ConvNet / CNN) that can receive an input image, assign importance to various aspects / objects in the image (e.g., via learnable weights and biases), and be able to distinguish them from each other. However, the embodiments of the present disclosure are not limited thereto. For example, the neural network may be a recurrent neural network (RNN) with convolution operations or a random forest network, etc. In some embodiments, the neural network additionally generates a spatial cell map 203a / e based on classification data.

[0059] As Figure 3 shown, in some embodiments, the first classifier 230a generates first cell map data 203a based on H&E image data 202a, and the second classifier 230b generates second cell map data 203b based on IHC image data 202e. However, the embodiments of the present invention are not limited thereto. For example, the first cell map data 203a and the second cell map data 203b may be generated by the same classifier 230a / b instead of two separate classifiers. Additionally, in some embodiments, only one of the first cell map data 203a and the second cell map data 203b may be generated and utilized by the prediction system 200. Although classifiers 230a and 230b are shown as being external to the prediction system 200, the embodiments of the present disclosure are not limited thereto, and one or more of the first classifier 230a and the second classifier 230b may be included in the prediction system 200 (e.g., be part of the prediction system).

[0060] The prediction system may include a multimodal fusion model (e.g., a multimodal deep orthogonal fusion (DOF) model) configured to correlate multiple input modalities with a viability score. In some embodiments, the prediction system 200 includes multiple neural networks (e.g., multiple unimodal neural networks) for processing multiple input modalities 202. Each modality may have a corresponding neural network of a suitable type.

[0061] Figure 4 is a block diagram showing the internal architecture of a prediction system 200 according to some embodiments of the present disclosure.

[0062] According to some embodiments, the prediction system 200 includes a plurality of neural networks (also referred to as unimodal networks / submodels or feature extraction models) corresponding to a plurality of modalities (e.g., having a one-to-one correspondence with the plurality of modalities). Each neural network generates one or more intermediate features (IFs, also referred to as data modality feature representations or modality feature vectors) corresponding to the input modality. The intermediate features from the respective neural networks are fused and utilized by the prediction system 200 to generate the viability score 204. In some examples, the intermediate features are self-identified by the prediction system 200 during training. However, embodiments of the present disclosure are not limited thereto, and during the training of the prediction system 200, a user of the prediction system 200 may manually select one or more of the intermediate features.

[0063] In some embodiments, the prediction system 200 includes a first convolutional neural network (e.g., U-Net) 210 for processing H&E image data of the first modality 202a to generate a corresponding one or more first intermediate data (e.g., histological features; IF1) 220; a first feedforward network 212 for processing genetic data of the second modality 202b to generate a corresponding one or more second intermediate data (e.g., genetic features; IF2) 222; a second feedforward network (or fully connected (FC) network, etc.) 214 for processing clinical data of the third modality 202c to generate a corresponding one or more third intermediate data (e.g., clinical features; IF3) 224; a second convolutional neural network 216 for processing MRI image data of the fourth modality 202d to generate a corresponding one or more fourth intermediate data (e.g., MRI features; IF4) 226; and a third convolutional neural network 218 for processing IHC image data of the fifth modality 202e to generate a corresponding one or more fifth intermediate data (e.g., IHC features; IF5) 228. In some examples, one or more input modalities 202 may be combined and fed into the same network. For example, genetic data 202b and clinical data 202c may be combined and then provided to a single feedforward network, FC network, etc.

[0064] In some embodiments, the prediction system 200 includes a graph convolutional neural network 219 for processing spatial cell graph data of the sixth modality 203 (e.g., 203a / b), which may correspond to (e.g., be based on or extracted from) H&E image data or IHC image data, to generate a corresponding one or more sixth intermediate data (e.g., graph features; IF6) 229. In embodiments where cell graph data is extracted from two or more images (e.g., both H&E images and IHC images), the prediction system 200 may use one or more additional graph convolutional networks to generate additional intermediate features.

[0065] Thus, each input data modality 202 is processed by machine learning (e.g., dedicated deep learning sub-model / single model) 210 / ... / 219, which is trained to generate intermediate data (e.g., modality-specific feature representation or feature representation vector) 220 / ... / 229.

[0066] In some embodiments, the prediction system 200 further includes a fusion network 240 that combines the generated feature representation vectors into a single fused representation (e.g., a single fused vector) and outputs a viability score (also referred to as a fused prognostic risk score) using the fused representation vector. The fusion network 240 may include a data fusion block (e.g., a fusion circuit) 242 and a fused layer neural network 244. The fusion block 242 is configured to combine (e.g., fuse) the respective intermediate features (e.g., IF1 - IF6) generated by multiple unimodal neural networks (e.g., 210, 212, 214, 216, 218, and 219) via attention-gated tensor fusion, which partially controls the expressiveness of each modality 202. The fused layer neural network 244 is configured to generate a viability score based on the fusion of multiple intermediate features, e.g., by utilizing a Cox partial likelihood loss function. The fused layer neural network 244 may include an integrated model of cross-modal pairwise feature interactions. In some examples, the fused layer neural network 244 may be an FC network (composed of several fully connected layers), etc.

[0067] The neural networks constituting the prediction system 200 can be trained all at once in an end-to-end training session with a large training dataset, or the different neural networks corresponding to the modalities can be trained separately and then combined together to form a large system, i.e., the prediction system 200. For example, generally, the unimodal networks 202 can first be trained separately for survival prediction. However, one or more of the image modality feature extraction models 202a, 202d, and 202e can be pre-trained models trained on a separate common dataset (e.g., ImageNet) for different tasks. Doing so can utilize non-medical data to learn richer feature representations. Then, during the multi-modal network training phase, the fusion network (including the fusion block 242 and the fused layer neural network 244) is trained using the features extracted by the trained unimodal networks 202 in the first step as input. During multi-modal network training, the unimodal networks 202 can be frozen, unfrozen, or frozen and then unfrozen in the first few epochs to allow further optimization of unimodal feature extraction.

[0068] The overall approach of the prediction system 200 that utilizes multiple modalities associated with a patient allows the system 200 to make more accurate prediction scores for each individual patient.

[0069] Table 1 compares the performance of the prediction system 200 with the performance of other methods based on the concordance index (CI), which is a standard performance metric for model evaluation in survival analysis. In Table 1, the baseline represents the statistical technique used to predict overall survival, while the other models use deep learning for prediction. As shown in the example of Table 1, although the fusion method of the prediction system 200 uses only three modalities, its prediction ability is better than that of the single-modal method of histological H&E and the dual-modal method of omics (clinical and mRNA).

[0070] Table 1:

[0071]

[0072] Figure 5 is a flowchart showing a process 500 for predicting a patient's overall survival ability by a machine learning-based prediction system according to some embodiments of the present disclosure.

[0073] According to some embodiments, the prediction system 200 receives a plurality of input modalities 202 corresponding to a patient, and the input modalities belong to different types from each other (S502). The plurality of input modalities may include one or more of the following: a first modality 202a, which includes histological hematoxylin and eosin (H&E) image data; a second modality 202b, which includes gene sequencing data; a third modality 202c, which includes clinical data; a fourth modality 202d, which includes radiological magnetic resonance imaging (MRI) data; and a fifth modality 202e, which includes immunohistochemistry (IHC) image data.

[0074] In some embodiments, the prediction system 200 then generates a plurality of intermediate features based on the plurality of input modalities, and each input modality of the plurality of input modalities corresponds to one or more features among the plurality of intermediate features (S504). The prediction system 200 combines the intermediate features to form a fusion of the intermediate features via attention gating tensor fusion.

[0075] The prediction system 200 determines a survival ability score corresponding to the overall survival ability of the patient based on the fusion of the intermediate features (S506).

[0076] As described above, according to some embodiments, the prediction system uses a single deep learning framework for patient survival prediction to fuse heterogeneous multi-modal data. The prediction system can better characterize the interactions between different modalities and allows clinicians to gain more insights from rich clinical and diagnostic information to build integrated personalized healthcare solutions. In addition, the prediction system provides a more accurate prediction score for each individual patient than single or oligoparameters (such as demographics, histology, PD-L1 status, and omics, etc.). This score can be used to select the best treatment plan for the patient.

[0077] According to various embodiments of the present disclosure, the prediction system 200 is implemented using one or more processing circuits or electronic circuits configured to perform the various operations described above. The types of electronic circuits may include a central processing unit (CPU), a graphics processing unit (GPU), an artificial intelligence (AI) accelerator (e.g., a vector processor, which may include a vector arithmetic logic unit configured to efficiently perform operations common to neural networks such as dot products and softmax), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a digital signal processor (DSP), etc. For example, in some cases, aspects of embodiments of the present disclosure are implemented as program instructions stored in a non-volatile computer-readable memory, where when executed by an electronic circuit (e.g., a CPU, a GPU, an AI accelerator, or a combination thereof), the program instructions perform the described operations. The operations performed by the prediction system 200 may be performed by a single electronic circuit (e.g., a single CPU, a single GPU, etc.) or may be distributed among multiple electronic circuits (e.g., multiple GPUs or a CPU in combination with a GPU). The multiple electronic circuits may be local to each other (e.g., located on the same die, within the same package, or within the same embedded device or computer system) and / or may be remote from each other (e.g., communicate via a network such as a local personal area network such as communicate via a local area network such as a local wired and / or wireless network and / or via a wide area network such as the Internet, in which case some operations are performed locally and other operations are performed on a server hosted by a cloud computing service). The one or more electronic circuits that implement the prediction system 200 are referred to herein as a computer or a computer system, which may include a memory storing instructions that, when executed by the one or more electronic circuits, implement the systems and methods described herein.

[0078] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the invention. As used herein, the singular forms "a" and "an" are also intended to include the plural forms unless the context clearly indicates otherwise. It should be further understood that the terms "comprises", "comprising", "includes" and "including", when used in this specification, specify the presence of the stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. Expressions such as "at least one of...", when preceding a list of elements, modify the entire list of elements and not individual elements in the list.

[0079] It should be understood that although the terms "first", "second", "third", etc. may be used herein to describe various elements, components, regions, layers and / or sections, these elements, components, regions, layers and / or sections should not be limited by these terms. These terms are used to distinguish one element, component, region, layer or section from another element, component, region, layer or section. Thus, without departing from the spirit and scope of the present invention, the first element, component, region, layer or section described hereinafter may be referred to as the second element, component, region, layer or section.

[0080] As used herein, the terms "substantially", "about" and similar terms are used as approximate terms and not as terms of degree, and are intended to account for the inherent deviations that would be recognized by one of ordinary skill in the art in the values being measured or calculated. Further, when describing embodiments of the present invention, "may" means "one or more embodiments of the present invention". As used herein, the terms "use", "using" and "used" may be considered to be synonymous with the terms "utilize", "utilizing" and "utilized", respectively.

[0081] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. It should be further understood that terms (such as those defined in common dictionaries) should be interpreted as having a meaning that is consistent with their meaning in the relevant art and / or the context of this specification, and should not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0082] Although aspects of some example embodiments of a system and method for quantifying pathology slides using a cell-based scoring system have been described and shown herein, as will be understood by one of ordinary skill in the art, various modifications and variations can be made without departing from the spirit and scope of the embodiments according to the present disclosure. Thus, it should be understood that the pathology slide manufacturing system and method according to the principles of the present disclosure can be embodiments other than those specifically described herein. The present disclosure is also defined in the appended claims and their equivalents.

Claims

1. A method for predicting the overall survival ability of a patient through a machine learning-based prediction system, the method comprises: receiving, by the prediction system, a plurality of input modalities corresponding to the patient, the prediction system including a processor and a memory, and the input modalities belonging to different types from each other; generating, by the prediction system, a plurality of intermediate features based on the plurality of input modalities, each input modality in the plurality of input modalities corresponding to one or more features in the plurality of intermediate features; and determining, by the prediction system, a survival ability score corresponding to the overall survival ability of the patient based on the fusion of the plurality of intermediate features.

2. The method according to claim 1, wherein the plurality of input modalities comprises: a first modality, the first modality including histological hematoxylin and eosin (H&E) image data; a second modality, the second modality including gene sequencing data; a third modality, the third modality including clinical data; a fourth modality, the fourth modality including radiological magnetic resonance imaging (MRI) data; and a fifth modality, the fifth modality including immunohistochemical (IHC) image data.

3. The method according to claim 2, wherein the histological H&E image data includes digital images of tissue samples of the patient stained with hematoxylin and eosin dyes.

4. The method according to claim 3, wherein the digital images include a plurality of images of different tumor regions of the tissue sample.

5. The method according to claim 2, wherein the gene sequencing data includes mRNA gene expression extracted from the patient's neoplastic tissue.

6. The method according to claim 2, wherein the clinical data comprises: the age of the patient; the gender of the patient; the tumor stage of the patient; and the performance status of the patient.

7. The method according to claim 2, wherein the IHC image data includes digital images of tissue samples of the patient stained with the PD-L1 biomarker.

8. The method according to claim 1, wherein the prediction system comprises: a first convolutional neural network configured to receive the first modality among the plurality of input modalities and generate one or more first intermediate features among the plurality of intermediate features; a first feedforward neural network configured to receive the second modality among the plurality of input modalities and generate one or more second intermediate features among the plurality of intermediate features; a second feedforward network configured to receive the third modality among the plurality of input modalities and generate one or more third intermediate features among the plurality of intermediate features; a second convolutional neural network configured to receive the fourth modality among the plurality of input modalities and generate one or more fourth intermediate features among the plurality of intermediate features; and a third convolutional neural network configured to receive the fifth modality among the plurality of input modalities and generate one or more fifth intermediate features among the plurality of intermediate features.

9. The method according to claim 8, further comprising: combining the first intermediate feature, the second intermediate feature, the third intermediate feature, and the fourth intermediate feature to form the fusion of the plurality of intermediate features.

10. The method according to claim 8, wherein the prediction system further comprises: a fusion layer neural network configured to receive the fusion of the plurality of intermediate features and generate the viability score.

11. The method according to claim 1, further comprising: receiving, by a classifier, an input image of a tissue sample from the patient; and extracting, by the classifier, cell spatial map data from the input image, the cell spatial map data including the cell type and location of each cell in the input image.

12. The method according to claim 11, wherein the input image comprises one of a histological H&E image and an IHC image.

13. The method according to claim 11, wherein the extracting the cell spatial map data comprises: detecting cells within the input image; generating cell classification data of the cells detected within the input image, the cell classification data including the cell type and the location of each cell in the input image; and constructing the cell spatial map data based on the cell classification data.

14. The method according to claim 1, wherein the prediction system further comprises: a graph convolutional network configured to receive a sixth modality among the plurality of input modalities and generate one or more sixth intermediate features among the plurality of intermediate features, wherein the sixth modality comprises a cell spatial map corresponding to a histological hematoxylin and eosin (H&E) image or an immunohistochemistry (IHC) image of a tissue sample from the patient.

15. The method according to claim 1, wherein the prediction system comprises a multimodal fusion model configured to correlate the plurality of input modalities with the viability score.

16. The method according to claim 1, further comprising: transmitting the viability score to a display device for display to a user.

17. A method for predicting the overall viability of a patient by a machine learning-based prediction system, the method comprising: receiving, by the prediction system, a plurality of input modalities corresponding to the patient, the input modalities including: a first modality comprising histological hematoxylin and eosin (H&E) image data; a second modality comprising gene sequencing data; a third modality comprising clinical data; a fourth modality comprising radiological magnetic resonance imaging (MRI) data; and a fifth modality comprising immunohistochemistry (IHC) image data; generating, by the prediction system, a plurality of intermediate features based on the plurality of input modalities, each input modality among the plurality of input modalities corresponding to one or more features among the plurality of intermediate features; and The prediction system determines a viability score corresponding to the overall viability of the patient based on the fusion of the plurality of intermediate features.

18. The method according to claim 17, wherein the prediction system comprises: a first convolutional neural network configured to receive the first modality and generate one or more first intermediate features among the plurality of intermediate features; a first feedforward neural network configured to receive the second modality and generate one or more second intermediate features among the plurality of intermediate features; a second feedforward neural network configured to receive the third modality and generate one or more third intermediate features among the plurality of intermediate features; a second convolutional neural network configured to receive the fourth modality and generate one or more fourth intermediate features among the plurality of intermediate features; and a third convolutional neural network configured to receive the fifth modality and generate one or more fifth intermediate features among the plurality of intermediate features.

19. The method according to claim 17, further comprising: receiving, by a classifier of the prediction system, an input image of a tissue sample from the patient; and extracting, by the classifier, cell spatial map data from the input image, the cell spatial map data including the cell type and location of each cell in the input image, wherein the input image includes one of a histological H&E image and an IHC image.

20. A prediction system, which comprises: a first convolutional neural network configured to receive a histological hematoxylin and eosin (H&E) image and generate one or more first intermediate features; a first feedforward neural network configured to receive gene sequencing data and generate one or more second intermediate features; a second feedforward neural network configured to receive clinical data and generate one or more third intermediate features; a second convolutional neural network configured to receive radiological magnetic resonance imaging (MRI) data and generate one or more fourth intermediate features; a fusion circuit configured to combine the first intermediate features, the second intermediate features, the third intermediate features, and the fourth intermediate features to form a fusion of intermediate features via attention gating tensor fusion; and a fusion layer neural network configured to receive the fusion of the intermediate features and generate a viability score corresponding to the overall viability of the patient.

Citation Information

Patent Citations

  • Injection mold and injection molding equipment

    CN116811152A

  • Super-resolution photoetching resolution enhancement method based on transient illumination

    CN117008427A

  • Data processing method and device, electronic equipment and storage medium

    CN117076451A