Computer-implemented method for performing clinical predictions
Patent Information
- Application Number
- JP2024527542
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-12-17
- Filing Date
- 2022-12-15
- Publication Date
- 2025-12-09
AI Technical Summary
Existing clinical prediction algorithms struggle to effectively combine and utilize multiple data modalities, leading to suboptimal performance and accuracy in predicting patient outcomes due to the complexity of integrating diverse data types.
A computer-implemented method using a deep neural network architecture that employs trainable data embedders and aggregation networks with attention and transformer layers to generate embedded modality representations, allowing for optimal combination and fusion of clinical data from various sources such as histological images, gene expression, and radiological scans.
Enhances the performance and accuracy of clinical predictions by leveraging the unique information from multiple data modalities, improving the reliability and precision of patient outcome predictions.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a computer-implemented method for performing clinical prediction, as well as a computer program and a computer-readable storage medium for carrying out the method according to the invention. The method and device may be used, inter alia, in the field of clinical research and drug development. However, other fields of application for the invention are also feasible. [Background technology]
[0002] In the field of clinical research and drug development, there is a large amount of patient data available from different sources and in different modalities. Typical data types available are patient clinical data, full-body microscopic images of biopsies and surgical specimens, gene expression data, proteomics, radiological images such as magnetic resonance imaging (MRI) and computed tomography (CT), and demographics. Personalized medicine aims to identify and match patients with the best drugs that can be most beneficial to them. To do this, algorithms have been developed to predict patient survival and response to certain treatments. Other algorithms can be developed to predict or confirm a patient's diagnosis, or to curate and / or complete patient data by predicting missing patient data points.
[0003] So far, these types of prediction algorithms have typically been trained on only one modality of input data to avoid the complexities inherent in the process of combining different types of data modalities. However, each modality can contain different information, and each of these pieces of information can bring additional value to the final prediction and performance of the system. Furthermore, using multiple modalities can be a better approximation of clinician behavior when analyzing patient data. Therefore, combining different patient modalities can improve the quality of the prediction system and help it gain more trust from experts.
[0004] Methods for multimodal fusion are known, for example, from Cheerla, Anika, and Olivier Gevaert. “Deep learning with multimodal representation for pancancer prognosis prediction.” Bioinformatics 35.14(2019):i446-i454, Vale-Silva, Luis A., and Karl Rohr. “Long-term cancer survival prediction using multimodal deep learning.” Scientific Reports 11.1(2021):1-12, and Sun, Li, et al. “Brain tumor segmentation and survival prediction using multimodal MRI scans with deep learning.” Frontiers in neuroscience 13(2019):810.
[0005] Despite the achievements of known methods for multimodal fusion, finding an appropriate mechanism for data multimodal fusion remains challenging. As with multiple instance learning (MIL), the fusion of multimodal representations (instance representations in MIL) is a critical step in the algorithm and can significantly affect performance and accuracy. Summary of the Invention
[0006] It is therefore desirable to provide a method that addresses the above-mentioned technical challenges. In particular, a method should be proposed that allows for improved performance and accuracy of data multimodal fusion. [Means for solving the problem]
[0007] This problem is addressed by a computer implemented method for performing clinical prediction comprising the features of the independent claims. Advantageous embodiments which may be implemented alone or in any combination are set out in the dependent claims as well as in the whole specification.
[0008] When used below, the terms "having", "comprises" or "includes" or any grammatical variants thereof are used in a non-exclusive manner. These terms may therefore refer both to the situation where, apart from the features introduced by these terms, no further features are present in the entity described in this context, and to the situation where one or more further features are present. As an example, the expressions "A has B", "A comprises B" and "A includes B" may all refer to the situation where, apart from B, no other elements are present in A (i.e., the situation where A is exclusively composed of B), and to the situation where, apart from B, one or more further elements are present in entity A, such as element C, elements C and D, or even further elements.
[0009] Furthermore, it should be noted that the terms "at least one" or "one or more" or similar expressions indicating that a feature or element may be present one or more times are typically used only once when introducing each feature or element. In the following, in most cases, when referring to each feature or element, the expressions "at least one" or "one or more" will not be repeated, despite the fact that each feature or element may be present one or more times.
[0010] Furthermore, when used hereinafter, the terms "preferably", "more preferably", "particularly", "more particularly", "particularly" or "more particularly" or similar terms are used in relation to optional features and do not limit alternative possibilities. Features introduced by these terms are therefore optional features and are not intended to limit the scope of the claims in any way. The invention may be implemented by using alternative features, as would be understood by a person skilled in the art. Similarly, features introduced by "in an embodiment of the invention" or similar expressions are intended to be optional features and do not entail any limitations in relation to alternative embodiments of the invention, any limitations in relation to the scope of the invention, or any limitations in relation to the possibility of combining a feature introduced in such a manner with other optional or non-optional features of the invention.
[0011] In a first aspect of the present invention, a computer-implemented method for performing clinical prediction is disclosed.
[0012] As used herein, the term "computer-implemented method" is a broad term and should be given its general and ordinary meaning to those skilled in the art and should not be limited to a special or special meaning. The term may specifically, but not limited to, refer to a method involving at least one computer and / or at least one computer network or cloud. The computer and / or computer network and / or cloud may comprise at least one processor configured to perform at least one of the method steps of the method according to the invention. Preferably, each of the method steps is performed by the computer and / or computer network and / or cloud. The method may be performed completely automatically, in particular without requiring user interaction. As used herein, the term "automatically" is a broad term and should be given its general and ordinary meaning to those skilled in the art and should not be limited to a special or special meaning. The term may specifically, but not limited to, refer to a process performed completely by at least one computer and / or computer network and / or cloud and / or machine, in particular without requiring manual action and / or user interaction.
[0013] As used herein, the term "clinical prediction" is a broad term and should be given its common and ordinary meaning to those skilled in the art and should not be limited to a special or special meaning. The term may specifically refer to, but is not limited to, the estimation of at least one patient endpoint. The patient endpoint may include one or more of at least one measure of efficacy of a treatment, at least one measure of tolerability of a treatment, at least one measure of usefulness of a treatment, at least one measure of harmfulness of a treatment, mortality, morbidity, side effects, health-related quality of life, and the like. The clinical prediction may include at least one predictive value. For example, the clinical prediction may include one or more of predicting response to treatment of a disease, predicting the risk of a patient having a disease, predicting outcomes, deriving biomarkers, and identifying targets for drug development, and the like. The disclosed technology can be used to treat various types of diseases, such as various types of cancer, and / or answer other clinical questions, and the like. For example, the clinical prediction may include predicting whether a patient may be resistant or sensitive to a drug treatment, for example, for cancer. For example, clinical predictions may include predicting patient survival rates for various types of treatments (e.g., immunotherapy, chemotherapy, etc.) for diseases such as cancer. The technology can be applied to other disease areas and other clinical hypotheses. Clinical predictions may be generated and / or provided, for example, as a histogram that represents the clinical prediction or shows the progression over time of at least one variable associated with the clinical prediction.
[0014] As used herein, the term "patient" is a broad term and should be given its common and ordinary meaning to those skilled in the art, and should not be limited to a specific or special meaning. The term may specifically refer to, but is not limited to, a human or an animal, regardless of whether the human or animal is in a healthy state or suffers from one or more diseases.
[0015] The method includes, by way of example, the following steps, which may be performed in the given order. It should be noted, however, that different orders are also possible. Furthermore, it is also possible to perform one or more of the method steps once or repeatedly. Furthermore, it is possible to perform two or more method steps simultaneously or overlapping in time. The method may include further method steps not listed.
[0016] The method comprises the following steps: i) retrieving input data comprising a plurality of different modalities of a patient via at least one communication interface of a processing device; ii) processing the input data by using a processing device, the processing including generating embedded modality representations from the input data by using at least one trainable data embedder, the processing including generating a clinical prediction by combining the embedded modality representations using at least one aggregation network, the aggregation network including at least one attention layer and / or at least one transformer layer; and iii) generating an output of the clinical prediction by using the processing device. Includes.
[0017] As generally used herein, the term "processing device" is a broad term and should be given its common and ordinary meaning to those skilled in the art and should not be limited to a specific or special meaning. The term may specifically, but not limited to, refer to any logic circuitry configured to perform basic operations of a computer or system, and / or may generally refer to a device configured to perform calculations or logical operations. In particular, the processing device may be configured to process basic instructions that drive a computer or system. By way of example, the processing device may include at least one arithmetic logic unit (ALU), at least one floating point unit (FPU), such as a numeric coprocessor or numeric coprocessor, a number of registers, specifically registers configured to provide operands to the ALU and store operation results, and memories, such as L1 and L2 cache memories. In particular, the processing device may be a multi-core processor. In particular, the processing device may be or comprise a central processing unit (CPU) or a graphics processing unit (GPU). Additionally or alternatively, the processing device may be or comprise a microprocessor, and thus, in particular, the elements of the processing device may be included on one single integrated circuit (IC) chip. Additionally or alternatively, the processing device may be or comprise one or more application specific integrated circuits (ASICs) and / or one or more field programmable gate arrays (FPGAs), etc. The processing device may be, in particular, configured, such as by software programming, to perform one or more evaluation operations.
[0018] As used herein, the term "communication interface" is a broad term and should be given its general and ordinary meaning to those skilled in the art and should not be limited to a special or special meaning. The term may specifically refer to, but is not limited to, an item or element forming a boundary configured to transfer information. In particular, the communication interface may be configured to transfer information from a computing device, such as a computer, for example, to transmit or output information to another device. Additionally or alternatively, the communication interface may be configured to transfer information to a computing device, such as a computer, for example, to receive information. The communication interface may specifically provide a means for transferring or exchanging information. In particular, the communication interface may provide a data transfer connection, such as Bluetooth, NFC, inductive coupling, etc. By way of example, the communication interface may be or may comprise at least one port comprising one or more of a network or Internet port, a USB port, and a disk drive. The communication interface may further comprise at least one display device. The communication interface may be at least one web interface.
[0019] As used herein, the term "retrieving" is a broad term and should be given its general and ordinary meaning to those skilled in the art and should not be limited to a special or special meaning. The term may specifically refer to a system, specifically a computer system process, that generates data and / or retrieves data from any data source, such as, but not limited to, a data storage device, a network, or a further computer or computer system. Retrieving may specifically be performed by at least one computer interface, for example, via a port, such as a serial port or a parallel port. Retrieving may include several substeps, such as obtaining one or more items of primary information, for example, by using a processor, and generating secondary information by utilizing the primary information, for example, by applying one or more algorithms to the primary information. Retrieving may include performing at least one measurement using at least one medical device, such as, for example, a magnetic resonance imaging (MRI), a computed tomography (CT), or the like.
[0020] As used herein, the term "input data" is a broad term and should be given its common and ordinary meaning to those skilled in the art and should not be limited to a specific or special meaning. The term may specifically refer to, but is not limited to, at least one value, parameter, image data, etc., that can be processed in step ii).
[0021] The input data includes multiple different modalities of a patient. The input data may include multimodal clinical data. As used herein, the term "modality" is a broad term and should be given its general and ordinary meaning to those skilled in the art and should not be limited to a special or special meaning. The term may specifically refer to, but is not limited to, a channel of input, such as an independent channel of input. The input data may include a single data point from a single modality, or multiple data points from the same modality for a single patient. The multiple modalities of a patient may include one or more of at least one histological tissue image, at least one full-surface microscopic image of a biopsy and / or surgical specimen, radiological images such as magnetic resonance imaging (MRI) and computed tomography (CT), genomic data, gene expression data, proteomics, patient clinical data, and demographics. The multimodal clinical data may generally refer to various types of clinical data, such as molecular data, biopsy image data, etc.
[0022] The processing of the input data in step ii) includes generating an embedded modality representation from the input data by using at least one trainable data embedder. The embedded modality representation may be a patient-level representation, also referred to as a patient-level embedding. As used herein, the term "data embedder", also referred to as "embedding", is a broad term and should be given its general and ordinary meaning to those skilled in the art and should not be limited to a special or special meaning. The term may specifically, but is not limited to, refer to at least one network layer configured to convert the input data into a continuous vector representation. For example, the input data may be an image and the data embedder is designed to convert the image into a low-dimensional representation of the image. The output of the trainable data embedder may be a generic embedding representation for each modality or a multiple instance embedding for each modality.
[0023] As used herein, the term "trainable data embedder" is a broad term and should be given its common and ordinary meaning to those skilled in the art and should not be limited to a special or special meaning. The term may specifically, but not be limited to, refer to the fact that the embedder can be further trained and / or updated based on additional training data. Specifically, the embedder is trained on a training dataset. The embedder may be trained using machine learning. At least one embedder for each modality may be used. Each embedder for a modality may be trained on historical data from the modality. For example, each embedder may be trained on historical data from histological tissue imaging, whole slide microscopic imaging, or radiological imaging, or on historical genomic data, historical gene expression data, historical proteomics, historical patient clinical data, or historical demographics. The embedder may be updated by using newly received input data.
[0024] The training data may include data from multiple patients, each with a different data modality and known ground truth outcome. For example, the training data may include at least one histological whole slide image (e.g., H&E). The slide may have expert annotations and / or tissue detection masks. The whole slide image is a high-resolution image, and tile images may be extracted from the entire slide and / or from specific expert annotations and / or from the tissue mask to generate image modality data points. The patient may further have genomic or proteomic data points. These may be vectors of raw or normalized floating point values. The training data may include patient metadata, such as age, sex, and clinical data, such as diagnosis, HER2 positive status, and the like. The training data may include one or more patient embeddings generated with different analysis systems.
[0025] The process includes generating a clinical prediction by combining the embedded modality representations using at least one aggregation network. The embedded modality representations may be introduced as inputs to an attention layer and / or a transformer layer. As used herein, the term "combining" is a broad term and should be given its general and ordinary meaning to one skilled in the art and should not be limited to a special or special meaning. The term may specifically refer to, but is not limited to, data fusion and / or data aggregation. As used herein, the term "aggregation network" is a broad term and should be given its general and ordinary meaning to one skilled in the art and should not be limited to a special or special meaning. The term may specifically refer to, but is not limited to, a deep neural network architecture designed to generate a clinical prediction by combining embedded modality representations from multiple modalities. The combining may include one or more of: considering a sum of transformed data points, considering a maximum of transformed data points, using an attention MIL model, and using a vision transformer.
[0026] The aggregation network comprises at least one attention layer and / or at least one transformer layer. The attention layer and / or the transformer layer may implement a self-attention mechanism. The self-attention mechanism may enable generating an optimal combination strategy for the multimodal data.
[0027] As used herein, the term "attention layer" is a broad term and should be given its general and ordinary meaning to those skilled in the art and should not be limited to a special or special meaning. The term may specifically, but not be limited to, refer to a layer of a neural network designed to enhance at least one important part of input data and fade out other parts. The importance of the input data may be modality dependent. The importance of the input data may be learned through training data. The attention layer may use dot product attention and / or multi-head attention.
[0028] As used herein, the term "transformer layer" is a broad term and should be given its general and ordinary meaning to those skilled in the art and should not be limited to a special or special meaning. The term may specifically refer to, but is not limited to, a layer of a deep learning model that employs a mechanism of self-attention and gives different weights to the importance of each part of the input data. For example, the transformer layer may be based on a vision transformer model. The vision transformer model is an image classification model that is based on a transformer encoder architecture and uses embeddings of image patches as inputs. In the vision transformer model, an image may be split into patches, which are then flattened, projected into a low-dimensional embedding, added to a position embedding, and fed into a transformer encoder network. The output of the transformer encoder may be used as an input to a multi-layer perceptron (MLP) head to generate a final prediction. The MLP head may comprise a series of linear transformation layers. The transformer encoder may comprise n encoders. Each encoder may comprise a multi-head attention layer, a normalization layer, and an MLP layer. As used herein, the term "multi-layer perceptron neural network" or "MLP neural network" is a broad term and should be given its general and ordinary meaning to those skilled in the art, and should not be limited to a special or special meaning. The term may specifically refer to, but is not limited to, a type of feed-forward artificial neural network. Residual skip connections may further be used between sublayers of the encoder to enable interaction between different level representations and prevent the vanishing gradient problem. Multi-head attention may be based on running a self-attention mechanism multiple times.For further design of multi-head attention, one can refer to Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017), “Attention is all you need in Advances in neural information processing systems”, pp. 5998-6008. Self-attention is a mechanism that allows learning relationships between different inputs and taking these relationships into account during model training. The use of vision transformers as secondary networks to aggregate different modality embeddings is a novel way of using models of this category. The present invention proposes to discover relevant relationships between different modalities and use them as additional information in training.
[0029] The present invention proposes to use a deep neural network architecture dedicated to predicting patient endpoints from multiple modalities. Specifically, the present invention proposes to fuse embedded modality representations, also called multimodal data representations, by using attention-based pooling methods for multimodal representation fusion and / or transformer-based pooling methods for multimodal representation fusion. The patient's multimodal data representation may be introduced as input to an attention layer and / or a transformer layer that may learn through backpropagation the optimal combination and / or attention strategy (parameters) for this dataset. The proposed network architecture may utilize an attention MIL module and / or a transformer module architecture. This may make it possible to create an optimal combination strategy for multimodal data via, for example, the self-attention mechanism available in the transformer module architecture.
[0030] The MIL network is a type of weakly supervised learning model in which training instances are grouped into bags and labels are assigned to the bags rather than to single instances. Attention-based MIL networks can aggregate different instance outputs according to an attention mechanism, thus making it possible to evaluate the contribution of each instance in a bag to the final bag output. Both the prediction network and the attention network can be trained simultaneously.
[0031] The method may comprise an attention layer and / or a transformer layer that learns, e.g., trains and / or optimizes, by backpropagation, an optimal combination and / or attention strategy for determining parameters for a dataset.
[0032] The input data may include at least one data point from each of a plurality of different modalities. For example, the at least one data point may include at least one image tile from a single biopsy, at least one gene sequence, at least one tile from a radiology image, etc. The input data may include multiple data points from the same modality of a single patient, such as, for example, multiple image tiles from a single biopsy, multiple gene sequences, multiple tiles from a radiology image, etc. For example, the input data may include a single data point from each of a plurality of different modalities. For example, the input data may include multiple data points from each of a plurality of different modalities. For example, the input data may include a single data point from one or more modalities and multiple data points from at least one other modality.
[0033] The method may include generating an embedded modality representation from each data point. For example, in the case of a single data point from each of the different modalities, the method may include generating an embedded modality representation from each of the single data points. The method may further include generating a clinical prediction from the embedded modality representations of the different modalities using an aggregation network.
[0034] The method may include generating embedded modality representations from each data point, combining the generated embedded modality representations for each of the modalities separately, and generating a clinical prediction from the combined embedded modality representations using an aggregation network.
[0035] For example, the input data may include multiple data points from the same modality for a single patient. The method may include generating an embedded modality representation from each data point, combining the generated embedded modality representations separately for each of the different modalities, and generating a clinical prediction from the combined embedded modality representations using an aggregation network. Additionally or alternatively, the method may include combining the multiple data points, generating a global embedded modality representation from the combined data points for each of the different modalities, and generating a clinical prediction from the global embedded modality representation using an aggregation network.
[0036] The method may include combining multiple data points, generating a global embedded modality representation from the combined data points for each of the modalities, and generating a clinical prediction from the global embedded modality representation using an aggregation network. For example, in the case of multiple data points from the same modality for a single patient, the method may include creating an embedding from each data point, combining the embeddings for each modality separately, and using a network layer to convert the patient-level multi-modal data points to a patient-level prediction. Additionally or alternatively, in the case of multiple data points from the same modality for a single patient, the multiple data points from the multi-modal data may be directly combined with each other and a network layer may be used to convert them to a patient-level prediction without first combining them into one patient-level modality data point.
[0037] In one embodiment, the different modalities are converted to patient-level embeddings by a first-order attention MIL network layer, which then feeds into a second-order attention MIL network that combines the multimodal patient-level data into patient-level predictions. Using backpropagation, the network can be trained and optimized.
[0038] In one embodiment, the different modalities are transformed into patient-level embeddings by a first-order attention MIL network layer, which are then input to a second-order vision transformer network that combines the multimodal patient-level data into patient-level predictions. Using backpropagation, the network can be trained and optimized.
[0039] In one embodiment, the different modalities are transformed into patient-level embeddings by a first-order vision transformer network layer, which are then input to a second-order vision transformer network that combines the multimodal patient-level data into patient-level predictions. Using backpropagation, the network can be trained and optimized.
[0040] In one embodiment, the different modalities are transformed into patient-level embeddings by a first-order vision transformer network layer, which are then fed into a two-attention MIL network that combines the multi-modal patient-level data into patient-level predictions. Using backpropagation, the network can be trained and optimized.
[0041] In one embodiment, the different modalities are input directly into an embedder network, and the resulting embeddings are input into a first-order attention MIL network layer that combines the multi-modal raw data into patient-level predictions. Using backpropagation, the network can be trained and optimized.
[0042] In one embodiment, the different modalities are input directly into an embedder network, and the resulting embeddings are input into a primary vision transformer network layer that combines the multimodal raw data into patient-level predictions. Using backpropagation, the network can be trained and optimized.
[0043] In one embodiment, depending on the data type, each modality is converted to a patient-level embedding by a first-order attention MIL network layer or can be directly input to the embedder network. All resulting embeddings are input to a second-order attention MIL network that combines the multi-modal data into a patient-level prediction. Using backpropagation, the network can be trained and optimized.
[0044] In one embodiment, depending on the data type, each modality is converted to a patient-level embedding by a first-order attention MIL network layer or can be directly input to the embedder network. All resulting embeddings are input to a second-order vision transformer network that combines the multimodal data into a patient-level prediction. Using backpropagation, the network can be trained and optimized.
[0045] The method may further include at least one pre-processing step. The pre-processing step may include converting the raw data into a new format. This step may depend on the modality type. For example, by using a pre-trained embedder, it is possible to project the histological tile images into a different space. This may help to capture relevant information in the raw data and possibly speed up the training process.
[0046] The output of the clinical prediction may include one or more of the following: information on medication for the patient, information on the patient's survival, information on the response to at least one specific treatment, information confirming the patient's diagnosis, information curating and / or completing the patient data by predicting missing patient data points. The method includes at least one output step including providing the clinical prediction via at least one output interface. The term "output interface" as used herein relates to any unit configured to transfer information from the processing device to another entity, which may be a further data processing device and / or a user. Thus, the output interface may comprise a user interface such as a suitably configured display or may be a printer.
[0047] Further disclosed and proposed herein is a computer program comprising computer executable instructions which, when executed on a computer or a computer network, executes the method according to the invention in one or more of the embodiments contained herein. In particular, the computer program may be stored on a computer readable data carrier and / or a computer readable storage medium.
[0048] As used herein, the terms "computer-readable data carrier" and "computer-readable storage medium" may specifically refer to non-transitory data storage means such as hardware storage media that store computer-executable instructions. A computer-readable data carrier or storage medium may specifically be or include a storage medium such as a random access memory (RAM) and / or a read-only memory (ROM).
[0049] Thus, in particular, one, several or even all of the method steps i) to iii) set out above may be implemented using a computer or a computer network, preferably by using a computer program.
[0050] Further disclosed and proposed herein is a computer program product, which comprises program code means for executing the method according to the invention in one or more of the embodiments contained herein when the program is executed on a computer or a computer network. In particular, the program code means may be stored on a computer readable data carrier and / or a computer readable storage medium.
[0051] Further disclosed and proposed herein is a data carrier having stored thereon a data structure which, after being loaded into a computer or computer network, for example into a working or main memory of the computer or computer network, is capable of performing a method according to one or more of the embodiments disclosed herein.
[0052] Further disclosed and proposed herein is a computer program product having program code means stored on a machine-readable carrier for executing the method according to one or more of the embodiments disclosed herein when the program is executed on a computer or computer network. As used herein, a computer program product refers to a program as a tradeable product. The product may generally be present in any format, such as a paper format, or may be present on a computer-readable data carrier and / or a computer-readable storage medium. In particular, the computer program product may be distributed via a data network.
[0053] Finally, a modulated data signal containing instructions readable by a computer system or computer network for carrying out a method according to one or more of the embodiments disclosed herein is disclosed and proposed herein.
[0054] With respect to computer-implemented aspects of the present invention, one or more or all of the method steps of the method according to one or more of the embodiments disclosed herein may be performed by using a computer or a computer network.Thus, in general, any of the method steps including providing and / or manipulating data may be performed by using a computer or a computer network.In general, these method steps may include any method steps, except for those method steps that typically require manual operations, such as providing a sample and / or performing the actual measurement in certain aspects.
[0055] Specifically, in this specification, a computer or computer network comprising at least one processor, the processor being configured to execute a method according to one of the embodiments described herein; a computer-loadable data structure configured, when executed on a computer, to perform a method according to one of the embodiments described herein; A computer program configured, when it is run on a computer, to carry out a method according to one of the embodiments described herein, a computer program comprising program means for carrying out a method according to one of the embodiments described herein when said computer program is run on a computer or a computer network; A computer program comprising program means according to the preceding embodiment, the program means being stored on a computer readable storage medium; a storage medium storing a data structure, the data structure being configured to perform a method according to one of the embodiments described herein after being loaded into a main memory and / or a working memory of a computer or a computer network; and A computer program product comprising program code means storable or stored on a storage medium, which, when executed on a computer or on a computer network, performs a method according to one of the embodiments described herein. is further disclosed.
[0056] In a further aspect of the present invention, a clinical prediction device is disclosed. The clinical prediction device comprises at least one processing device having at least one communication interface configured to retrieve input data. The input data comprises a plurality of different modalities of a patient. The processing device is configured to process the input data. The processing comprises generating an embedded modality representation from the input data by using at least one trainable data embedder. The processing comprises generating a clinical prediction by combining the embedded modality representations using at least one aggregation network. The aggregation network comprises at least one attention layer and / or at least one transformer layer. The processing device is configured to generate an output of the clinical prediction.
[0057] The clinical prediction device may be configured to execute the method for performing clinical prediction according to the invention. With regard to the definitions and embodiments of the clinical prediction device, reference is therefore made to the definitions and embodiments described with regard to the method.
[0058] In summary, without excluding further embodiments, the following embodiments can be envisaged:
[0059] Example 1. A computer-implemented method for performing clinical predictions, comprising: i) retrieving input data comprising a plurality of different modalities of a patient via at least one communication interface of a processing device; ii) processing the input data by using a processing device, the processing including generating embedded modality representations from the input data by using at least one trainable data embedder, the processing including generating a clinical prediction by combining the embedded modality representations using at least one aggregation network, the aggregation network including at least one attention layer and / or at least one transformer layer; iii) generating a clinical prediction output by using the processing device; and A method comprising:
[0060] Example 2. A method according to the preceding embodiments, wherein the output of the clinical prediction includes one or more of: information regarding medications for the patient, information regarding the patient's survival, information regarding response to at least one particular treatment, information confirming the patient's diagnosis, information curating and / or completing patient data by predicting missing patient data points.
[0061] Example 3. A method according to any one of the preceding embodiments, comprising at least one output step comprising providing a clinical prediction via at least one output interface.
[0062] Example 4. The method according to any one of the preceding embodiments, wherein the output of the trainable data embedder is a generic patient-level embedding representation per modality, or a multiple instance embedding for each modality.
[0063] Example 5. A method according to any one of the preceding embodiments, wherein the embedded modality representation is introduced as input to an attention layer and / or a transformer layer.
[0064] Example 6. The method according to any one of the preceding embodiments, wherein the multiple modalities of the patient include one or more of at least one histological tissue image, at least one full-face microscopic image of a biopsy and / or surgical specimen, radiological images such as magnetic resonance imaging (MRI) and computed tomography (CT), genomic data, gene expression data, proteomics, patient clinical data, and demographics.
[0065] Example 7. A method according to any one of the preceding embodiments, comprising an attention layer and / or a transformer layer learning an optimal combination and / or attention strategy through backpropagation.
[0066] Example 8. The method according to any one of the preceding embodiments, wherein the input data includes at least one data point from each of the different modalities, and the method includes generating an embedded modality representation from each of the data points, and generating a clinical prediction from the embedded modality representations of the different modalities using an aggregation network.
[0067] Example 9. The method according to any one of the preceding embodiments, wherein the input data includes, for at least one of the different modalities, multiple data points from the same modality for a single patient.
[0068] Example 10. A method according to the preceding embodiments, comprising generating an embedded modality representation from each data point, combining the embedded modality representations generated separately for each of the different modalities, and generating a clinical prediction from the combined embedded modality representations using an aggregation network.
[0069] Example 11. A method according to the prior embodiment, comprising combining a plurality of data points and generating a global embedded modality representation from the combined data points for each of the different modalities, and generating a clinical prediction from the global embedded modality representation using an aggregation network.
[0070] Example 12. The method according to any one of the preceding embodiments, wherein the different modalities are converted into embedded modality representations by a first-order attention MIL network layer and then input into a second-order attention MIL network that combines the embedded modality representations into a clinical prediction.
[0071] Example 13. The method according to any one of the preceding embodiments, wherein the different modalities are transformed into embedded modality representations by a first-order attention MIL network layer and then input into a second-order vision transformer network that combines the embedded modality representations into a clinical prediction.
[0072] Example 14. A method according to any one of the preceding embodiments, wherein the different modalities are transformed into embedded modality representations by a primary vision transformer network layer and then input into a secondary vision transformer network that combines the embedded modality representations into a clinical prediction.
[0073] Example 15. The method according to any one of the preceding embodiments, wherein the different modalities are transformed into embedded modality representations by a primary vision transformer network layer and then input into a secondary attention MIL network that combines the embedded modality representations into a clinical prediction.
[0074] Example 16. The method according to any one of the preceding embodiments, wherein the different modalities are input into an embedder network and the resulting embedded modality representations are input into a primary attention MIL network layer that combines the embedded modality representations into a clinical prediction.
[0075] Example 17. The method according to any one of the preceding embodiments, wherein the different modalities are input into an embedder network and the resulting embedded modality representations are input into a primary vision transformer network layer that combines the multi-modal raw data into a clinical prediction.
[0076] Example 18. A method according to any one of the preceding embodiments, wherein depending on the data type, each modality is converted into an embedded modality representation by a first-order attention MIL network layer or input into an embedder network, and the resulting embedded modality representations are input into a second-order attention MIL network that combines the embedded modality representations into a clinical prediction.
[0077] Example 19. A method according to any one of the preceding embodiments, wherein depending on the data type, each modality is converted into an embedded modality representation by a first attention MIL network layer or input into an embedder network, and the resulting embedded modality representations are input into a second-order vision transformer network that combines the embedded modality representations into a clinical prediction.
[0078] Example 20. A method according to any one of the preceding embodiments, comprising at least one pre-processing step, the pre-processing step comprising converting the raw data into a new format.
[0079] Example 21. A computer program comprising instructions, which when executed by a processing device, cause the processing device to perform steps i) to iii) of a method according to any one of the preceding embodiments.
[0080] Example 22. A computer-readable storage medium comprising instructions which, when executed by a processor, cause the processor to perform steps i) to iii) of a method according to any one of the preceding method embodiments.
[0081] Example 23. A clinical prediction device comprising at least one processing device having at least one communication interface configured to retrieve input data, the input data including a plurality of different modalities of a patient, the processing device configured to process the input data, the processing including generating an embedded modality representation from the input data by using at least one trainable data embedder, the processing including generating a clinical prediction by combining the embedded modality representations using at least one aggregation network, the aggregation network including at least one attention layer and / or at least one transformer layer, and the processing device configured to generate an output of the clinical prediction.
[0082] Example 24. A clinical prediction device according to the preceding embodiments, configured to perform a method for performing clinical prediction, the method being according to any one of the preceding embodiments.
[0083] BRIEF DESCRIPTION OF THE DRAWINGS Further optional features and embodiments are disclosed in more detail in the following description of the embodiments, preferably in conjunction with the dependent claims. In the disclosure, each optional feature may be realized in an independent manner as well as in any possible combination, as understood by a person skilled in the art. The scope of the present invention is not limited by the preferred embodiments. The embodiments are illustrated diagrammatically in the figures, where the same reference numbers in these figures refer to the same or functionally equivalent elements. [Brief description of the drawings]
[0084] [Figure 1] 2 illustrates an embodiment of the workflow of the method according to the invention. [Diagram 2] 13 shows a further exemplary workflow with three selected modalities for each patient. [Diagram 3] 1 shows an example of a fusion setup based on a mixed global multimodal transformer. [Figure 4] 1 shows the architecture of a visual encoder. [Figure 5A] 4 shows a further example of a method according to the invention. [Figure 5B] 4 shows a further example of a method according to the invention. [Figure 5C] 4 shows a further example of a method according to the invention. [Figure 5D] 4 shows a further example of a method according to the invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0085] FIG. 1 shows a general high-level workflow of a computer-implemented method for performing clinical prediction according to the present invention. In this embodiment, step i) 110 includes receiving multiple modalities for each patient. The modalities can be of different types and can carry different information, such as, for example, histological tissue images (e.g., H&E and / or IHC and / or fluorescent stained slide images), gene sequence data (e.g., RNA-Seq, mRNA), clinical data (e.g., tumor type, tissue type), etc. FIG. 1 shows a pre-processing step 112 that includes converting the original raw data into a new format. This step 112 is optional and depends on the type of modality. For example, by using a pre-trained embedder, it is possible to project the histological tile images into a different space. This can help to capture relevant information in the raw data and possibly speed up the training process. Step ii) 114 may include inputting each of the modalities into a trainable data embedder to obtain a meaningful embedded modality representation. In this step 114, the output can be either a generic patient-level embedding representation per modality or multiple instance embeddings for each modality, depending on the nature of the data and the embodiment selected. Then, in step 116, all these output embeddings are combined using an aggregation network. There are many aggregation options, e.g., average or max operators, attention networks, transformers, etc. Furthermore, step iii) 118 includes generating an output of clinical prediction.
[0086] FIG. 2 shows an example workflow with three selected modalities for each patient: histological whole slide images (WSI) 120, gene sequences 122, and clinical data 124. The method may include generating an embedded modality representation from each data point, combining the generated embedded modality representations for each of the modalities separately, and generating a clinical prediction from the combined embedded modality representation using an aggregation network. In this example, each modality goes through a series of steps 126 to generate an embedded modality representation. For WSI 120, high-resolution WSIs are typically very large and cannot fit in memory, so the method may include tiling the slide into non-overlapping patches. Each patch is then projected into a different space using a pre-trained embedder (e.g., Resnet). These tile-level embeddings are input into a trainable aggregation network (e.g., attention MIL) to generate patient-level embeddings of the WSI. Each of the gene sequences 122 and clinical data 124 is passed to a trainable embedder to generate a representative patient-level embedding for each modality. All patient-level embeddings are then passed to a second aggregate model 128 to generate an overall patient-level embedding and a final patient prediction.
[0087] Figure 3 shows an example of a fusion setup based on a mixed global multimodal transformer. In the upper part of Figure 3, a general example of the components used is shown, and in the lower part, an application is shown using three selected modalities for each patient: histological whole slide images (WSIs) 120, gene sequences 122, and clinical data 124. Modality A 129 (WSIs 120 at the bottom of Figure 3) may be passed through an embedding projection, e.g. an embedder 130, and a trainable attention network 132 to generate a patient-level representation 134. Modalities B 136 (clinical data 124 at the bottom of Figure 3) and C 138 (gene sequences 122 at the bottom of Figure 3) may be passed through trainable embedders 140, 142 to generate patient-level representations 144, 146. The patient-level representations 134, 144, 146 are then passed to a global multimodal vision transformer 148 to generate a multimodal patient representation 150 and a patient-level prediction 152.
[0088] Figure 4 shows the architecture of the transformer encoder as described in Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., …&Houlsby, N. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929.
[0089] 5A-5D show further examples of the method according to the invention for three modalities.
[0090] 5A shows received modality A 129, modality B 136, and modality C 138. Each of modality A 129, modality B 136, and modality C 138 may be passed through embedders 130, 140, 142 and trainable attention networks 132, 154, 156 to generate modality representations 134, 144, 146 for each modality. The modality representations 134, 144, 146 are then passed to a global multi-modality embedder 158 and a global modality attention network 160 to generate a multi-modal patient representation 150 and patient-level predictions.
[0091] 5B shows received modality A 129, modality B 136, and modality C 138. Each of modality A 129, modality B 136, and modality C 138 may be passed through embedders 130, 140, 142 and trainable attention networks 132, 154, 156 to generate modality representations 134, 144, 146 for each modality. The modality representations 134, 144, 146 are then passed to a global multimodal transformer 162 to generate a multimodal patient representation 150 and patient-level predictions.
[0092] 5C shows received modality A 129, modality B 136, and modality C 138. Each of modality A 129, modality B 136, and modality C 138 may pass through embedders 130, 140, 142 to generate modality representations 134, 144, 146 for each modality. The modality representations 134, 144, 146 are then passed to a global multimodal transformer 162 to generate a multimodal patient representation 150 and patient-level predictions.
[0093] 5D shows received modality A 129, modality B 136, and modality C 138. Each of modality A 129, modality B 136, and modality C 138 may pass through embedders 130, 140, 142 to generate modality representations 134, 144, 146 for each modality. The modality representations 134, 144, 146 are then passed to a global multi-modality embedder 158 and a global modality attention network 160 to generate a multi-modal patient representation 150 and a patient-level prediction.
[0094] As shown highly diagrammatically in Figures 5A-5D, the input data is received via at least one communication interface 164 of a processing device 166. The processing device 166 is configured to process the input data. The clinical prediction may be provided, for example, via an output interface 168 of the processing device 166 or via a further device such as a display and / or a printer. Figures 5A-5D further show an embodiment of a clinical prediction device 170 comprising the communication interface 164 and the processing device 166. The clinical prediction device 170 may further comprise an output interface 168. [Explanation of symbols]
[0095] 110 Step i) 112 Pre-processing step 114 Step ii) 116 combinations 118 Step iii) 120 Full slide image 122 Gene Sequence 124 Clinical Data 126 Steps for Generating Embedded Modality Representations 126 Generating Global Patient-Level Embeddings and Predictions 128 Aggregation Model 129 Modality A 130 Embedder 132 Trainable Attention Networks 134 Modality (Patient-Level) Representation 136 Modality B 138 Modality C 140 Embedder 142 Embedder 144 Modality (Patient-Level) Representation 146 Modality (Patient-Level) Representation 148 Global Multimodal Vision Transformer 150 Multimodal Patient Representations 152 Patient-level predictions 154 Attention Network 156 Attention Network 158 Global Multimodality Implant 160 Global Modality Attention Network 162 Global Multimodal Transformer 164 Communication Interface 166 Processing Equipment 168 Output Interface 170 Clinical Prediction Device
Claims
1. 1. A computer-implemented method for performing clinical predictions, comprising: i) retrieving (110) input data comprising a plurality of different modalities of a patient via at least one communication interface (164) of a processing device (166); ii) processing (114) said input data by using said processing device (166), the processing includes generating an embedded modality representation from the input data by using at least one trainable data embedder; the process includes generating the clinical prediction by combining the embedded modality representations using at least one aggregation network; the aggregation network includes at least one attention layer and / or at least one transformer layer; Steps and iii) generating (118) an output of said clinical prediction by using said processing device (166); A method comprising:
2. 10. The method of claim 1, wherein the output of the clinical prediction includes one or more of: information regarding medications for the patient; information regarding patient survival; information regarding response to at least one particular treatment; information confirming the patient's diagnosis; and information curating and / or completing patient data by predicting missing patient data points.
3. 3. The method of claim 1 or 2, wherein the method comprises: at least one output step including providing said clinical prediction via at least one output interface (168); The method wherein the output of the trainable data embedder is a generic patient-level embedding representation for each modality, or a multiple instance embedding for each modality.
4. 3. The method according to claim 1 or 2, The method, wherein the multiple modalities of the patient include one or more of at least one histological tissue image, at least one full-body microscopic image of a biopsy and / or surgical specimen, radiological images such as magnetic resonance imaging (MRI) and computed tomography (CT), genomic data, gene expression data, proteomics, patient clinical data, and demographics.
5. 3. The method according to claim 1 or 2, The method includes a step in which the attention layer and / or the transformer layer learns an optimal combination and / or attention strategy through backpropagation.
6. 3. The method according to claim 1 or 2, the input data comprises at least one data point from each of the different modalities, the method comprising generating an embedded modality representation from each of the data points, and using the aggregation network to generate the clinical prediction from the embedded modality representations of the different modalities; and / or the input data includes, for at least one of the different modalities, multiple data points from the same modality for a single patient; the method comprising generating an embedded modality representation from each data point, combining the generated embedded modality representations separately for each of the different modalities, and generating the clinical prediction from the combined embedded modality representations using the aggregation network; and / or The method includes combining the plurality of data points and generating a global embedded modality representation from the combined data points for each of the different modalities; and using the aggregation network to generate the clinical prediction from the global embedded modality representation. Including, method.
7. 3. The method according to claim 1 or 2, The method, wherein the different modalities are converted into embedded modality representations by a first-order attention multi-instance learning (MIL) network layer, and then input into a second-order attention MIL network that combines the embedded modality representations into a clinical prediction.
8. 3. The method according to claim 1 or 2, The different modalities are transformed into embedded modality representations by a first-order attention MIL network layer, and then input into a second-order vision transformer network that combines the embedded modality representations into a clinical prediction.
9. 3. The method according to claim 1 or 2, The method, wherein the different modalities are transformed into embedded modality representations by a first-order vision transformer network layer and then input into a second-order vision transformer network that combines the embedded modality representations into a clinical prediction.
10. 3. The method according to claim 1 or 2, The method, wherein the different modalities are transformed into embedded modality representations by a first-order vision transformer network layer, and then input into a second-order attention MIL network that combines the embedded modality representations into a clinical prediction.
11. 3. The method according to claim 1 or 2, The different modalities are input to an embedder network, and the resulting embedded modality representations are input to a primary attention MIL network layer that combines the embedded modality representations into a clinical prediction.
12. 3. The method according to claim 1 or 2, The different modalities are input to an embedder network and the resulting embedded modality representations are input to a primary vision transformer network layer that combines the multi-modal raw data into a clinical prediction.
13. 3. The method according to claim 1 or 2, Depending on the data type, each modality is either converted into an embedded modality representation by the first attention MIL network layer or input into the embedder network; The resulting embedded modality representations are input into a second-order attention MIL network that combines the embedded modality representations into clinical predictions.
14. 3. The method according to claim 1 or 2, Depending on the data type, each modality is either converted into an embedded modality representation by the first attention MIL network layer or input into the embedder network; The method, wherein the resulting embedded modality representations are input into a second-order vision transformer network that combines the embedded modality representations into a clinical prediction.
15. A clinical prediction device (170) comprising: at least one processing unit (166) having at least one communication interface (164) configured to retrieve input data; the input data includes a plurality of different modalities of the patient; the processing unit (166) is configured to process the input data; the processing includes generating an embedded modality representation from the input data by using at least one trainable data embedder; the process includes generating the clinical prediction by combining the embedded modality representations using at least one aggregation network; the aggregation network includes at least one attention layer and / or at least one transformer layer; The processing unit (166) generates an output of the clinical prediction. A clinical prediction device (170) configured to: