Medical image processing analysis based on machine learning using few-shot learning with task instructions

By encoding medical images and task instructions into features and using a machine learning-based task network for medical image processing analysis, the system addresses the challenge of requiring large annotated datasets, achieving efficient and accurate analysis with fewer examples.

DE102023211592A1Pending Publication Date: 2025-05-22SIEMENS HEALTHINEERS AG
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
DE102023211592
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-17
Filing Date
2023-11-21
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Conventional machine learning models for medical image processing analysis require a large amount of annotated medical images, which is time-consuming, costly, and prone to variations between annotators.

Method used

The system uses a machine learning-based task network that encodes medical input images and task instructions into image features and text features, respectively, and performs medical image processing analysis tasks based on these features, utilizing few shot learning with task instructions.

Benefits of technology

This approach reduces the need for extensive annotated datasets, allows for more efficient training with fewer examples, and incorporates contextual information from task instructions to improve analysis accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Systems and methods for performing a medical image processing analysis task using a machine-learning task network based on task instructions are provided. One or more medical input images of a patient and task instructions for performing a medical image processing analysis task are received. The one or more medical input images are encoded into image features using an image encoding network. The task instructions are encoded into text features using a text encoding network. The medical image processing analysis task is performed based on the image features and the text features using a machine-learning task network. Results of the medical image processing analysis task are output.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention relates generally to medical image processing analysis and, more particularly, to machine learning-based medical image processing analysis using few-shot learning with task instructions. STATE OF THE ART

[0002] Recently, machine learning models have been proposed for performing various medical image processing analysis tasks, such as segmentation, registration, classification, detection, diagnosis, etc. Typically, such machine learning models are trained using supervised learning to perform a specific medical image processing analysis task. Supervised learning requires a large amount of task-specific medical images with expert annotations. However, obtaining such a large amount of annotated medical images is time-consuming, costly, and subject to inter-annotator variability. SUMMARY OF THE INVENTION

[0003] In accordance with one or more embodiments, systems and methods are provided for performing a medical image processing analysis task using a machine learning-based task network based on task instructions. One or more input medical images of a patient and task instructions for performing a medical image processing analysis task are received. The one or more input medical images are encoded into image features using an image encoding network. The task instructions are encoded into text features using a text encoding network. The medical image processing analysis task is performed based on the image features and the text features using a machine learning-based task network. Results of the medical image processing analysis task are output.

[0004] In one embodiment, the task instructions include references to image regions in at least one of the one or more medical input images. The task instructions may include anatomical knowledge and task knowledge. The task knowledge may include at least one of a description of an anatomical anomaly, the manner in which the anatomical anomaly is represented, and the manner in which the anatomical anomaly may be detected in the one or more medical input images. The task instructions may be user-defined.

[0005] In one embodiment, the text features and the image features are aligned in the same latent space.

[0006] In one embodiment, the machine learning-based task network is trained using self-supervised learning based on unannotated medical training images and text. In one embodiment, the machine learning-based task network is trained using few-shot learning using annotated medical training images and annotated task descriptions. In one embodiment, the machine learning-based task network comprises a large language model (LLM)-based task network.

[0007] In one embodiment, text-based medical data of the patient is received. The task instructions and the text-based medical data are encoded into the text features using the text encoding network.

[0008] In accordance with one or more embodiments, systems and methods are provided for training a machine learning-based task network to perform a medical image processing analysis task based on task instructions. One or more training medical images and training task instructions for performing a medical image processing analysis task are received. The one or more medical training images are encoded into image features using a pre-trained image encoding network. The training task instructions are encoded into text features using a pre-trained text encoding network. A machine learning-based task network is trained to perform the medical image processing analysis task based on the image features and the text features. The trained machine learning-based task network is output.

[0009] In one embodiment, the machine learning-based task network is trained using self-supervised learning based on unannotated medical training images and text. In one embodiment, the machine learning-based task network is trained using few-shot learning using annotated medical training images and annotated task descriptions.

[0010] These and other advantages of the invention will be apparent to those of ordinary skill in the art upon reference to the following detailed description and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1 shows a method for performing a medical image processing analysis task using a machine learning model based on task instructions according to one or more embodiments; Fig. 2 shows exemplary task instructions for detecting pneumothorax in chest x-ray images according to one or more embodiments; Fig. 3 shows a method for training a machine learning-based task network to perform a medical image processing analysis task based on task instructions according to one or more embodiments; Fig. 4 shows a workflow for training a machine learning-based task network to perform a medical image processing analysis task based on training task instructions according to one or more embodiments; Fig. 5 shows an exemplary artificial neural network that may be used to implement one or more embodiments; Fig. 6 shows a convolutional neural network that may be used to implement one or more embodiments; Fig. 7 shows a schematic structure of a recurrent machine learning model that can be used to implement one or more embodiments; and Fig. Figure 8 shows a high-level block diagram of a computer that may be used to implement one or more embodiments. DETAILED DESCRIPTION

[0011] The present invention generally relates to methods and systems for machine learning-based medical image processing analysis using few-shot learning with task instructions. Embodiments of the present invention are described herein to provide a visual understanding of such methods and systems. A digital image often consists of digital representations of one or more objects (or shapes). The digital representation of an object is often described herein in terms of identifying and manipulating the objects. Such manipulations are virtual manipulations achieved in memory or other circuitry / hardware of a computer system. Accordingly, it is understood that embodiments of the present invention may be performed within a computer system using data stored within the computer system.Furthermore, a reference herein to pixels of an image may equally refer to voxels of an image and vice versa.

[0012] Conventional machine learning models are trained to perform medical image processing analysis tasks through supervised learning using a large set of annotated medical images. Such a large set of annotated medical images is required in supervised learning to implicitly encode all the details and all possible variations for performing medical image processing analysis tasks through the examples.

[0013] According to embodiments described herein, machine learning models for performing medical image processing analysis tasks are trained using few-shot learning with task instructions. The task instructions provide a "recipe" about how to perform the medical image processing analysis tasks. By encoding the task instructions and medical images, the text encoding of the task instruction is used to draw conclusions about tasks, thereby eliminating the complexity of scaling the amount of annotated medical images needed for supervised learning, as well as using more context-dependent information provided by the task instructions.

[0014] Fig. 1 shows a method 100 for performing a medical image processing analysis task using a machine learning model based on task instructions according to one or more embodiments. The steps of the method 100 may be performed by one or more suitable computing devices, such as a computer 902 of Fig. 9, be carried out.

[0015] At step 102 of Fig. 1 receives 1) one or more medical input images of a patient and 2) task instructions for performing a medical image processing analysis task.

[0016] The one or more medical input images may depict any anatomical object of interest of the patient, such as organs, vessels, tumors, abnormalities, etc. The one or more medical input images may comprise any suitable modality or modalities, such as CT (computed tomography), MRI (magnetic resonance imaging), US (ultrasound), X-ray, or any other medical imaging modality or combination of medical imaging modalities. The one or more medical input images may include 2D images (two-dimensional images) and / or 3D volumes (three-dimensional volumes).

[0017] The task instructions act as a "recipe" or instructions on how to perform the medical image processing analysis task. The task instructions may include anatomical knowledge, task knowledge, and / or any other suitable information that can be used to perform the medical image processing analysis task. In one embodiment, the task instructions are associated with the one or more medical input images. For example, the task instructions may include explicit primers or references to anatomical image regions and / or task-specific image regions in at least one of the one or more medical input images. In one embodiment, the task instructions are custom task instructions generated by a user. However, the task instructions may be generated using any other suitable approach.Example task instructions are in . Fig. 2. The task knowledge may include, for example, a description of an anatomical anomaly, the way the anatomical anomaly is represented, the way the anatomical anomaly can be detected in the one or more medical input images, etc.

[0018] Fig. 2 shows exemplary task instructions 202 for detecting pneumothorax in chest x-ray images according to one or more embodiments. As in Fig. 2, the task instructions 202 include anatomical knowledge and task knowledge. The anatomical knowledge describes anatomical knowledge of the pleura and pleural spaces. The task knowledge describes the radiological features of pleural disease, pneumothorax, and pleural thickening. The task instructions 202 include primers or references 206-A, 206-B, and 206-C (collectively, primers 206) to anatomical image regions and / or task-specific image regions in the medical input images 204-A, 204-B, and 204-C (collectively, medical input images 204). In particular, the primer 206 primes or references the pleura and pleural spaces described in the task instructions 202 with anatomical regions of the pleura and pleural spaces shown in the medical input image 204-A.Primer 206-B primes or references the radiological features of the pneumothorax ("visible pleural edge" and "lung markings not visible beyond that edge") in the task instructions 202 with anatomical regions shown in the input medical image 204-B. Primer 206-C primes or references the radiological features of pleural thickening ("shading over the entire right lung") in the task instructions 202 with anatomical regions shown in the input medical image 204-C. In one example, the input medical images 204 and the task instructions 202 are the one or more input medical images and the images obtained in step 102 of FIG. Fig. 1 received task instructions.

[0019] Back on Fig. 1, in one embodiment, non-imaging medical data of the patient may also be received at step 102. Such non-imaging medical data may include, for example, text-based medical data such as radiology reports, lab reports, medical records, imaging / exam notes, demographic information, administrative data, etc.

[0020] The one or more medical input images and task instructions (and non-imaging medical data) may be obtained, for example, by loading the medical input data, task instructions, and / or non-imaging medical data from a storage or memory of one or more computer systems and / or receiving the medical input images, task instructions, and / or non-imaging medical data from one or more remote computer systems. Such computer systems may include an EHR (electronic health record), EHR (electronic patient record), EHR (electronic case record), HIS (health information system), RIS (radiology information system), PACS (image storage and communication system), LIMS (laboratory information management system), or any other suitable database or system.In some embodiments, the input medical images may be received directly from an image acquisition device, such as a CT scanner, while the medical images are being acquired.

[0021] At step 104 of Fig. 1, the one or more medical input images are encoded into image features using an image coding network. The image coding network may be a machine learning-based image coding network, such as an autoencoder, a VAE (Variational Autoencoder), or may be implemented in accordance with any other suitable machine learning-based architecture. The image coding network receives the one or more medical input images as input and generates the image features as output. The image features are embeddings that represent low-dimensional, dense vector representations of the relatively higher-dimensional medical input images in a latent space. The image coding network is pre-trained during a previous offline or training phase using a large set of training images.Once the image coding network is trained, it is applied during an online or inference phase, for example, to perform step 104 of . Fig. 1 to be carried out.

[0022] At step 106 of Fig. 1, the task instructions are encoded into text features using a text encoding network. In one embodiment, if at step 102 of Fig. 1 If non-imaging medical data of the patient is also received, the non-imaging medical data (e.g., the text-based medical data) is encoded into the image features using the image coding network with the task instructions.

[0023] In one embodiment, the text encoding network is an LLM-based (Large Language Model-based) text encoding network. However, the text encoding network may be implemented according to any suitable machine learning-based architecture. The LLM-based text encoding network may be any suitable pre-trained, deep learning-based LLM. For example, the LLM-based text encoding network may be based on the Transformer architecture, which uses a self-attention mechanism to capture long-range dependencies in text. An example of a transformer-based architecture is GPT (Generative Pre-Training Transformer), which has a multi-layer transformer-decoder architecture that can be pre-trained to optimize the next token prediction task and then fine-tuned with labeled data for various downstream tasks.GPT-based LLMs can be trained using reinforcement learning with human feedback to perform various linguistic computing tasks. Other example transformer-based architectures include BLOOM (BigScience Large Open-science Open-access Multilingual Language Model) and BERT (Bidirectional Encoder Representations from Transfers).

[0024] In one embodiment, the LLM-based text coding network is restricted to a specific medical domain. For example, the LLM-based text coding network may be restricted for the use case of detecting pneumothorax in chest x-ray images. To restrict the LLM-based text coding network, the LLM-based text coding network may be updated (e.g., trained, retrained, or fine-tuned) using, for example, clinical data and data extracted from medical images using AI-based systems. Such extracted data may include, for example, clinical measurements (e.g., diameters, volumes, distances, etc.), anatomical locations, detections, etc.

[0025] The LLM-based text encoding network receives the task instructions (and possibly the non-imaging medical data) as input (e.g., as one or more prompts) and generates the text features as output. In one embodiment, the task instructions include specific prompts for the LLM-based text encoding network (and possibly the LLM-based task network generated at step 108 of Fig. 1) to perform the task. Alternatively, at least some of the task instructions may be separately / directly incorporated into the machine learning-based task network (which is used at step 108 of Fig. 1 is used). The text features are embeddings that represent low-dimensional, dense vector representations of the relatively higher-dimensional task instructions (and possibly the non-imaging medical data) in the latent space. The LLM-based text encoding network is pre-trained during a previous offline or training phase using a large amount of text-based training data. Once the LLM-based text encoding network is trained, it is applied during an online or inference phase, for example, to perform step 106 of Fig. 1 to be carried out.

[0026] The text features and the image features are represented in the same latent space. The latent space can be refined to match corresponding image features and text features such that similar features represent similar concepts. In one embodiment, the text and image features in the latent space can be matched by pre-training the image coding network and text coding network on image-text pairs (such as corresponding image reports or other semantically annotated image / image region-text pairs). Pre-training can be performed, for example, by minimizing the distance in the latent space between corresponding pairs and maximizing the distance between non-corresponding pairs. By matching the image features and the text features, inference can be performed in the same latent feature space using interchangeable features (i.e., either the text features or the image features).

[0027] At step 108 of Fig. 1, the medical image processing analysis task is performed based on the image features and the text features using a machine learning-based task network. The medical image processing analysis task may include any suitable medical image processing analysis task, such as segmentation, registration, classification, detection, diagnosis, medical data summarization, etc.

[0028] In one embodiment, the machine learning-based task network is an LLM-based task network. However, the machine learning-based task network may be implemented in accordance with any suitable machine learning-based architecture. The LLM-based task network may be any suitable pre-trained deep learning-based LLM, such as a transformer-based network (e.g., GPT-based LLMs, BMOOM, BERT). In one embodiment, the LLM-based task network is restricted to a specific medical domain (e.g., detecting pneumothorax in chest x-ray images).

[0029] The LLM-based task network receives the image features and the text features as input (e.g., as one or more prompts) and generates results of the medical image processing analysis task as output. The LLM-based task network also receives instructions for performing the medical image processing analysis task. The instructions may be received indirectly via the text encoding network and / or the image encoding network or directly as separate task instructions. In one embodiment, the LLM-based task network may include a plurality of heads, each for performing a respective medical image processing analysis task (e.g., text decoding, classification evaluation, segmentation, recognition, etc.). The LLM-based task network is trained during a previous offline or training phase.For example, the LLM-based task network can be trained using a relatively small set of training images and annotated task instructions using few-shot learning and self-supervised learning, as described in more detail below with respect to . Fig. 3 and Fig. 4. Once the LLM-based task network is trained, it is applied during an online or inference phase, for example, to perform step 108 of Fig. 1 to be carried out.

[0030] At step 110 of Fig. 1, results of the medical image processing analysis task are output. For example, the results of the medical image processing analysis task can be output by displaying the results of the medical image processing analysis task on a display device of a computer system, storing the results of the medical image processing analysis task in a memory or storage of a computer system, or sending the results of the medical image processing analysis task to a remote computer system.

[0031] In one embodiment, the (at step 102 of Fig. 1) task instructions include verification steps to be performed by the final system (for self-verification) to ensure consistency of information for increased operational robustness, which can be combined with well-known uncertainty estimation techniques.

[0032] In one embodiment, additional instructions may be received from a user via an interactive user interface. The additional instructions may be parsed and validated based on the one or more medical input images and current knowledge. The interactive user interface may be used to query the user to explain the current medical image processing analysis task. This may be used for improved workflow efficiency using LLMs.

[0033] Fig. 3 shows a method 300 for training a machine learning-based task network to perform a medical image processing analysis task based on training task instructions, according to one or more embodiments. The steps of method 300 may be performed by one or more suitable computing devices, such as a computer 902 of Fig. 9, be carried out. Fig. 4 shows a workflow 400 for training a machine learning-based task network to perform a medical image processing analysis task based on training task instructions according to one or more embodiments. Fig. 3 and Fig. 4 are described together. The steps of method 300 of Fig. 3 and workflow 400 of Fig. 4 are performed during an offline or training phase to train the machine learning-based task network. Once the machine learning-based task network is trained, it is applied during an online or inference phase to perform the medical image processing analysis task, e.g., at step 108 of Fig. 1, to be carried out.

[0034] At step 302 of Fig. 3, 1) one or more medical training images and 2) training task instructions for performing a medical image processing analysis task are received. In one embodiment, non-imaging medical training data (e.g., text-based medical data) may also be received at step 302. In one example, as in workflow 400 of Fig. 4, the one or more medical training images are medical training images 404 selected from a data lake 402, and the training task instructions are training task instructions 406.

[0035] The one or more medical training images may depict any anatomical item of interest of a patient and may comprise any suitable modality or modalities, such as CT, MRI, US, X-ray, or any other medical imaging modality or combination of medical imaging modalities. The one or more medical training images may include 2D images and / or 3D volumes. The one or more medical training images may include one or more annotated medical images (e.g., annotated by a user) and one or more unannotated medical images.

[0036] The training task instructions may include anatomical knowledge, task knowledge, and / or any other suitable information that can be used to perform the medical image processing analysis task. During the training phase, the training task instructions are associated with the one or more medical training images. For example, the training task instructions may include explicit primers or references to anatomical image regions and / or task-specific image regions in at least one of the one or more medical training images. The primers are used to align embeddings of corresponding image-text concepts in latent space. The training task instructions may be extracted from textbooks, publications, websites, etc., or may be user-defined.

[0037] The one or more medical training images and training task instructions (and non-imaging medical training data) may be obtained, for example, by loading the medical training data, training task instructions, and / or non-imaging medical training data from a storage or memory of one or more computer systems and / or receiving the medical training images, training task instructions, and / or non-imaging medical training data from one or more remote computer systems. In some embodiments, the medical training images may be received directly from an image acquisition device while the medical images are being acquired.

[0038] At step 304 of Fig. 3, the one or more medical training images are encoded into image features using a pre-trained image coding network. The pre-trained image coding network may, for example, be an autoencoder, a VAE, or may be implemented in accordance with any other suitable machine learning-based architecture. The pre-trained image coding network receives the one or more medical training images as input and generates the image features as output. The pre-trained image coding network is pre-trained during a previous offline or training phase using a large set of training images. Once the pre-trained image coding network is trained, it is applied during an online or inference phase, for example, to perform step 304 of Fig. 3. In an example, as in workflow 400 of Fig. 4, the pre-trained image coding network is an image AI (artificial intelligence) encoder 408 that receives the medical training images 404 as input and generates image features 410 as output.

[0039] At step 306 of Fig. 3, the training task instructions are encoded into text features using a pre-trained text encoding network. In one embodiment, when at step 302 of Fig. 3 If non-imaging medical training data is also received, the non-imaging medical training data (e.g., the text-based medical data) is encoded into the image features using the pre-trained image encoding network with the training task instructions.

[0040] In one embodiment, the text coding network is an LLM-based text coding network. However, the text coding network may be implemented in accordance with any suitable machine learning-based architecture. In one embodiment, the LLM-based text coding network is restricted to a specific medical domain. The LLM-based text coding network receives the training task instructions (and possibly the non-imaging medical training data) as input and generates the text features as output. The image features 410 and the text features 414 are aligned within the latent space. The LLM-based text coding network is pre-trained during a previous offline or training phase using a large amount of text-based training data. Once the LLM-based text coding network is trained, it is applied during an online or inference phase, for example, to perform step 306 of Fig. 3 to be carried out.

[0041] In an example, as in workflow 400 of Fig. 4, the pre-trained text encoding network is an LLM text AI encoder 412 that receives the training task instructions 406 as input and generates the text features 414 as output.

[0042] At step 308 of Fig. 3, a machine learning-based task network is trained to perform the medical image processing analysis task based on the image features and the text features. In one embodiment, the machine learning-based task network is an LLM-based task network. However, the machine learning-based task network may be implemented in accordance with any suitable machine learning-based architecture. In one embodiment, the LLM-based task network is restricted to a specific medical domain (e.g., detecting pneumothorax in chest x-ray images).

[0043] The machine learning-based task network receives image features and text features as input and generates results of the medical image processing analysis task as output. The machine learning-based task network is trained using few-shot learning and self-supervised learning.

[0044] In few-shot learning, the machine learning task network is trained to accurately perform the medical image processing analysis task using one or more annotated training images and annotated task descriptions. Few-shot learning applies metalearning so that the machine learning task network learns to learn. During the metatraining phase, the machine learning task network is trained on a relatively small number of related tasks. During the metatesting phase, the machine learning task network can generalize to unfamiliar (but related) tasks.

[0045] In self-supervised learning, the machine learning-based task network is trained using the one or more unannotated medical training images and text. The matched image and text features enable the text features to be used to generate task instructions and pseudolabels for the one or more unannotated medical training images and text.

[0046] In one embodiment, multiple output "heads" can be decoded in a generative framework. Example decoding heads can include text decoding, classification evaluation, segmentation, detection, etc. When using the LLM, decoding is performed implicitly as text. Training for this can be performed, for example, by sequentially predicting the same tokens from baseline knowledge (e.g., using a cross-entropy loss function) or by reinforcement learning (with / without human feedback), with some text decodings being more preferred than others.

[0047] In an example, as in workflow 400 of Fig. 4, the machine learning-based task network is an RL-LLM 416 that receives the image features 410 and the text features 414 as input and generates results 418 of the medical image processing analysis task as output.

[0048] At step 310 of Fig. 3, the trained machine learning-based task network is output. For example, the trained machine learning-based task network may be output by storing the trained machine learning-based task network in a memory or storage of a computer system or by sending the trained machine learning-based task network to a remote computer system. In one example, the trained machine learning-based task network is applied to, for example, step 108 of Fig. 1 to be carried out.

[0049] Embodiments described herein are described with respect to both the claimed systems and the claimed methods. Features, advantages, or alternative embodiments herein may be assigned to the other claimed subject matter, and vice versa. In other words, claims and embodiments for the systems may be enhanced with features described or claimed in the context of the respective methods. In this case, the functional features of the method are implemented by physical units of the system.

[0050] Furthermore, certain embodiments described herein are described with respect to methods and systems employing trained machine learning models, as well as with respect to methods and systems for providing trained machine learning models. Features, advantages, or alternative embodiments herein may be assigned to the other claimed subject matter, and vice versa. In other words, claims and embodiments for providing trained machine learning models may be enhanced with features described or claimed in the context of employing trained machine learning models, and vice versa.In particular, data sets used in the methods and systems for deploying trained machine learning models may have the same properties and characteristics as the corresponding data sets used in the methods and systems for providing trained machine learning models, and the trained machine learning models provided by the respective methods and systems may be used in the methods and systems for deploying the trained machine learning models.

[0051] In general, a trained machine learning model mimics cognitive functions that humans associate with a different kind of human mind. Specifically, through training based on training data, the machine learning model can adapt to new situations and recognize and extrapolate patterns. Another term for "trained machine learning model" is "trained function."

[0052] In general, the parameters of a machine learning model can be adjusted using training. In particular, supervised training, semi-supervised training, unsupervised training, reinforcement learning, and / or active learning can be used. Furthermore, representation learning (an alternative term is "feature learning") can be used. In particular, the parameters of the machine learning model can be adjusted iteratively through multiple training steps. In particular, certain cost functions can be minimized during training. In particular, the error feedback algorithm can be used during the training of a neural network.

[0053] In particular, a machine learning model, such as the image coding network used in step 104, the text coding network used in step 106 and the one used in step 108 of Fig. 1, the pre-trained image coding network used in step 304, the pre-trained text coding network used in step 306 and the pre-trained text coding network used in step 308 of Fig. 3, as well as the Image AI Encoder 408, the LLM Text AI Encoder 412 and the RL-LLM 416 of Fig. 4, for example, a neural network, a support vector machine, a decision tree, and / or a Bayesian network, and / or the machine learning model can be based, for example, on k-means clustering, Q-learning, genetic algorithms, and / or association rules. In particular, a neural network can be, for example, a deep neural network, a convolutional neural network, or a deep convolutional neural network. Furthermore, a neural network can be, for example, an adversarial network, a deep adversarial network, and / or a generative adversarial network.

[0054] Fig. 5 illustrates one embodiment of an artificial neural network 500 that may be used to implement one or more machine learning models described herein. Alternative terms for "artificial neural network" are "neural network," "artificial neural network," or "neural network."

[0055] The artificial neural network 500 comprises nodes 520, ..., 532 and edges 540, ..., 542, where each edge 540, ..., 542 is a directed connection from a first node 520, ..., 532 to a second node 520, ..., 532. In general, the first node 520, ..., 532 and the second node 520, ..., 532 are different nodes 520, ..., 532, but it is also possible that the first node 520, ..., 532 and the second node 520, ..., 532 are identical. Fig. 5, for example, edge 540 is a directed connection from node 520 to node 523, and edge 542 is a directed connection from node 530 to node 532. An edge 540, ..., 542 from a first node 520, ..., 532 to a second node 520, ..., 532 is also referred to as an "incoming edge" for the second node 520, ..., 532 and an "outgoing edge" for the first node 520, ..., 532.

[0056] In this embodiment, the nodes 520, ..., 532 of the artificial neural network 500 can be arranged in layers 510, ..., 513, where the layers can include an intrinsic ordering introduced by the edges 540, ..., 542 between the nodes 520, ..., 532. In particular, the edges 540, ..., 542 can only exist between adjacent node layers. In the embodiment shown, there is an input layer 510 that includes only nodes 520, ..., 522 without an incoming edge, an output layer 513 that includes only nodes 531, 532 without outgoing edges, and hidden layers 511, 512 between the input layer 510 and the output layer 513. In general, the number of hidden layers 511, 512 can be chosen arbitrarily. The number of nodes 520, ..., 522 within the input layer 510 is typically associated with the number of input values ​​of the neural network, and the number of nodes 531, 532 within the output layer 513 is typically associated with the number of output values ​​of the neural network.

[0057] In particular, each node 520, ..., 532 of the neural network 500 can be assigned a (real) number as a value. Here, x denotes (n) i the value of the i-th node 520, ..., 532 of the n-th layer 510, ..., 513. The values ​​of the nodes 520, ..., 522 of the input layer 510 are equivalent to the input values ​​of the neural network 500, the values ​​of the nodes 531, 532 of the output layer 513 are equivalent to the output values ​​of the neural network 500. Furthermore, each edge 540, ..., 542 can include a weight that is a real number, in particular, the weight is a real number within the interval [-1, 1] or within the interval [0, 1]. Here, w denotes(m,n) i,j the weight of the edge between the i-th node 520, ..., 532 of the m-th layer 510, ..., 513 and the j-th node 520, ..., 532 of the n-th layer 510, ..., 513. Furthermore, the abbreviation w (n) i,j for the weight w (n, n+1) i,j defined.

[0058] In particular, to calculate the output values ​​of the neural network 500, the input values ​​are propagated through the neural network. Specifically, the values ​​of the nodes 520, ..., 532 of the (n+1)th layer 510, ..., 513 can be calculated based on the values ​​of the nodes 520, ..., 532 of the nth layer 510, ..., 513 as follows: x(n+1)j=f(∑ix(n)i⋅w(n)i,j).

[0059] Here, the function f is a transfer function (another term is "activation function"). Common transfer functions include step functions, sigmoid functions (e.g., the logistic function, the generalized logistic function, the hyperbolic tangent, the arctangent function, the error function, the smoothstep function), or rectifier functions. The transfer function is primarily used for normalization purposes.

[0060] In particular, the values ​​are propagated layer by layer through the neural network, where values ​​of the input layer 510 are given by the input of the neural network 500, where values ​​of the first hidden layer 511 can be calculated based on the values ​​of the input layer 510 of the neural network, where values ​​of the second hidden layer 512 can be calculated based on the values ​​in the first hidden layer 511, etc.

[0061] To get the values ​​w (m,n) i,jfor the edges, the neural network 500 must be trained using training data. In particular, training data includes training input data and training output data (as t i For a training step, the neural network 500 is applied to the training input data to generate computed output data. Specifically, the training data and the computed output data comprise a number of values, where the number is equal to the number of nodes of the output layer.

[0062] In particular, a comparison between the calculated output data and the training data is used to recursively adjust the weights within the neural network 500 (error feedback algorithm). In particular, the weights in Changed to match the following: where y is a learning rate and the digits δ (n) j can be calculated recursively as follows: based on δ (n+1) j , if the (n+1)th layer is not the output layer, and if the (n+1)-th layer is the output layer 513, where f' is the first derivative of the activation function and t (n+1) j is the comparison training value for the j-th node of the output layer 513.

[0063] A convolutional neural network is a neural network that uses a convolution operation instead of general matrix multiplication in at least one of its layers (called a "convolutional layer"). Specifically, a convolutional layer performs a dot product of one or more convolutional kernels on the input data / image of the convolutional layer, where the entries of the one or more convolutional kernels are the parameters or weights that are adjusted through training. In particular, the Frobenius dot product and the ReLU activation function may be used. A convolutional neural network may include additional layers, such as pooling layers, fully connected layers, and normalization layers.

[0064] Using convolutional neural networks, input images can be processed very efficiently, as a convolution operation based on different kernels can extract different image features. By adjusting the weights of the convolution kernel, the relevant image features can be found during training. Furthermore, by sharing the weights in the convolution kernels, fewer parameters need to be trained, preventing overfitting during the training phase and allowing for faster training or having more layers in the network, thus improving the network's performance.

[0065] Fig. 6 shows one embodiment of a convolutional neural network 600 that can be used to implement one or more machine learning models described herein. In the embodiment shown, the convolutional neural network 600 includes an input node layer 610, a convolutional layer 611, a pooling layer 613, a fully connected layer 614, and an output node layer 616, as well as hidden node layers 612, 614. Alternatively, the convolutional neural network 600 may include multiple convolutional layers 611, multiple pooling layers 613, and multiple fully connected layers 615, as well as other types of layers. The order of the layers can be chosen arbitrarily; typically, fully connected layers 615 are used as the last layers before the output layer 616.

[0066] In particular, nodes 620, 622, 624 of a node layer 610, 612, 614 within a convolutional neural network 600 can be viewed as being arranged in the form of a d-dimensional matrix or a d-dimensional image. In particular, in the two-dimensional case, the value of node 620, 622, 624 indexed by i and j in the nth node layer 610, 612, 614 can be denoted as x(n)[i, j]. However, the arrangement of nodes 620, 622, 624 of a node layer 610, 612, 614 has no effect on the computations performed within the convolutional neural network 600 as such, since these are determined exclusively by the structure and weights of the edges.

[0067] A convolutional layer 611 is a connecting layer between a front node layer 610 (with node values ​​x(n-1)) and a back node layer 612 (with node values ​​x(n)). In particular, a convolutional layer 611 is characterized by the structure and weights of the incoming edges, which form a convolution operation based on a certain number of kernels. In particular, the structure and weights of the edges of the convolutional layer 611 are chosen such that the values ​​x(n) of the nodes 622 of the back node layer 612 are calculated as a convolution x(n) = K * x(n-1) based on the values ​​x(n-1) of the nodes 620 of the front node layer 610, where the convolution * is defined as follows in the two-dimensional case:

[0068] Here, the kernel K is a d-dimensional matrix (in this embodiment, a two-dimensional matrix) that is typically small compared to the number of nodes 620, 622 (e.g., a 3x3 matrix or a 5x5 matrix). In particular, this means that the weights of the edges in the convolution layer 611 are not independent, but are chosen to generate the convolution equation. In particular, for a kernel that is a 3x3 matrix, there are only 9 independent weights (where each entry of the kernel matrix corresponds to an independent weight), regardless of the number of nodes 620, 622 in the front node layer 610 and the back node layer 612.

[0069] Generally, convolutional neural networks 600 use node layers 610, 612, 614 with a plurality of channels, particularly due to the use of a plurality of kernels in the convolutional layers 611. In these cases, the node layers can be considered as (d+1)-dimensional matrices (where the first dimension indexes the channels). The effect of a convolutional layer 611 is then a two-dimensional example, defined as follows: x(n)b[i, j]=∑aKa,b*x(n−1)a[i, j]=∑a∑i'∑j'Ka,b[i', j']⋅x(n−1)a[i−i', j−j'], where x (n-1)a corresponds to the a-th channel of the front node layer 610, x (n)b corresponds to the b-th channel of the back node layer 612 and K a,b corresponds to one of the kernels. If a convolutional layer 611 acts on a front node layer 610 with A channels and outputs a back node layer 612 with B channels, there are A · B independent d-dimensional kernels K a,b .

[0070] Activation functions are generally used in convolutional neural networks 600. In this embodiment, ReLU (Rectified Linear Units) is used, where R(z) = max(0, z), so that in the two-dimensional example, convolutional layer 611 acts as follows: x(n)b[i, j]=R(∑a(Ka,b*x(n−1)a)[i,j]) =R(∑a∑i'∑j'Ka,b[i', j']⋅x(n−1)a[i−i', j−j'])

[0071] It is also possible to use other activation functions, such as ELU (abbreviation for “exponential linear unit”), LeakyReLU, Sigmoid, Tanh or Softmax.

[0072] In the embodiment shown, the input layer 610 comprises 36 nodes 620 arranged as a two-dimensional 6x6 matrix. The first hidden node layer 612 comprises 72 nodes 622 arranged as two two-dimensional 6x6 matrices, each of the two matrices being the result of convolving the input layer values ​​with a 3x3 kernel within the convolutional layer 611. Similarly, the nodes 622 of the first hidden node layer 612 can be interpreted as arranged as a three-dimensional 2x6x6 matrix, where the first dimension corresponds to the channel dimension.

[0073] The advantage of using convolutional layers 611 is that a spatially local correlation of the input data can be exploited by enforcing a local connectivity pattern between nodes of neighboring layers, in particular each node being connected only to a small range of the nodes of the previous layer.

[0074] A pooling layer 613 is a connecting layer between a front node layer 612 (with node values ​​x(n-1)) and a back node layer 614 (with node values ​​x(n)). In particular, a pooling layer 613 can be characterized by the structure and weights of the edges and the activation function, which form a pooling operation based on a nonlinear pooling function f. For example, in the two-dimensional case, the values ​​x(n) of the nodes 624 of the back node layer 614 can be calculated based on the values ​​x(n-1) of the nodes 622 of the front node layer 612 as follows: x(n)b[i, j]=f(x(n−1)[id1, jd2], …, x(n−1)b[(i+1)d1−1, (j+1)d2−1])

[0075] In other words, by using a pooling layer 613, the number of nodes 622, 624 can be reduced by replacing a number d1 d2 of neighboring nodes 622 in the front node layer 612 with a single node 622 in the back node layer 614, which is calculated as a function of the values ​​of the number of neighboring nodes. In particular, the pooling function f can be the max function, the mean, or the L2 norm. In particular, for a pooling layer 613, the weights of the incoming edges are fixed and are not modified by training.

[0076] The advantage of using a pooling layer 613 is that the number of nodes 622, 624 and the number of parameters are reduced. This reduces the amount of computation in the network and regulates overfitting.

[0077] In the illustrated embodiment, pooling layer 613 is a max-pooling layer that replaces four neighboring nodes with only one node, where the value is the maximum of the values ​​of the four neighboring nodes. Max-pooling is applied to each d-dimensional matrix of the previous layer; in this embodiment, max-pooling is applied to each of the two two-dimensional matrices, reducing the number of nodes from 72 to 18.

[0078] In general, the last layers of a convolutional neural network 600 are fully connected layers 615. A fully connected layer 615 is a connecting layer between a front node layer 614 and a back node layer 616. A fully connected layer 613 can be characterized in that a majority, in particular all edges, are present between the nodes 614 of the front node layer 614 and the nodes 616 of the back node layer, and wherein the weight of each of these edges can be individually adjusted.

[0079] In this embodiment, the nodes 624 of the front node layer 614 of the fully connected layer 615 are shown both as two-dimensional matrices and additionally as unrelated nodes (shown as a row of nodes, with the number of nodes reduced for clarity). This operation is also referred to as "flattening." In this embodiment, the number of nodes 626 in the back node layer 616 of the fully connected layer 615 is less than the number of nodes 624 in the front node layer 614. Alternatively, the number of nodes 626 may be equal to or greater.

[0080] Furthermore, in this embodiment, the softmax activation function is used within the fully connected layer 615. By applying the softmax function, the sum of the values ​​of all nodes 626 of the output layer 616 is 1, and all values ​​of all nodes 626 of the output layer 616 are real numbers between 0 and 1. In particular, if the convolutional neural network 600 is used to categorize input data, the values ​​of the output layer 616 can be interpreted as the probability that the input data falls into one of the various categories.

[0081] In particular, the convolutional neural network 600 can be trained based on the error feedback algorithm. To prevent overfitting, regularization techniques can be used, e.g., dropping nodes 620, ..., 624, stochastic pooling, using artificial data, weight decay based on the L1 or L2 norm, or maximum norm constraints.

[0082] According to one aspect, the machine learning model may comprise one or more residual networks (ResNet). In particular, a ResNet is an artificial neural network comprising at least one jump or skip connection used to skip at least one layer of the artificial neural network. In particular, a ResNet may be a convolutional neural network comprising one or more skip connections, each skipping one or more convolutional layers. According to some examples, the ResNets may be represented as m-layer ResNets, where m is the number of layers in the respective architecture and, according to some examples, may take values ​​of 34, 50, 101, or 152. According to some examples, such an m-layer ResNet may each comprise (m-2) / 2 skip connections.

[0083] A skip connection can be thought of as a bypass that feeds the output of a previous layer through one or more bypassed layers directly into a layer following the one or more bypassed layers. Instead of having to be directly adapted to a desired mapping, the bypassed layers would then have to be adapted to a residual mapping, which "balances" the directly fed output.

[0084] Fitting the residual map is computationally easier to optimize than the direct map. Furthermore, it mitigates the problem of vanishing / exploding gradients during optimization when training machine learning models: If a bypassed layer encounters such problems, its contribution can be skipped by regularizing the directly fed output. Using ResNets thus offers the advantage of allowing much deeper networks to be trained.

[0085] Specifically, a recurrent machine learning model is a machine learning model whose output depends not only on the input value and the machine learning model's parameters adjusted during the training process, but also on a hidden state vector, where the hidden state vector is based on previous inputs used for the recurrent machine learning model. In particular, the recurrent machine learning model may include additional memory states or additional structures that incorporate time delays or include feedback loops.

[0086] In particular, the underlying structure of a recurrent machine learning model can be a neural network, which can be referred to as a recurrent neural network. Such a recurrent neural network can be described as an artificial neural network, where connections between nodes form a directed graph along a temporal sequence. In particular, a recurrent neural network can be interpreted as a directed acyclic graph. In particular, the recurrent neural network can be a finite-pulse recurrent neural network or an infinite-pulse recurrent neural network (where a finite-pulse network can be rolled out and replaced by a pure feedforward neural network, and an infinite-pulse network can be unrolled and replaced by a pure feedforward neural network).

[0087] In particular, the training of a recurrent neural network can be based on the BPTT algorithm (abbreviation for "backpropagation through time"), the RTRL algorithm (abbreviation for "realtime recurrent learning") and / or on genetic algorithms.

[0088] When using a recurrent machine learning model, input data comprising sequences of variable length can be used. In particular, this means that the method can be used not only for a fixed number of input datasets (and must be trained differently for each different number of input datasets used as input), but for any number of input datasets. This means that the entire set of training data, regardless of the number of input datasets contained in different sequences, can be used within the training, and this training data is not reduced to training data corresponding to a specific number of consecutive input datasets.

[0089] Fig. Figure 7 shows a schematic structure of a recurrent machine learning model F, both in a recurrent representation 702 and in an unfolded representation 704, which can be used to implement one or more machine learning models described herein. The recurrent machine learning model takes multiple input data sets x, x 1 , ..., x N 706 as input and generates a corresponding set of output records y, y 1 , ..., y N 708. Furthermore, the output depends on a so-called hidden vector h, h 1 , ..., h N 710, which implicitly includes information about input data sets previously used as input to the recurrent machine learning model F 712. By using these hidden vectors h, h 1 , ..., h N 710, a sequentiality of the input data records can be used.

[0090] In a single processing step, the recurrent machine learning model F 712 takes the hidden vector h n-1 , which was generated in the previous step, and an input data set x n as input. In this step, the recurrent machine learning model F generates an updated hidden vector h n and an output data set y n as output. In other words, a processing step (y n , h n ) = F(x n , h n-1 ), or, when dividing the recurrent machine learning model F 712 into a part F(y), which calculates the output data, and F(h), which calculates the hidden vector, a processing step y calculates n = F (y) (x n , h n-1 ) and h n = F (h) (x n , h n-1 ). For the first processing step, h 0randomly selected or filled with all entries being zero. The parameters of the recurrent machine learning model F 712, which were previously trained based on training data sets, do not change between the different processing steps.

[0091] In particular, the output data and the hidden vector of a processing step depend on all previous input data sets used in the previous steps. n = F (y) (x n , F (h) (x n-1 , h n-2 )) and h n = F(h)(x n , F (h) (x n-1 , h n-2 )).

[0092] Systems, devices, and methods described herein may be implemented using digital circuitry or using one or more computers using well-known computer processors, memory units, storage devices, computer software, and other components. Typically, a computer includes a processor for executing instructions and one or more memories for storing instructions and data. A computer may also include or be coupled to one or more mass storage devices, such as one or more magnetic disks, internal hard disks and removable media, magneto-optical disks, optical disks, etc.

[0093] Systems, devices, and methods described herein may be implemented using computers operating in a client-server relationship. Typically, in such a system, the client computers are remote from the server computer and interact over a network. The client-server relationship may be defined and controlled by computer programs running on the respective client and server computers.

[0094] Systems, devices, and methods described herein may be implemented within a network-based cloud computing system. In such a network-based cloud computing system, a server or other processor connected to a network communicates with one or more client computers over a network. A client computer may communicate with the server, for example, through a network browser application residing and executing on the client computer. A client computer may store data on the server and access the data over the network. A client computer may send requests for data or requests for online services to the server over the network. The server may perform requested services and provide data to the one or more client computers. The server may also send data configured to cause a client computer to perform a specified function, e.g.,performs a calculation, displays specified data on a screen, etc. For example, the server may send a request configured to cause a client computer to perform one or more of the steps or functions of the methods and workflows described herein, including one or more of the steps or functions of the . Fig. 1, Fig. 3 or Fig. 4. Certain steps or functions of the methods and workflows described herein, including one or more of the steps or functions of Fig. 1, Fig. 3 or Fig. 4, may be performed by a server or other processor in a network-based cloud computing system. Certain steps or functions of the methods and workflows described herein, including one or more of the steps of Fig. 1, Fig. 3 or Fig. 4, may be performed by a client computer in a network-based cloud computing system. The steps or functions of the methods and workflows described herein, including one or more of the steps of Fig. 1, Fig. 3 or Fig. 4, may be performed by a server and / or a client computer in a network-based cloud computing system, in any combination.

[0095] Systems, devices, and methods described herein may be implemented using a computer program product tangibly embodied in an information carrier, e.g., a non-transitory machine-readable storage device, for execution by a programmable processor; and the method and workflow steps described herein, including one or more of the steps or functions of Fig. 1, Fig. 3 or Fig. 4, may be implemented using one or more computer programs executable by such a processor. A computer program is a set of computer program instructions that can be used directly or indirectly in a computer to perform a specific activity or achieve a specific result. A computer program may be written in any form of programming language, including compiled or interpreted languages, and the program may be implemented in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0096] A high-level block diagram of an exemplary computer 802 that may be used to implement systems, devices, and methods described herein is shown in Fig. 8. The computer 802 includes a processor 804 operatively coupled to a data storage device 812 and a memory 810. The processor 804 controls the overall operation of the computer 802 by executing computer program instructions that define such operations. The computer program instructions may be stored in the data storage device 812 or another computer-readable medium and loaded into the memory 810 when execution of the computer program instructions is desired. Therefore, the method and workflow steps or functions of the Fig. 1, Fig. 3 or Fig. 4 are defined by the computer program instructions stored in the memory 810 and / or the data storage device 812 and controlled by the processor 804, which executes the computer program instructions. For example, the computer program instructions may be implemented as computer-executable code programmed by those skilled in the art to implement the method and workflow steps or functions of the Fig. 1, Fig. 3 or Fig. 4. Accordingly, by executing the computer program instructions, the processor 804 performs the method and workflow steps or functions of the Fig. 1, Fig. 3 or Fig. 4. Computer 802 may also include one or more network interfaces 806 for communicating with other devices over a network. Computer 802 may also include one or more input / output devices 808 that enable user interaction with computer 802 (e.g., a display, a keyboard, a mouse, speakers, buttons, etc.).

[0097] Processor 804 may include both general-purpose and special-purpose microprocessors and may be the sole processor or one of multiple processors of computer 802. For example, processor 804 may include one or more central processing units (CPUs). Processor 804, data storage device 812, and / or memory 810 may include, be augmented by, or integrated with one or more application-specific integrated circuits (ASICs) and / or one or more field-programmable gate arrays (FPGAs).

[0098] The data storage device 812 and the memory 810 each include a tangible non-transitory computer-readable storage medium.The data storage device 812 and the memory 810 may each include high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), synchronous double data rate dynamic random access memory (DDR-RAM), or other solid-state random access memory devices, and may include non-volatile memory, such as one or more magnetic disk storage devices such as internal hard disks and removable disks, magneto-optical disk storage devices, optical disk storage devices, flash memory devices, semiconductor storage devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM), digital versatile disc read-only memory (DVD-ROM) disks, or other non-volatile solid-state storage devices.

[0099] The input / output devices 808 may include peripherals such as a printer, a scanner, a display screen, etc. For example, the input / output devices 808 may include a display device such as a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor for displaying information to the user, a keyboard, and a pointing device such as a mouse or trackball through which the user can provide input to the computer 802.

[0100] An image capture device 814 may be connected to the computer 802 to input image data (e.g., medical images) into the computer 802. It is possible to implement the image capture device 814 and the computer 802 as a single device. It is also possible for the image capture device 814 and the computer 802 to communicate wirelessly via a network. In one possible embodiment, the computer 802 may be remotely located with respect to the image capture device 814.

[0101] Any or all of the systems, devices, and methods discussed herein may be implemented using one or more computers, such as computer 802.

[0102] Those skilled in the art will recognize that an implementation of an actual computer or computer system may have different structures and may also include different components, and that Fig.Figure 8 is a high-level diagram of some of the components of such a computer for illustrative purposes.

[0103] Regardless of the grammatical usage of the term, persons with male, female or other gender identity are included in the term.

[0104] The foregoing detailed description is to be considered in all respects illustrative and exemplary, but not restrictive, and the scope of the invention disclosed herein is not to be determined from the detailed description, but rather from the claims as interpreted in accordance with the full breadth permitted by the patent laws. It is to be understood that the embodiments shown and described herein are merely illustrative of the principles of the present invention, and that various modifications may be implemented by those skilled in the art without departing from the scope and spirit of the invention. Those skilled in the art could implement various other combinations of features without departing from the scope and spirit of the invention.

[0105] The following is a list of non-limiting illustrative embodiments disclosed herein:

[0106] Illustrative Embodiment 1. A computer-implemented method comprising: receiving 1) one or more input medical images of a patient and 2) task instructions for performing a medical image processing analysis task; encoding the one or more input medical images into image features using an image encoding network; encoding the task instructions into text features using a text encoding network; performing the medical image processing analysis task based on the image features and the text features using a task network based on machine learning; and outputting results of the medical image processing analysis task.

[0107] Illustrative Embodiment 2. The computer-implemented method of illustrative embodiment 1, wherein the task instructions comprise references to image regions in at least one of the one or more input medical images.

[0108] Illustrative Embodiment 3. The computer-implemented method of any preceding embodiment, wherein the task instructions comprise anatomical knowledge and task knowledge. The task knowledge comprises at least one of a description of an anatomical anomaly, the manner in which the anatomical anomaly is represented, and the manner in which the anatomical anomaly may be detected in the one or more medical input images.

[0109] Illustrative Embodiment 4. The computer-implemented method of any preceding embodiment, wherein the task instructions are user-defined.

[0110] Illustrative Embodiment 5. The computer-implemented method of any preceding embodiment, wherein the text features and the image features are aligned in the same latent space.

[0111] Illustrative Embodiment 6. The computer-implemented method of any preceding embodiment, wherein the machine learning-based task network is trained with self-supervised learning based on unannotated medical training images and text.

[0112] Illustrative Embodiment 7. The computer-implemented method of any preceding embodiment, wherein the machine learning-based task network is trained with few-shot learning using annotated medical training images and annotated task descriptions.

[0113] Illustrative Embodiment 8. The computer-implemented method of any preceding embodiment, wherein the machine learning-based task network comprises a large language model (LLM)-based task network.

[0114] Illustrative Embodiment 9. The computer-implemented method of any preceding embodiment, wherein: receiving 1) one or more input medical images of a patient and 2) task instructions for performing a medical image processing analysis task comprises receiving text-based medical data of the patient; and encoding the task instructions into text features using a text encoding network comprises encoding the task instructions and the text-based medical data into the text features using the text encoding network.

[0115] Illustrative Embodiment 10. An apparatus comprising: means for receiving 1) one or more input medical images of a patient and 2) task instructions for performing a medical image processing analysis task; means for encoding the one or more input medical images into image features using an image encoding network; means for encoding the task instructions into text features using a text encoding network; means for performing the medical image processing analysis task based on the image features and the text features using a task network based on machine learning; and means for outputting results of the medical image processing analysis task.

[0116] Illustrative Embodiment 11. The apparatus of illustrative embodiment 10, wherein the task instructions include references to image regions in at least one of the one or more input medical images.

[0117] Illustrative Embodiment 12. The device of any of illustrative embodiments 10-11, wherein the task instructions comprise anatomical knowledge and task knowledge. The task knowledge comprises at least one of a description of an anatomical anomaly, the manner in which the anatomical anomaly is represented, and the manner in which the anatomical anomaly may be detected in the one or more medical input images.

[0118] Illustrative Embodiment 13. The device of any of illustrative embodiments 10-12, wherein the task instructions are user-defined.

[0119] Illustrative Embodiment 14. A non-transitory computer-readable medium comprising instructions that, when executed by a computer, cause the computer to perform operations comprising: receiving 1) one or more input medical images of a patient and 2) task instructions for performing a medical image processing analysis task; encoding the one or more input medical images into image features using an image encoding network; encoding the task instructions into text features using a text encoding network; performing the medical image processing analysis task based on the image features and the text features using a task network based on machine learning; and outputting results of the medical image processing analysis task.

[0120] Illustrative Embodiment 15. The non-transitory computer-readable medium of illustrative embodiment 14, wherein the text features and the image features are aligned in the same latent space.

[0121] Illustrative Embodiment 16. The non-transitory computer-readable medium of any of illustrative embodiments 14-15, wherein the machine learning-based task network is trained with self-supervised learning based on unannotated medical training images and text.

[0122] Illustrative Embodiment 17. The non-transitory computer-readable medium of any of illustrative embodiments 14-16, wherein the machine learning-based task network is trained with few-shot learning using annotated medical training images and annotated task descriptions.

[0123] Illustrative Embodiment 18. A computer-implemented method, comprising: receiving 1) one or more medical training images of a patient and 2) training task instructions for performing a medical image processing analysis task; encoding the one or more medical training images into image features using a pre-trained image encoding network; encoding the training task instructions into text features using a pre-trained text encoding network; training a machine-learning-based task network for performing the medical image processing analysis task based on the image features and the text features; and outputting the trained machine-learning-based task network.

[0124] Illustrative Embodiment 19. The computer-implemented method of illustrative embodiment 18, wherein the machine learning-based task network is trained with self-supervised learning based on unannotated medical training images and text.

[0125] Illustrative Embodiment 20. The computer-implemented method of any of illustrative embodiments 18-19, wherein the machine learning-based task network is trained with few-shot learning using annotated medical training images and annotated task descriptions.

Claims

[1] Computer-implemented method comprising: Receiving 1) one or more input medical images of a patient and 2) task instructions for performing a medical image processing analysis task; encoding the one or more medical input images into image features using an image encoding network; Encoding the task instructions into text features using a text coding network; Performing the medical image processing analysis task based on the image features and the text features using a task network based on machine learning; and Outputting results of the medical image processing analysis task. [2] The computer-implemented method of claim 1, wherein the task instructions comprise references to image regions in at least one of the one or more input medical images. [3] The computer-implemented method of claim 1, wherein the task instructions comprise anatomical knowledge and task knowledge, wherein the task knowledge comprises at least one of a description of an anatomical anomaly, the manner in which the anatomical anomaly is represented, and the manner in which the anatomical anomaly can be detected in the one or more input medical images. [4] The computer-implemented method of claim 1, wherein the task instructions are user-defined. [5] The computer-implemented method of claim 1, wherein the text features and the image features are laid out on the same latent space. [6] The computer-implemented method of claim 1, wherein the machine learning-based task network is trained with self-supervised learning based on unannotated medical training images and text. [7] The computer-implemented method of claim 1, wherein the machine learning-based task network is trained with few-shot learning using annotated medical training images and annotated task descriptions. [8] The computer-implemented method of claim 1, wherein the machine learning-based task network comprises a LLM-based (Large Language Model-based) task network. [9] A computer-implemented method according to claim 1, wherein: receiving 1) one or more medical input images of a patient and 2) task instructions for performing a medical image processing analysis task comprises receiving text-based medical data of the patient; and encoding the task instructions into text features using the text coding network comprises encoding the task instructions and the text-based medical data into the text features using the text coding network. [10] Facility comprising: Means for receiving 1) one or more medical input images of a patient and 2) task instructions for performing a medical image processing analysis task; means for encoding the one or more medical input images into image features using an image encoding network; means for encoding the task instructions into text features using a text encoding network; means for performing the medical image processing analysis task based on the image features and the text features using a task network based on machine learning; and Means for outputting results of the medical image processing analysis task. [11] A non-transitory computer-readable medium comprising instructions which, when executed by a computer, cause the computer to perform operations comprising: Receiving 1) one or more input medical images of a patient and 2) task instructions for performing a medical image processing analysis task; encoding the one or more medical input images into image features using an image encoding network; Encoding the task instructions into text features using a text coding network; Performing the medical image processing analysis task based on the image features and the text features using a task network based on machine learning; and Outputting results of the medical image processing analysis task.

Citation Information

Patent Citations

  • Chest radiograph automatic diagnosis system with enhanced medical knowledge

    CN116665877A

  • Training of text and image models

    EP4266195A1

  • Apparatus and method for diagnosing a medical condition from a medical image

    US20220375576A1

  • CN000116665877A