Machine learning based medical imaging analysis using sample-less learning with task instructions
By using a machine learning-based task network in medical imaging analysis, combining image encoding and text encoding, the dependence problem on a large number of labeled images in the prior art is solved, and more efficient and accurate medical imaging analysis is achieved.
Patent Information
- Application Number
- CN202411633383.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-17
- Filing Date
- 2024-11-15
- Publication Date
- 2025-05-20
AI Technical Summary
The prior art requires a large number of labeled medical images in medical imaging analysis. Acquisition and labeling of these images is time-intensive and expensive, and is greatly affected by the difference in labelers.
Using a machine learning-based task network, input medical images and task instructions are encoded into imaging features and text features through image encoder networks and text encoder networks, and medical imaging analysis tasks are performed based on these features.
Reduces dependence on a large number of labeled medical images, shortens training time and cost, and improves the accuracy and consistency of the analysis tasks.
Smart Images

Figure CN120020868A_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to medical imaging analysis, and more particularly to machine learning-based medical imaging analysis using few shot learning with task instructions. Background Art
[0002] Machine learning models have recently been proposed to perform various medical imaging analysis tasks, such as for example segmentation, registration, classification, detection, diagnosis, etc. Generally, such machine learning models use supervised learning to train for performing a specific medical imaging analysis task. Supervised learning requires a large number of task-specific medical images with expert annotations. However, obtaining such a large number of annotated medical images is time-consuming, expensive, and affected by differences between annotators. Summary of the Invention
[0003] According to one or more embodiments, there are provided systems and methods for performing a medical imaging analysis task using a machine learning-based task network based on task instructions. One or more input medical images of a patient and task instructions for performing a medical imaging analysis task are received. The one or more input medical images are encoded into imaging features using an image encoder network. The task instructions are encoded into text features using a text encoder network. A machine learning-based task network performs a medical imaging analysis task based on the imaging features and the text features. The result of the medical imaging analysis task is output.
[0004] In one embodiment, the task instructions include a reference to an image region in at least one of the one or more input medical images. The task instructions may include anatomical knowledge and task knowledge. The task knowledge may include a description of an anatomical abnormality, how to represent the anatomical abnormality, and at least one of how the anatomical abnormality can be detected in the one or more input medical images. The task instructions may be user-defined.
[0005] In one embodiment, the text features and the imaging features are aligned in the same latent space.
[0006] In one embodiment, the machine learning-based task network is trained using self-supervised learning based on unlabeled training medical images and text. In one embodiment, the machine learning-based task network is trained using few shot learning with labeled training medical images and labeled task descriptions. In one embodiment, the machine learning-based task network includes a large language model (LLM)-based task network.
[0007] In one embodiment, text-based medical data of a patient is received. The task instructions and the text-based medical data are encoded into text features using a text encoder network.
[0008] According to one or more embodiments, systems and methods are provided for training a machine learning-based task network for performing a medical imaging analysis task based on a task instruction. One or more training medical images and training task instructions for performing a medical imaging analysis task are received. The one or more training medical images are encoded into imaging features using a pre-trained image encoder network. The training task instructions are encoded into text features using a pre-trained text encoder network. The machine learning-based task network is trained based on the imaging features and the text features for performing a medical imaging analysis task. The trained machine learning-based task network is output.
[0009] In one embodiment, the machine learning-based task network is trained using self-supervised learning based on unlabeled training medical images and text. In one embodiment, the machine learning-based task network is trained using few-shot learning using labeled training medical images and labeled task descriptions.
[0010] These and other advantages of the present invention will be apparent to those of ordinary skill in the art by reference to the following detailed description and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 A method for performing a medical imaging analysis task using a machine learning model based on a task instruction according to one or more embodiments is shown; Figure 2 Exemplary task instructions for detecting pneumothorax in a chest x-ray image according to one or more embodiments are shown; Figure 3 A method for training a machine learning-based task network for performing a medical imaging analysis task based on training task instructions according to one or more embodiments is shown; Figure 4 A workflow for training a machine learning-based task network for performing a medical imaging analysis task based on training task instructions according to one or more embodiments is shown; Figure 5 An exemplary artificial neural network that can be used to implement one or more embodiments is shown; Figure 6 A convolutional neural network that can be used to implement one or more embodiments is shown; Figure 7 A schematic structure of a recurrent machine learning model that can be used to implement one or more embodiments is shown; and Figure 8 A high-level block diagram of a computer that can be used to implement one or more embodiments is shown. Detailed implementation manners
[0012] The present invention generally relates to methods and systems for machine - learning - based medical imaging analysis using few - shot learning with task instructions. Embodiments of the present invention are described herein to give a visual understanding of such methods and systems. Digital images generally consist of digital representations of one or more objects (or shapes). The digital representations of the objects are generally described herein in terms of identifying and manipulating the objects. Such manipulation is virtual manipulation done in the memory of a computer system or other circuitry / hardware. Thus, it is to be understood that embodiments of the present invention can be implemented within a computer system using data stored within the computer system. Additionally, references to image pixels herein can equivalently refer to voxels of an image, and vice versa.
[0013] Traditional machine - learning models are trained via supervised learning using a large number of labeled medical images for performing medical imaging analysis tasks. Such a large number of labeled medical images are required in supervised learning to implicitly encode all the details and possible variations of performing medical imaging analysis tasks by example.
[0014] According to embodiments described herein, a machine - learning model for performing medical imaging analysis tasks is trained using few - shot learning with task instructions. The task instructions provide a "recipe" for how to perform the medical imaging analysis task. By encoding the task instructions and the medical images, the text encoding of the task instructions is used for task reasoning, thereby overcoming the complexity of scaling the amount of labeled medical images required for supervised learning and using more context information provided by the task instructions.
[0015] Figure 1 Method 100 for performing a medical imaging analysis task using a machine - learning model based on task instructions according to one or more embodiments is shown. The steps of method 100 can be performed by one or more suitable computing devices (such as, for example Figure 8 computer 802).
[0016] In Figure 1 step 102, 1) receive one or more input medical images of a patient and 2) task instructions for performing a medical imaging analysis task.
[0017] One or more input medical images can depict any anatomical object of interest to a patient, such as, for example, an organ, a blood vessel, a tumor, an abnormality, etc. The one or more input medical images can belong to any suitable one or more modalities, such as, for example, CT (Computed Tomography), MRI (Magnetic Resonance Imaging), US (Ultrasound), x-ray, or any other medical imaging modality or combination of medical imaging modalities. The one or more input medical images can include 2D (two-dimensional) images and / or 3D (three-dimensional) volumes.
[0018] Task instructions act as a "recipe" or instructions on how a medical imaging analysis task is to be performed. The task instructions can include anatomical knowledge, task knowledge, and / or any other suitable information that can be used to perform the medical imaging analysis task. In one embodiment, the task instructions are associated with one or more input medical images. For example, the task instructions can include explicit groundings or references to anatomical image regions and / or task-specific image regions in at least one of the one or more input medical images. In one embodiment, the task instructions are user-defined task instructions generated by a user. However, any other suitable method can be used to generate the task instructions. Figure 2 Exemplary task instructions are shown. The task knowledge can include, for example, a description of an anatomical abnormality, how to represent an anatomical abnormality, how an anatomical abnormality can be detected in one or more input medical images, etc.
[0019] Figure 2 Exemplary task instructions 202 for detecting pneumothorax in a chest x-ray image are shown. As Figure 2As shown, task instruction 202 includes anatomical knowledge and task knowledge. The anatomical knowledge describes the anatomy of the pleura and pleural space. The task knowledge describes the radiological features of pleural diseases, pneumothorax, and pleural thickening. Task instruction 202 includes groundings or references 206-A, 206-B, and 206-C (collectively, groundings 206) to anatomical image regions and / or task-specific image regions in input medical images 204-A, 204-B, and 204-C (collectively, input medical images 204). In particular, grounding 206-A grounds or references the pleura and pleural space described in task instruction 202 to the anatomical region of the pleura and pleural space shown in input medical image 204-A. Grounding 206-B grounds or references the radiological features of pneumothorax in task instruction 202 ("visible pleural edge" and "lung markings not visible beyond that edge") to the anatomical region shown in input medical image 204-B. Grounding 206-C grounds or references the radiological feature of pleural thickening ("shadow over the entire right lung") to the anatomical region shown in input medical image 204-C. In one example, input medical images 204 and task instruction 202 are respectively one or more input medical images and task instructions received at step 102 of Figure 1 One or more input medical images and task instructions received at step 102.
[0020] Referring back to Figure 1 , in one embodiment, non-imaging medical data of the patient may also be received at step 102. Such non-imaging medical data may include, for example, text-based medical data such as radiology reports, laboratory reports, medical records, instructions for imaging / examinations, demographic information, administrative data, and the like.
[0021] One or more input medical images and task instructions (and non-imaging medical data) may be received, for example, by loading the input medical images, task instructions, and / or non-imaging medical data from a storage device or memory of one or more computer systems and / or receiving the input medical images, task instructions, and / or non-imaging medical data from one or more remote computer systems. Such computer systems may include an EHR (Electronic Health Record), EMR (Electronic Medical Record), PHR (Personal Health Record), HIS (Health Information System), RIS (Radiology Information System), PACS (Picture Archiving and Communication System), LIMS (Laboratory Information Management System), or any other suitable database or system. In some embodiments, the input medical images may be received directly from an image acquisition device (such as a CT scanner) when acquiring the medical images.
[0022] In Figure 1At step 104, an image encoder network is used to encode one or more input medical images into imaging features. The image encoder network can be a machine learning-based image encoder network, such as, for example, an autoencoder, a VAE (Variational Autoencoder), or can be implemented according to any other suitable machine learning-based architecture. The image encoder network receives one or more input medical images as input and generates imaging features as output. The imaging features are embeddings that represent the relatively high-dimensional input medical images as low-dimensional, dense vector representations in the latent space. The image encoder network is pre-trained using a large set of training images during a previous offline or training phase. Once trained, the image encoder network is applied during the online or inference phase, for example, to perform Figure 1 step 104.
[0023] At Figure 1 step 106, a text encoder network is used to encode the task instructions into text features. In one embodiment, in Figure 1 the case where non-imaging medical data of the patient is also received at step 102, the non-imaging medical data (e.g., text-based medical data) is encoded together with the task instructions into imaging features using the image encoder network.
[0024] In one embodiment, the text encoder network is an LLM (Large Language Model)-based text encoder network. However, the text encoder network can be implemented according to any suitable machine learning-based architecture. The LLM-based text encoder network can be any suitable pre-trained deep learning-based LLM. For example, the LLM-based text encoder network can be based on a Transformer architecture that uses self-attention mechanisms to capture long-range dependencies in the text. An example of a Transformer-based architecture is GPT (Generative Pretrained Transformer), which has a multi-layer Transformer decoder architecture that can be pre-trained to optimize the next token prediction task and then fine-tuned using labeled data for various downstream tasks. The GPT-based LLM can be trained using reinforcement learning with human feedback for performing various natural language processing tasks. Other exemplary Transformer-based architectures include BLOOM (BigScience Large Open Science Open Access Multilingual Language Model) and BERT (Bidirectional Encoder Representations from Transformers).
[0025] In one embodiment, the LLM-based text encoder network is constrained to a specific medical domain. For example, the LLM-based text encoder network may be constrained for a use case of detecting pneumothorax in chest x-ray images. To constrain the LLM-based text encoder network, clinical data and data extracted from medical images using an AI-based system can be used to update (e.g., train, retrain, or fine-tune) the LLM-based text encoder network. Such extracted data can include, for example, clinical measurements (e.g., diameter, volume, distance, etc.), anatomical locations, detections, etc.
[0026] The LLM-based text encoder network receives task instructions (and possibly non-imaging medical data) as input (e.g., as one or more prompts) and generates text features as output. In one embodiment, the task instructions include specific prompts for the LLM-based text encoder network (and possibly the LLM-based task network utilized at step 108 of Figure 1 to perform a task. Alternatively, at least some of the task instructions can be separately / directly input into the machine learning-based task network (utilized at step 108 of Figure 1 ). The text features are embeddings that represent the task instructions (and possibly non-imaging medical data) as low-dimensional, dense vector representations in a latent space that is relatively high-dimensional. The LLM-based text encoder network is pre-trained during a previous offline or training phase using a large set of text-based training data. Once trained, the LLM-based text encoder network is applied during an online or inference phase, e.g., to perform step 106 of Figure 1 .
[0027] The text features and the image features are represented in the same latent space. The latent space can be refined to align the corresponding image features and text features such that similar features represent similar concepts. In one embodiment, the image encoder network and the text encoder network can be pre-trained on image-text pairs (such as, for example, corresponding image-reports or other semantically annotated image / image-region-text pairs) to align the text and image features in the latent space. The pre-training can be performed, for example, by minimizing the distance in the latent space between the corresponding pairs and maximizing the distance between the non-corresponding pairs. By aligning the image features and the text features, inference can be performed in the same latent feature space with interchangeable features (i.e., from text features or image features).
[0028] In Figure 1At step 108, a machine learning-based task network performs a medical imaging analysis task based on imaging features and text features. The medical imaging analysis task can include any suitable medical imaging analysis task, such as, for example, segmentation, registration, classification, detection, diagnosis, medical data aggregation, etc.
[0029] In one embodiment, the machine learning-based task network is an LLM-based task network. However, the machine learning-based task network can be implemented according to any suitable machine learning-based architecture. The LLM-based task network can be any suitable pre-trained deep learning-based LLM, such as, for example, a transformer-based network (e.g., GPT-based LLM, BMOOM, BERT). In one embodiment, the LLM-based task network is constrained to a specific medical domain (e.g., detecting pneumothorax in chest x-ray images).
[0030] The LLM-based task network receives imaging features and text features as inputs (e.g., as one or more prompts) and generates the result of the medical imaging analysis task as an output. The LLM-based task network also receives instructions for performing the medical imaging analysis task. These instructions can be received indirectly through a text encoder network and / or an image encoder network, or directly as separate task instructions. In one embodiment, the LLM-based task network can include multiple heads, each head for performing a corresponding medical imaging analysis task (e.g., text decoding, classification scoring, segmentation, detection, etc.). The LLM-based task network is trained during a previous offline or training phase. For example, the LLM-based task network can be trained using few-shot learning and self-supervised learning with a relatively small number of labeled training images and task instructions, as described in further detail below with respect to Figure 3 and 4 Once trained, the LLM-based task network is applied during an online or inference phase, for example, to perform Figure 1 step 108.
[0031] In Figure 1 At step 110, the result of the medical imaging analysis task is output. For example, the result of the medical imaging analysis task can be output by displaying the result of the medical imaging analysis task on a display device of a computer system, storing the result of the medical imaging analysis task in the memory or storage device of the computer system, or by transmitting the result of the medical imaging analysis task to a remote computer system.
[0032] In one embodiment, the task instructions (in Figure 1The verification steps (for self-verification) to be performed by the final system, which may be received at step 102 of , can be included to ensure information consistency, thereby increasing operational robustness, and this can be combined with known uncertainty estimation techniques.
[0033] In one embodiment, additional instructions can be received from a user via an interactive user interface. The additional instructions can be parsed and verified based on one or more input medical images and current knowledge. The interactive user interface can be utilized to probe the user for an interpretation of the current medical imaging analysis task. This can be used to improve workflow efficiency using an LLM.
[0034] Figure 3 FIG. 300 shows a method 300 for training a machine learning-based task network for performing a medical imaging analysis task based on training task instructions. The steps of method 300 can be performed by one or more suitable computing devices (such as, for example, Figure 8 computer 802 of . Figure 4 FIG. 400 shows a workflow 400 for training a machine learning-based task network for performing a medical imaging analysis task based on training task instructions according to one or more embodiments. Figure 3 and Figure 4 will be described together. Figure 3 The steps of method 300 of Figure 4 and the workflow 400 of are performed during an offline or training phase for training the machine learning-based task network. Once trained, the machine learning-based task network is applied during an online or inference phase to perform a medical imaging analysis task, such as at Figure 1 step 108 of .
[0035] At Figure 3 step 302 of , 1) one or more training medical images and 2) training task instructions for performing a medical imaging analysis task are received. In one embodiment, non-imaging training medical data (e.g., text-based medical data) can also be received at step 302. In one example, as Figure 4 shown in workflow 400 of , the one or more training medical images are curated training medical images 404 selected from a data lake 402, and the training task instructions are training task instructions 406.
[0036] One or more training medical images can depict any anatomical object of interest of a patient and can belong to any suitable one or more modalities, such as, for example, CT, MRI, US, x-ray, or any other medical imaging modality or combination of medical imaging modalities. One or more training medical images can include 2D images and / or 3D volumes. One or more training medical images can include one or more annotated medical images (e.g., annotated by a user) and one or more unannotated medical images.
[0037] Training task instructions can include anatomical knowledge, task knowledge, and / or any other suitable information that can be used to perform a medical imaging analysis task. During the training phase, the training task instructions are associated with one or more training medical images. For example, the training task instructions can include explicit grounding or references to anatomical image regions and / or task-specific image regions in at least one of the one or more training medical images. The grounding is used to align the embeddings of corresponding image-text concepts in the latent space. The training task instructions can be extracted from textbooks, publications, websites, etc., or can be user-defined.
[0038] One or more training medical images and training task instructions (and non-imaging training medical data) can be received, for example, by loading the training medical images, training task instructions, and / or non-imaging training medical data from a storage device or memory of one or more computer systems, and / or receiving the training medical images, training task instructions, and / or non-imaging training medical data from one or more remote computer systems. In some embodiments, the training medical images can be received directly from an image acquisition device when acquiring the medical images.
[0039] At Figure 3 step 304, one or more training medical images are encoded into imaging features using a pre-trained image encoder network. The pre-trained image encoder network can be, for example, an autoencoder, a VAE, or can be implemented according to any other suitable machine learning-based architecture. The pre-trained image encoder network receives one or more training medical images as input and generates imaging features as output. The pre-trained image encoder network is pre-trained using a large set of training images during a previous offline or training phase. Once trained, the pre-trained image encoder network is applied during the online or inference phase, for example, to perform Figure 3 step 304. In one example, as shown in Figure 4 workflow 400, the pre-trained image encoder network is image AI (artificial intelligence) encoder 408, and encoder 408 receives training medical image 404 as input and generates image feature 410 as output.
[0040] At Figure 3At step 306, a pre-trained text encoder network is used to encode the training task instructions into text features. In one embodiment, in Figure 3 In the case where non-imaging training medical data is also received at step 302, a pre-trained image encoder network is used to encode the non-imaging training medical data (e.g., text-based medical data) together with the training task instructions into imaging features.
[0041] In one embodiment, the text encoder network is an LLM-based text encoder network. However, the text encoder network can be implemented according to any suitable machine learning-based architecture. In one embodiment, the LLM-based text encoder network is constrained to a specific medical domain. The LLM-based text encoder network receives the training task instructions (and possibly non-imaging training medical data) as input and generates text features as output. The image features 410 and the text features 414 are aligned within the latent space. The LLM-based text encoder network is pre-trained using a large set of text-based training data during a previous offline or training phase. Once trained, the LLM-based text encoder network is applied during the online or inference phase, e.g., to perform Figure 3 step 306.
[0042] In one example, as Figure 4 shown in the workflow 400, the pre-trained text encoder network is the LLM text AI encoder 412, and the encoder 412 receives the training task instructions 406 as input and generates the text features 414 as output.
[0043] At Figure 3 step 308, a machine learning-based task network is trained based on the imaging features and the text features for performing a medical imaging analysis task. In one embodiment, the machine learning-based task network is an LLM-based task network. However, the machine learning-based task network can be implemented according to any suitable machine learning-based architecture. In one embodiment, the LLM-based task network is constrained to a specific medical domain (e.g., detecting pneumothorax in chest x-ray images).
[0044] The machine learning-based task network receives the image features and the text features and generates the result of the medical imaging analysis task as output. The machine learning-based task network is trained using few-shot learning and self-supervised learning.
[0045] In few-shot learning, one or more labeled training images and a labeled task description are used to train a machine learning-based task network to accurately perform a medical imaging analysis task. Few-shot learning applies meta-learning such that the machine learning-based task network learns to learn. During the meta-training phase, the machine learning-based task network is trained on a relatively small number of relevant tasks. During the meta-testing phase, the machine learning-based task network can generalize to unseen (but relevant) tasks.
[0046] In self-supervised learning, one or more unlabeled training medical images and text are used to train a machine learning-based task network. Aligned image and text features enable the text features to be used to generate task instructions and pseudo-labels for one or more unlabeled training medical images and text.
[0047] In one embodiment, several output “heads” can be decoded in a generation framework. Example decoding heads can be text decoding, classification scoring, segmentation, detection, etc. Using an LLM, this decoding is implicitly performed as text. Training for this can be performed, for example, by sequentially predicting the same tokens from ground truth (e.g., using a cross-entropy loss function) or by reinforcement learning (with / without human feedback) — where some text decodings are more preferred than others.
[0048] In one example, as Figure 4 shown in the workflow 400, the machine learning-based task network is an RL-LLM 416 that receives image features 410 and text features 414 as inputs and generates a result 418 of a medical imaging analysis task as an output.
[0049] In Figure 3 step 310, a trained machine learning-based task network is output. For example, the trained machine learning-based task network can be output by storing the trained machine learning-based task network in a memory or storage device of a computer system or by transmitting the trained machine learning-based task network to a remote computer system. In one example, the trained machine learning-based task network is applied to, for example, perform Figure 1 step 108.
[0050] The embodiments described herein are described with respect to the claimed system and the claimed method. Features, advantages, or alternative embodiments herein can be assigned to other claimed subject matter and vice versa. In other words, the claims and embodiments for the system can be improved using features described or claimed in the context of the corresponding method. In such cases, the functional features of the method are implemented by the physical units of the system.
[0051] In addition, certain embodiments described herein are described in terms of methods and systems for utilizing trained machine learning models and methods and systems for providing trained machine learning models. Features, advantages, or alternative embodiments herein may be assigned to other claimed subject matter and vice versa. In other words, claims and embodiments for providing trained machine learning models may be improved using features described or claimed in the context of utilizing trained machine learning models and vice versa. In particular, a dataset used in methods and systems for utilizing trained machine learning models may have the same attributes and characteristics as a corresponding dataset used in methods and systems for providing trained machine learning models, and a trained machine learning model provided by the corresponding methods and systems may be used in methods and systems for utilizing trained machine learning models.
[0052] Generally, a trained machine learning model mimics cognitive functions associated with human-to-human thinking. In particular, through training based on training data, a machine learning model is able to adapt to new environments and detect and infer patterns. Another term for "trained machine learning model" is "trained function". Generally, the parameters of a machine learning model can be adapted by means of training. In particular, supervised training, semi-supervised training, unsupervised training, reinforcement learning, and / or active learning can be used. In addition, representation learning (an alternative term is "feature learning") can be used. In particular, the parameters of a machine learning model can be iteratively adapted through a number of training steps. In particular, within training, a certain cost function can be minimized. In particular, within the training of a neural network, the backpropagation algorithm can be used.
[0053] In particular, machine learning models such as Figure 1 the image encoder network used at step 104, the text encoder network used at step 106, and the machine learning-based task network used at step 108, Figure 3 the pre-trained image encoder network used at step 304, the pre-trained text encoder network used at step 306, and the machine learning-based task network used at step 308, and Figure 4The image AI encoder 408, the LLM text AI encoder 412, and the RL-LLM 416 may include, for example, neural networks, support vector machines, decision trees, and / or Bayesian networks, and / or the machine learning models may be based on, for example, k-means clustering, Q-learning, genetic algorithms, and / or association rules. In particular, the neural network may be, for example, a deep neural network, a convolutional neural network, or a convolutional deep neural network. Additionally, the neural network may be, for example, an adversarial network, a deep adversarial network, and / or a generative adversarial network.
[0054] Figure 5 An embodiment of an artificial neural network 500 that can be used to implement one or more of the machine learning models described herein is shown. Alternative terms for "artificial neural network" are "neural network", "artificial neural net", or "neural net".
[0055] The artificial neural network 500 includes nodes 520...532 and edges 540...542, where each edge 540...542 is a directed connection from a first node 520...532 to a second node 520...532. Generally, the first node 520...532 and the second node 520...532 are different nodes 520...532, and it is also possible that the first node 520...532 and the second node 520...532 are the same. For example, in Figure 5 the edge 540 is a directed connection from node 520 to node 523, and the edge 542 is a directed connection from node 530 to node 532. The edge 540...542 from the first node 520...532 to the second node 520...532 is also represented as an "incoming edge" for the second node 520...532 and an "outgoing edge" for the first node 520...532.
[0056] In this embodiment, the nodes 520...532 of the artificial neural network 500 may be arranged in layers 510...513, where these layers may include an inherent order introduced by the edges 540...542 between the nodes 520...532. In particular, the edges 540...542 may only exist between adjacent layers of nodes. In the shown embodiment, there is an input layer 510 that includes only nodes 520...522 and has no incoming edges, an output layer 513 that includes only nodes 531, 532 and has no outgoing edges, and hidden layers 511, 512 between the input layer 510 and the output layer 513. Generally, the number of hidden layers 511, 512 can be arbitrarily selected. The number of nodes 520...522 in the input layer 510 is typically related to the number of input values of the neural network, and the number of nodes 531, 532 in the output layer 513 is typically related to the number of output values of the neural network.
[0057] Specifically, (real) numbers can be assigned as values to each node 520...532 of the neural network 500. Here, x (n) i represents the value of the i-th node 520...532 of the n-th layer 510...513. The values of the nodes 520...522 of the input layer 510 are equal to the input values of the neural network 500, and the values of the nodes 531, 532 of the output layer 513 are equal to the output values of the neural network 500. In addition, each edge 540...542 can include a weight that is a real number, specifically, the weight is a real number within the interval [-1, 1] or within the interval [0, 1]. Here, w (m,n) i,j represents the weight of the edge between the i-th node 520...532 of the m-th layer 510...513 and the j-th node 520...532 of the n-th layer 510...513. In addition, the abbreviation w (n) i,j is defined for the weight w (n,n+1) i,j .
[0058] Specifically, to calculate the output value of the neural network 500, the input values are propagated through the neural network. Specifically, the value of the node 520...532 of the (n + 1)-th layer 510...513 can be calculated based on the value of the node 520...532 of the n-th layer 510...513 by the following formula:
[0059] In this article, the function f is a transfer function (another term is "activation function"). Known transfer functions are step functions, sigmoid functions (e.g., logistic functions, generalized logistic functions, hyperbolic tangent functions, arctangent functions, error functions, smoothstep functions), or rectifier functions. Transfer functions are mainly used for normalization purposes.
[0060] Specifically, these values are propagated layer by layer through the neural network, where the value of the input layer 510 is given by the input of the neural network 500, where the value of the first hidden layer 511 can be calculated based on the value of the input layer 510 of the neural network, where the value of the second hidden layer 512 can be calculated based on the value of the first hidden layer 511, and so on.
[0061] To set the values of the edges the neural network 500 must be trained using training data. Specifically, the training data includes training input data and training output data (denoted as t i)。For the training step, the neural network 500 is applied to the training input data to generate the computed output data. In particular, the training data and the computed output data include a certain number of values, the number being equal to the number of nodes in the output layer.
[0062] In particular, the comparison between the computed output data and the training data is used to recursively adapt the weights within the neural network 500 (backpropagation algorithm). In particular, the weights are changed according to the following formula: where γ is the learning rate, and if the (n + 1)-th layer is not the output layer, δ can be based on (n+1) j recursively calculate the number δ (n) j as: and if the (n + 1)-th layer is the output layer 513, the number δ (n) j is calculated as: where f' is the first derivative of the activation function, and t (n+1) j is the comparison training value of the j-th node of the output layer 513.
[0063] A convolutional neural network is a neural network that uses a convolution operation instead of a general matrix multiplication in at least one of its layers (the so-called "convolutional layer"). In particular, the convolutional layer performs the dot product of one or more convolutional kernels and the input data / image of the convolutional layer, where the entries of one or more convolutional kernels are parameters or weights adapted through training. In particular, the Frobenius inner product and the ReLU activation function can be used. A convolutional neural network can include additional layers, such as pooling layers, fully connected layers, and normalization layers.
[0064] By using a convolutional neural network, the input image can be processed in a very efficient way, because the convolution operations based on different kernels can extract various image features, so that by adapting the weights of the convolutional kernels, relevant image features can be found during training. In addition, based on the weight sharing in the convolutional kernels, fewer parameters need to be trained, which prevents overfitting during the training phase and allows for faster training or more layers in the network, thus improving the performance of the network.
[0065] Figure 6FIG. 0 shows an embodiment of a convolutional neural network 600 that can be used to implement one or more machine learning models described herein. In the illustrated embodiment, the convolutional neural network 600 includes an input node layer 610, a convolutional layer 611, a pooling layer 613, a fully connected layer 614 and an output node layer 616, as well as hidden node layers 612, 614. Alternatively, the convolutional neural network 600 may include a number of convolutional layers 611, a number of pooling layers 613 and a number of fully connected layers 615, as well as other types of layers. The order of the layers can be arbitrarily selected, and typically the fully connected layer 615 is used as the last layer before the output layer 616.
[0066] Specifically, within the convolutional neural network 600, the nodes 620, 622, 624 of the node layers 610, 612, 614 can be regarded as being arranged as a d-dimensional matrix or a d-dimensional image. Specifically, in the two-dimensional case, the value of the node 620, 622, 624 indexed by i and j in the nth node layer 610, 612, 614 can be expressed as x(n)[i,j]. However, the arrangement of the nodes 620, 622, 624 of a node layer 610, 612, 614 itself has no effect on the calculations performed within the convolutional neural network 600, since these are given only by the weights and structure of the edges.
[0067] The convolutional layer 611 is a connection layer between the previous node layer 610 (with node values x(n-1)) and the subsequent node layer 612 (with node values x(n)). Specifically, the convolutional layer 611 is characterized by the structure and weights of the incoming edges that form a convolutional operation based on a certain number of kernels. Specifically, the structure and weights of the edges of the convolutional layer 611 are selected such that the value x(n) of the nodes 622 in the subsequent node layer 612 is calculated as a convolution x(n)=K*x(n-1) based on the value x(n-1) of the nodes 620 in the previous node layer 610, where the convolution * is defined in the two-dimensional case as:
[0068] Here, the kernel K is a d-dimensional matrix (in this embodiment, a two-dimensional matrix), which is typically small compared to the number of nodes 620, 622 (e.g., a 3×3 matrix or a 5×5 matrix). Specifically, this means that the weights of the edges in the convolutional layer 611 are not independent, but are selected such that they produce the said convolutional equation. Specifically, for a kernel that is a 3×3 matrix, there are only 9 independent weights (each entry of the kernel matrix corresponds to an independent weight), regardless of the number of nodes 620, 622 in the previous node layer 610 and the subsequent node layer 612.
[0069] Generally, the convolutional neural network 600 uses node layers 610, 612, 614 with multiple channels, especially due to the use of multiple kernels in the convolutional layer 611. In these cases, the node layer can be considered a (d + 1)-dimensional matrix (indexing the first dimension of the channels). Then, the operation of the convolutional layer 611 is a two-dimensional example, defined as: where corresponds to the ath channel of the previous node layer 610, corresponds to the bth channel of the subsequent node layer 612, and K a,b corresponds to one of the kernels. If the convolutional layer 611 acts on the previous node layer 610 with A channels and outputs the subsequent node layer 612 with B channels, there are A·B independent d-dimensional kernels K a,b .
[0070] Generally, in the convolutional neural network 600, an activation function is used. In this embodiment, ReLU (acronym for "Rectified Linear Unit") is used, where R(z) = max(0, z), such that the operation of the convolutional layer 611 in the two-dimensional example is:
[0071] It is also possible to use other activation functions, such as ELU (acronym for "Exponential Linear Unit"), LeakyReLU, Sigmoid, Tanh, or Softmax.
[0072] In the shown embodiment, the input layer 610 includes 36 nodes 620 arranged as a two-dimensional 6×6 matrix. The first hidden node layer 612 includes 72 nodes 622 arranged as two two-dimensional 6×6 matrices, each of which is the result of the convolution of the values of the input layer with a 3×3 kernel within the convolutional layer 611. Equivalently, the nodes 622 of the first hidden node layer 612 can be interpreted as arranged as a three-dimensional 2×6×6 matrix, where the first dimension corresponds to the channel dimension.
[0073] The advantage of using the convolutional layer 611 is that the spatial local correlation of the input data can be exploited by enforcing a local connectivity pattern between the nodes of adjacent layers, especially by connecting each node only to a smaller region of the nodes of the previous layer.
[0074] The pooling layer 613 is a connection layer between the previous node layer 612 (with node values x(n-1)) and the subsequent node layer 614 (with node values x(n)). In particular, the pooling layer 613 can be characterized by an activation function and the structure and weights of the edges that form a pooling operation based on a non-linear pooling function f. For example, in the two-dimensional case, the value x(n) of the node 624 in the subsequent node layer 614 can be calculated based on the value x(n-1) of the node 622 in the previous node layer 612 as follows:
[0075] In other words, by using the pooling layer 613, the number of nodes 622, 624 can be reduced in the following way: a single node 622 in the subsequent node layer 614 replaces a number d1·d2 of adjacent nodes 622 in the previous node layer 612, and this replacement is calculated as a function of the values of the said number of adjacent nodes. In particular, the pooling function f can be a maximum function, an average value, or an L2 norm. In particular, for the pooling layer 613, the weights of the incoming edges are fixed and not modified through training.
[0076] The advantage of using the pooling layer 613 is that it reduces the number of nodes 622, 624 and the number of parameters. This results in a reduction in the computational amount in the network and control of overfitting.
[0077] In the shown embodiment, the pooling layer 613 is a max pooling layer, which replaces four adjacent nodes with only one node, and this value is the maximum value among the values of the four adjacent nodes. Max pooling is applied to each d-dimensional matrix of the previous layer; in this embodiment, max pooling is applied to each of the two two-dimensional matrices, thereby reducing the number of nodes from 72 to 18.
[0078] Generally speaking, the last layer of the convolutional neural network 600 is the fully connected layer 615. The fully connected layer 615 is a connection layer between the previous node layer 614 and the subsequent node layer 616. The fully connected layer 613 can be characterized by the fact that most of the edges, especially all the edges, exist between the nodes 614 of the previous node layer 614 and the nodes 616 of the subsequent node layer, and the weights of each of these edges can be adjusted individually.
[0079] In this embodiment, the nodes 624 of the previous node layer 614 of the fully connected layer 615 are shown as a two-dimensional matrix and additionally as unconnected nodes (indicated as a row of nodes, where the number of nodes is reduced for better renderability). This operation is also referred to as "flattening". In this embodiment, the number of nodes 626 in the subsequent node layer 616 of the fully connected layer 615 is less than the number of nodes 624 in the previous node layer 614. Alternatively, the number of nodes 626 can be equal to or greater.
[0080] In addition, in this embodiment, the Softmax activation function is used within the fully connected layer 615. By applying the Softmax function, the sum of the values of all nodes 626 in the output layer 616 is 1, and all values of all nodes 626 in the output layer 616 are real numbers between 0 and 1. In particular, if the convolutional neural network 600 is used to classify input data, the values in the output layer 616 can be interpreted as the probabilities that the input data falls into one of the different classes.
[0081] In particular, the convolutional neural network 600 can be trained based on the backpropagation algorithm. To prevent overfitting, regularization methods can be used, such as dropout of nodes 620...624, random pooling, use of artificial data, weight decay based on L1 or L2 norms or max norm constraints.
[0082] According to one aspect, a machine learning model can include one or more Residual Networks (ResNet). In particular, a ResNet is an artificial neural network that includes at least one jump or skip connection for jumping on at least one layer of the artificial neural network. In particular, a ResNet can be a convolutional neural network that includes one or more skip connections that respectively skip one or more convolutional layers. According to some examples, a ResNet can be represented as an m-layer ResNet, where m is the number of layers in the corresponding architecture, and according to some examples, can take the values of 34, 50, 101, or 152. According to some examples, such an m-layer ResNet can correspondingly include (m - 2) / 2 skip connections.
[0083] A skip connection can be regarded as a bypass that directly feeds the output of a previous layer to a layer following the one or more bypassed layers through the one or more bypassed layers. Instead of having to directly fit the desired mapping, the bypassed layers will have to fit the residual mapping that "balances" the directly fed output.
[0084] The fitting residual mapping is computationally easier to optimize than the directional mapping. More importantly, this alleviates the problem of vanishing / exploding gradients during optimization when training machine learning models: if the bypassed layer encounters such a problem, its contribution can be skipped by regularizing the directly fed output. Therefore, the advantage of using ResNets is that much deeper networks can be trained.
[0085] In particular, a recurrent machine learning model is a machine learning model whose output depends not only on the input values and the machine learning model parameters adapted through the training process, but also on a hidden state vector, where the hidden state vector is based on previous inputs for the recurrent machine learning model. In particular, the recurrent machine learning model can include additional storage states or additional structures that incorporate time delays or include feedback loops.
[0086] In particular, the underlying structure of the recurrent machine learning model can be a neural network, which can be represented as a recurrent neural network. Such a recurrent neural network can be described as an artificial neural network where the connections between nodes form a directed graph along a time series. In particular, the recurrent neural network can be interpreted as a directed acyclic graph. In particular, the recurrent neural network can be a finite impulse recurrent neural network or an infinite impulse recurrent neural network (where the finite impulse network can be unfolded and replaced with a strictly feedforward neural network, and the infinite impulse network cannot be unfolded and replaced with a strictly feedforward neural network).
[0087] In particular, training a recurrent neural network can be based on the BPTT algorithm (acronym for "Backpropagation Through Time"), the RTRL algorithm (acronym for "Real - Time Recurrent Learning"), and / or genetic algorithms.
[0088] By using a recurrent machine learning model, input data including variable - length sequences can be used. In particular, this means that the method cannot be used only for a fixed number of input data sets (and needs to be trained differently for each other number of input data sets used as input), but can be used for any number of input data sets. This means that, independent of the number of input data sets included in different sequences, the entire training data set can be used during training, and the training data is not reduced to training data corresponding to a certain number of consecutive input data sets.
[0089] Figure 7 Both the recurrent representation 702 and the unfolded representation 704 show the schematic structure of a recurrent machine learning model F, which can be used to implement one or more of the machine learning models described herein. The recurrent machine learning model takes a number of input data sets x, x 1 ……x N706 is taken as input and a corresponding set of output data sets y, y 1 ……y N 708 are created. Additionally, the output depends on the so-called hidden vectors h, h 1 ……h N 710, and the hidden vectors implicitly include information about the input data set that was previously used as the input to the recurrent machine learning model F 712. By using these hidden vectors h, h 1 ……h N 710, the sequential nature of the input data set can be exploited.
[0090] In a single processing step, the recurrent machine learning model F 712 takes the hidden vector h n-1 created in the previous step, as well as the input data set x n as input. Within this step, the recurrent machine learning model F generates an updated hidden vector h n and an output data set y n as output. In other words, one processing step computes (y n , h n ) = F(x n , h n-1 ), or by splitting the recurrent machine learning model F 712 into a part F(y) that computes the output data and a part F(h) that computes the hidden vector, one processing step computes y n = F (y) (x n , h n-1 ) and h n = F (h) (x n , h n-1 ). For the first processing step, h 0 can be randomly selected or filled with all entries being zero. The parameters of the recurrent machine learning model F 712 that have been trained based on the training data set do not change between different processing steps.
[0091] In particular, the output data and the hidden vector of a processing step depend on all the previous input data sets used in the previous steps. y n = F (y) (x n , F (h) (x n-1 , h n-2 )) and h n = F(h)(x n , F (h) (x n-1 , h n-2 ))).
[0092] The systems, devices, and methods described herein can be implemented using digital circuitry, or using one or more computers employing well-known computer processors, memory units, storage devices, computer software, and other components. Generally, a computer includes a processor for executing instructions and one or more memories for storing instructions and data. The computer may also include or may be coupled to one or more mass storage devices, such as one or more disks, internal hard drives and removable disks, magneto-optical disks, optical disks, and the like.
[0093] The systems, devices, and methods described herein can be implemented using computers operating in a client-server relationship. Generally, in such systems, the client computers are located at a distance from the server computer and interact via a network. The client-server relationship can be defined and controlled by computer programs running on the respective client and server computers.
[0094] The systems, devices, and methods described herein can be implemented within a network-based cloud computing system. In such a network-based cloud computing system, a server or another processor connected to the network communicates with one or more client computers via the network. The client computers can communicate with the server via, for example, a web browser application resident and operating on the client computers. The client computers can store data on the server and access the data via the network. The client computers can transmit requests for data or requests for online services to the server via the network. The server can perform the requested services and provide the data to the client computer(s). The server can also transmit data adapted to cause the client computer to perform specified functions (e.g., perform calculations, display specified data on a screen, etc.). For example, the server can transmit requests adapted to cause the client computer to perform one or more steps or functions of the methods and workflows described herein (including Figure 1 , Figure 3 or Figure 4 of the one or more steps or functions). Certain steps or functions of the methods and workflows described herein (including Figure 1 , Figure 3 or Figure 4 of the one or more steps or functions) can be performed by the server or by another processor in the network-based cloud computing system. Certain steps or functions of the methods and workflows described herein (including Figure 1 , Figure 3 or Figure 4 of the one or more steps) can be performed by the client computers in the network-based cloud computing system. The steps or functions of the methods and workflows described herein (including Figure 1 , Figure 3 or Figure 4One or more steps) can be performed by the server and / or by client computers in a network-based cloud computing system in any combination.
[0095] The systems, devices, and methods described herein can be implemented using a computer program product tangibly embodied in an information carrier (e.g., embodied in a non-transitory machine-readable storage device) for execution by a programmable processor; and the methods and workflow steps described herein (including Figure 1 , Figure 3 or Figure 4 One or more steps or functions) can be implemented using one or more computer programs executable by such a processor. A computer program is a set of computer program instructions that can be used directly or indirectly in a computer to perform an activity or bring about a result. A computer program can be written in any form of programming language (including compiled or interpreted languages), and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0096] Figure 8 Figure 802 depicts a high-level block diagram of an example computer 802 that can be used to implement the systems, devices, and methods described herein. Computer 802 includes a processor 804 operatively coupled to a data storage device 812 and a memory 810. Processor 804 controls such operations by executing computer program instructions that define the overall operation of computer 802. The computer program instructions can be stored in data storage device 812, or other computer-readable media, and are loaded into memory 810 when desired for execution of the computer program instructions. Thus, Figure 1 , Figure 3 or Figure 4 The methods and workflow steps or functions can be defined by computer program instructions stored in memory 810 and / or data storage device 812, and can be controlled by processor 804 executing the computer program instructions. For example, the computer program instructions can be implemented as computer-executable code programmed by those skilled in the art to perform Figure 1 , Figure 3 or Figure 4 The methods and workflow steps or functions. Thus, by executing the computer program instructions, processor 804 performs Figure 1 , Figure 3 or Figure 4 The methods and workflow steps or functions. Computer 802 can also include one or more network interfaces 806 for communicating with other devices via a network. Computer 802 can also include one or more input / output devices 808 (e.g., a display, keyboard, mouse, speakers, buttons, etc.) that enable a user to interact with computer 802.
[0097] The processor 804 can include both a general - purpose microprocessor and a special - purpose microprocessor, and can be the sole processor of the computer 802 or one of multiple processors. For example, the processor 804 can include one or more central processing units (CPUs). The processor 804, the data storage device 812, and / or the memory 810 can include one or more application - specific integrated circuits (ASICs) and / or one or more field - programmable gate arrays (FPGAs), be supplemented by them, or be incorporated therein.
[0098] Both the data storage device 812 and the memory 810 include tangible non - transitory computer - readable storage media. The data storage device 812 and the memory 810 can both include high - speed random - access memory, such as dynamic random - access memory (DRAM), static random - access memory (SRAM), double - data - rate synchronous dynamic random - access memory (DDR RAM), or other random - access solid - state memory devices, and can include non - volatile memory, such as one or more disk storage devices (such as internal hard disks and removable disks), magneto - optical storage devices, optical disk storage devices, flash memory devices, semiconductor memory devices (such as erasable programmable read - only memory (EPROM), electrically erasable programmable read - only memory (EEPROM), compact disc read - only memory (CD - ROM), digital versatile disc read - only memory (DVD - ROM) discs) or other non - volatile solid - state storage devices.
[0099] The input / output device 808 can include peripheral devices, such as printers, scanners, displays, etc. For example, the input / output device 808 can include a display device (such as a cathode - ray tube (CRT) or a liquid - crystal display (LCD) monitor) for displaying information to the user, a keyboard, and a pointing device such as a mouse or a trackball, through which the user can provide input to the computer 802.
[0100] The image acquisition device 814 can be connected to the computer 802 to input image data (e.g., medical images) into the computer 802. It is possible to implement the image acquisition device 814 and the computer 802 as one device. It is also possible for the image acquisition device 814 and the computer 802 to communicate wirelessly through a network. In a possible embodiment, the computer 802 can be located remotely relative to the image acquisition device 814.
[0101] Any or all of the systems, apparatuses, and methods discussed herein can be implemented using one or more computers (such as the computer 802).
[0102] Those skilled in the art will recognize that the implementation of an actual computer or computer system may have other structures and may also include other components, and for illustrative purposes, Figure 8 is a high-level representation of some of the components in such a computer.
[0103] Individuals with male or female identities are also included in this term, independent of the use of grammatical terms.
[0104] The foregoing specific embodiments should be understood in every aspect as illustrative and exemplary, rather than restrictive, and the scope of the invention disclosed herein is not determined by this specific embodiment, but is determined from the claims as interpreted in accordance with the full breadth permitted by patent law. It is to be understood that the embodiments shown and described herein merely illustrate the principles of the invention, and those skilled in the art can implement various modifications without departing from the scope and spirit of the invention. Those skilled in the art can implement various other combinations of features without departing from the scope and spirit of the invention.
[0105] The following is a list of non-limiting illustrative embodiments disclosed herein: Illustrative Embodiment 1. A computer-implemented method, comprising: receiving 1) one or more input medical images of a patient and 2) task instructions for performing a medical imaging analysis task; encoding the one or more input medical images into imaging features using an image encoder network; encoding the task instructions into text features using a text encoder network; performing a medical imaging analysis task based on the imaging features and the text features using a machine learning-based task network; and outputting the result of the medical imaging analysis task.
[0106] Illustrative Embodiment 2. The computer-implemented method according to Illustrative Embodiment 1, wherein the task instructions include a reference to an image region in at least one of the one or more input medical images.
[0107] Illustrative Embodiment 3. The computer-implemented method according to one of the foregoing embodiments, wherein the task instructions include anatomical knowledge and task knowledge. The task knowledge includes at least one of a description of an anatomical abnormality, how to represent an anatomical abnormality, and how an anatomical abnormality can be detected in the one or more input medical images.
[0108] Illustrative Embodiment 4. The computer-implemented method according to one of the foregoing embodiments, wherein the task instructions are user-defined.
[0109] Illustrative Embodiment 5. The computer-implemented method according to one of the foregoing embodiments, wherein the text features and the imaging features are aligned in the same latent space.
[0110] Exemplary Embodiment 6. The computer-implemented method according to one of the foregoing embodiments, wherein the machine learning-based task network is trained using self-supervised learning based on unlabeled training medical images and text.
[0111] Exemplary Embodiment 7. The computer-implemented method according to one of the foregoing embodiments, wherein the machine learning-based task network is trained using few-shot learning with labeled training medical images and labeled task descriptions.
[0112] Exemplary Embodiment 8. The computer-implemented method according to one of the foregoing embodiments, wherein the machine learning-based task network includes a large language model (LLM)-based task network.
[0113] Exemplary Embodiment 9. The computer-implemented method according to one of the foregoing embodiments, wherein: receiving 1) one or more input medical images of a patient and 2) task instructions for performing a medical imaging analysis task includes receiving text-based medical data of the patient; and encoding the task instructions into text features using a text encoder network includes encoding the task instructions and the text-based medical data into text features using the text encoder network.
[0114] Exemplary Embodiment 10. An apparatus, comprising: means for receiving 1) one or more input medical images of a patient and 2) task instructions for performing a medical imaging analysis task; means for encoding the one or more input medical images into imaging features using an image encoder network; means for encoding the task instructions into text features using a text encoder network; means for performing a medical imaging analysis task based on the imaging features and the text features using a machine learning-based task network; and means for outputting the result of the medical imaging analysis task.
[0115] Exemplary Embodiment 11. The apparatus according to Exemplary Embodiment 10, wherein the task instructions include a reference to an image region in at least one of the one or more input medical images.
[0116] Exemplary Embodiment 12. The apparatus according to one of Exemplary Embodiments 10-11, wherein the task instructions include anatomical knowledge and task knowledge. The task knowledge includes at least one of a description of an anatomical abnormality, how to represent an anatomical abnormality, and how an anatomical abnormality can be detected in the one or more input medical images.
[0117] Exemplary Embodiment 13. The apparatus according to one of Exemplary Embodiments 10-12, wherein the task instructions are user-defined.
[0118] Exemplary Embodiment 14. A non-transitory computer-readable medium including instructions that, when executed by a computer, cause the computer to perform operations including: receiving 1) one or more input medical images of a patient and 2) task instructions for performing a medical imaging analysis task; encoding the one or more input medical images into imaging features using an image encoder network; encoding the task instructions into text features using a text encoder network; performing a medical imaging analysis task using a machine learning-based task network based on the imaging features and the text features; and outputting a result of the medical imaging analysis task.
[0119] Exemplary Embodiment 15. The non-transitory computer-readable medium according to Exemplary Embodiment 14, wherein the text features and the imaging features are aligned in the same latent space.
[0120] Exemplary Embodiment 16. The non-transitory computer-readable medium according to one of Exemplary Embodiments 14-15, wherein the machine learning-based task network is trained using self-supervised learning based on unlabeled training medical images and text.
[0121] Exemplary Embodiment 17. The non-transitory computer-readable medium according to one of Exemplary Embodiments 14-16, wherein the machine learning-based task network is trained using few-shot learning using labeled training medical images and labeled task descriptions.
[0122] Exemplary Embodiment 18. A computer-implemented method including: receiving 1) one or more training medical images and 2) training task instructions for performing a medical imaging analysis task; encoding the one or more training medical images into imaging features using a pre-trained image encoder network; encoding the training task instructions into text features using a pre-trained text encoder network; training a machine learning-based task network based on the imaging features and the text features for performing a medical imaging analysis task; and outputting the trained machine learning-based task network.
[0123] Exemplary Embodiment 19. The computer-implemented method according to Exemplary Embodiment 18, wherein the machine learning-based task network is trained using self-supervised learning based on unlabeled training medical images and text.
[0124] Exemplary Embodiment 20. The computer-implemented method according to one of Exemplary Embodiments 18-19, wherein the machine learning-based task network is trained using few-shot learning using labeled training medical images and labeled task descriptions.
Claims
1. A computer-implemented method comprising: receiving 1) one or more input medical images of a patient and 2) a task instruction for performing a medical imaging analysis task; encoding the one or more input medical images into imaging features using an image encoder network; Use a text encoder network to encode task instructions into text features; Performing medical imaging analysis tasks based on imaging features and text features using a machine learning-based task network; and Output the results of the medical imaging analysis task. 2 . The computer-implemented method of claim 1 , wherein the task instruction comprises a reference to an image region in at least one of the one or more input medical images.
3. A computer-implemented method according to claim 1, wherein the task instructions include anatomical knowledge and task knowledge, the task knowledge including a description of anatomical abnormalities, how to represent anatomical abnormalities, and how at least one of the anatomical abnormalities can be detected in the one or more input medical images. The computer-implemented method of claim 1 , wherein the task instructions are user-defined.
5. The computer-implemented method of claim 1, wherein the text features and the imaging features are aligned in the same latent space.
6. The computer-implemented method of claim 1, wherein the machine learning-based task network is trained using self-supervised learning based on unlabeled training medical images and text.
7. The computer-implemented method of claim 1, wherein the machine learning based task network is trained using few-shot learning using annotated training medical images and annotated task descriptions.
8. The computer-implemented method of claim 1, wherein the machine learning based task network comprises an LLM (Large Language Model) based task network.
9. The computer-implemented method of claim 1 , wherein: Receiving 1) one or more input medical images of a patient and 2) task instructions for performing a medical imaging analysis task includes receiving text-based medical data of the patient; and Encoding task instructions into text features using a text encoder network includes encoding task instructions and text-based medical data into text features using a text encoder network.
10. A device comprising: means for receiving 1) one or more input medical images of a patient and 2) task instructions for performing a medical imaging analysis task; means for encoding the one or more input medical images into imaging features using an image encoder network; means for encoding task instructions into text features using a text encoder network; Means for performing a medical imaging analysis task based on imaging features and text features using a machine learning based task network; as well as Means for outputting the results of a medical imaging analysis task.
11. The apparatus of claim 10, wherein the task instruction comprises a reference to an image region in at least one of the one or more input medical images.
12. The apparatus of claim 10, wherein the task instructions include anatomical knowledge and task knowledge, the task knowledge including at least one of a description of an anatomical abnormality, how to represent an anatomical abnormality, and how the anatomical abnormality can be detected in the one or more input medical images.
13. The apparatus of claim 10, wherein the task instructions are user defined.
14. A non-transitory computer-readable medium comprising instructions that, when executed by a computer, cause the computer to perform operations comprising: receiving 1) one or more input medical images of a patient and 2) a task instruction for performing a medical imaging analysis task; encoding the one or more input medical images into imaging features using an image encoder network; Use a text encoder network to encode task instructions into text features; Performing medical imaging analysis tasks based on imaging features and text features using a machine learning-based task network; and Output the results of the medical imaging analysis task.
15. The non-transitory computer-readable medium of claim 14, wherein the text features and the imaging features are aligned in the same latent space.
16. The non-transitory computer-readable medium of claim 14, wherein the machine learning based task network is trained using self-supervised learning based on unlabeled training medical images and text.
17. The non-transitory computer-readable medium of claim 14, wherein the machine learning based task network is trained using few-shot learning using annotated training medical images and annotated task descriptions.
18. A computer-implemented method comprising: receiving 1) one or more training medical images and 2) a training task instruction for performing a medical imaging analysis task; encoding the one or more training medical images into imaging features using a pre-trained image encoder network; Use a pre-trained text encoder network to encode training task instructions into text features; Training a machine learning-based task network based on imaging features and text features to perform medical imaging analysis tasks; as well as Outputs a trained machine learning-based task network.
19. The computer-implemented method of claim 18, wherein the machine learning based task network is trained using self-supervised learning based on unlabeled training medical images and text.
20. The computer-implemented method of claim 18, wherein the machine learning based task network is trained using few-shot learning using annotated training medical images and annotated task descriptions.