Anatomical perceptual visual language model for medical imaging analysis

By combining visual language models with semantic language representation, anatomical perception instructions are generated and visual language models are trained to perform medical imaging analysis tasks, which solves the shortcomings of traditional models in anatomical constraints and achieves higher accuracy and robustness.

CN120473092APending Publication Date: 2025-08-12SIEMENS HEALTHINEERS AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510154165.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-12
Filing Date
2025-02-12
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Traditional machine learning-based vision models are difficult to implement anatomical constraints, especially precise positioning and medical interpretation of complex anatomical structures, resulting in insufficient accuracy and robustness of medical imaging analysis tasks.

Method used

By combining the anatomically perceived visual language model (VLM) with semantic language representation, using text-based reports to generate instructions, the visual language model is trained to perform medical imaging analysis tasks, and the accurate positioning and association of anatomical features is achieved.

Benefits of technology

Improves the accuracy and anatomical robustness of medical imaging analysis tasks, enhances the model's positioning ability in complex anatomical structures, and reduces error occurrence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120473092A_ABST
    Figure CN120473092A_ABST
Patent Text Reader

Abstract

Systems and methods are provided for performing one or more medical imaging analysis tasks using a visual language model. One or more input medical images are received. An image embedding is extracted from one or more input medical images. The one or more medical imaging analysis tasks are performed based on image embedding extracted from the one or more input medical images using the trained visual language model. And outputting a result of the one or more medical imaging analysis tasks. A trained visual language model is trained by receiving one or more training medical images and a text-based report associated with the one or more training medical images, extracting an image embedding from the one or more training medical images, generating one or more instructions based on the text-based report using the language model, and outputting the one or more instructions. And training a visual language model to perform the one or more medical imaging analysis tasks based on the image embedding extracted from the one or more training medical images and the one or more generated instructions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to machine learning based medical imaging analysis, and in particular to an anatomically aware VLM (Visual Language Model) for medical imaging analysis. Background Art

[0002] Recently, machine learning-based visual models have been proposed for performing various medical imaging analysis tasks. Traditionally, such visual models rely on pixel-based imaging data of medical images to directly extract insights for performing medical imaging analysis tasks. However, such traditional visual models are inherently difficult to enforce even simple anatomical constraints (e.g., the precise positioning of the left atrium in the heart compared to the right atrium), let alone complex anatomical constraints (e.g., the relative positioning of one organ with respect to any other organ). In addition, such traditional visual models have difficulty enforcing constraints on specific medical interpretations. Summary of the Invention

[0003] Embodiments described herein provide a machine learning-based visual language model that associates anatomical visual representations with semantic language representations for performing medical imaging analysis tasks.

[0004] According to one or more embodiments, a system and method for performing one or more medical imaging analysis tasks using a visual language model are provided. One or more input medical images are received. Image embeddings are extracted from the one or more input medical images. One or more medical imaging analysis tasks are performed based on the image embeddings extracted from the one or more input medical images using the trained visual language model. Results of the one or more medical imaging analysis tasks are output. The trained visual language model is trained by: receiving one or more training medical images and text-based reports associated with the one or more training medical images, extracting image embeddings from the one or more training medical images, generating one or more instructions based on the text-based reports using a language model, and training the visual language model to perform the one or more medical imaging analysis tasks based on the image embeddings extracted from the one or more training medical images and the one or more generated instructions.

[0005] In one embodiment, the one or more instructions are further generated based on a plurality of predefined templates. The plurality of predefined templates include different initial instructions for extracting information from the text-based report and generating the one or more instructions.

[0006] In one embodiment, one or more instructions are generated for associating anatomical features depicted in one or more training medical images with textual anatomical descriptors, for associating anatomical features depicted in one or more training medical images with each other, and / or for associating textual anatomical descriptors with image findings of one or more training medical images. The image findings may include quantitative image findings of the one or more input medical images.

[0007] In one embodiment, an instruction embedding representing one or more instructions is generated, the image embedding and the instruction embedding are combined, and a result of one or more medical imaging analysis tasks is generated based on the combined image embedding and instruction embedding.

[0008] According to one embodiment, a system and method for training a visual language model to perform one or more medical imaging analysis tasks is provided. One or more training medical images and text-based reports associated with the one or more training medical images are received. Image embeddings are extracted from the one or more training medical images. One or more instructions are generated based on the text-based reports using a language model. The visual language model is trained to perform the one or more medical imaging analysis tasks based on the image embeddings extracted from the one or more training medical images and the one or more generated instructions.

[0009] In one embodiment, one or more instructions are generated for associating anatomical features depicted in one or more training medical images with textual anatomical descriptors, for associating anatomical features depicted in one or more training medical images with each other, and / or for associating textual anatomical descriptors with image findings of one or more training medical images. The image findings may include quantitative image findings of the one or more input medical images.

[0010] These and other advantages of the present invention will become apparent to those of ordinary skill in the art upon reference to the following detailed description and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 A workflow for training a visual language model for performing medical imaging analysis tasks according to one or more embodiments is shown;

[0012] Figure 2 A method for training a visual language model for performing medical imaging analysis tasks according to one or more embodiments is shown;

[0013] Figure 3 A method for performing medical imaging analysis tasks using a visual language model according to one or more embodiments is shown;

[0014] Figure 4An exemplary artificial neural network that can be used to implement one or more embodiments is shown;

[0015] Figure 5 A convolutional neural network that can be used to implement one or more embodiments is shown;

[0016] Figure 6 shows a schematic structure of a recurrent machine learning model that can be used to implement one or more embodiments; and

[0017] Figure 7 A high-level block diagram of a computer that can be used to implement one or more embodiments is shown. DETAILED DESCRIPTION

[0018] The present invention generally relates to methods and systems for anatomically perceptual visual language models for medical imaging analysis. Embodiments of the present invention are described herein to provide an intuitive understanding of such methods and systems. A digital image typically consists of digital representations of one or more objects (or shapes). In this document, digital representations of objects are typically described in terms of identifying and manipulating the objects. Such manipulations are virtual manipulations performed in the memory or other circuitry / hardware of a computer system. Therefore, it should be understood that embodiments of the present invention can be performed within a computer system using data stored within the computer system. Additionally, references herein to pixels of an image may equally refer to voxels of the image, and vice versa.

[0019] The embodiments described herein provide a general-purpose anatomically-aware visual language model for medical imaging analysis and clinical decision support. The embodiments described herein utilize instruction tuning to improve the ability of a visual language model to understand visual signals and associate them with textual anatomical descriptors. Utilizing medical reports, the embodiments described herein autonomously generate anatomically correct instructions for instructing a visual language model to perform medical imaging analysis tasks. By combining medical images with text-based reports using a large language model, the visual language model is trained for anatomical understanding, making the visual language model less prone to errors. Advantageously, the visual language model according to the embodiments described herein thereby performs medical imaging analysis tasks with increased accuracy and anatomical robustness compared to traditional approaches.

[0020] Figure 1 A workflow 100 is shown for training a visual language model for performing medical imaging analysis tasks in accordance with one or more embodiments. Figure 2 A method 200 for training a visual language model for performing medical imaging analysis tasks according to one or more embodiments is shown. The steps of the method 200 may be performed by one or more suitable computing devices (such as, for example, Figure 7 The computer 702) is executed. Figure 1and Figure 2 . Figure 1 Workflow 100 and Figure 2 The steps / operations of method 200 are performed during a previous offline or training phase for training a visual language model. Once trained, the trained visual language model is applied during an online or inference phase, for example, to perform Figure 3 Method 300.

[0021] exist Figure 2 At step 202, one or more training medical images and a text-based report associated with the one or more training medical images are received. In one example, Figure 1 As shown in workflow 100 , the one or more training medical images are medical images 102 , and the text-based report is report 104 .

[0022] The one or more training medical images may depict one or more anatomical objects of interest, such as, for example, an organ, a bone, a blood vessel, a tumor or an abnormality or pathology, or any other suitable anatomical object(s) of interest of the patient. The one or more training medical images may have any suitable imaging modality, such as, for example, CT (computed tomography), MRI (magnetic resonance imaging), US (ultrasound), x-ray (e.g., angiography or fluoroscopy), PET (positron emission tomography), SPECT (single photon emission computed tomography), or any other medical imaging modality or combination of medical imaging modalities. The one or more training medical images may be 2D (two-dimensional) images and / or 3D (three-dimensional) volumes. The one or more training medical images may be annotated. For example, the annotations may be expert annotations, predictions generated by a visual model, information extracted from text-based reports, longitudinal clinical data (e.g., disease progression or events), etc.

[0023] The text-based report includes text associated with one or more training medical images. The text-based report may include, for example, the patient's demographic information, vital signs, medical history, family history, laboratory results, medications, measurements, and information extracted from the medical images. In one embodiment, the text-based report may be a medical report that includes medical findings from one or more training medical images. In one embodiment, the text-based report includes textual anatomical descriptors of anatomical objects depicted in one or more training medical images. The text-based report may be manually generated by a user (e.g., a clinician) or automatically generated by, for example, a machine learning-based model.

[0024] The one or more training medical images and the text-based report may be received, for example, by directly from the image acquisition device (e.g., Figure 7The image acquisition device 714 receives one or more training medical images by reading from a storage device or memory (e.g., Figure 7 710 or storage device 712 of computer 702), or by loading one or more training medical images and / or text-based reports from a remote computer system (e.g., Figure 7 The computer system 702 receives one or more training medical images and / or text-based reports. Such a computer system or remote computer system may include one or more patient databases, such as, for example, EHR (electronic health records), EMR (electronic medical records), PHR (personal health records), HIS (health information system), RIS (radiology information system), PACS (picture archiving and communication system), LIMS (laboratory information management system), or any other suitable database or system.

[0025] exist Figure 2 At step 204, image embeddings are extracted from one or more training medical images. In one embodiment, image embeddings are extracted from one or more training medical images using a machine learning-based feature extraction model. In one example, Figure 1 As shown in the workflow 100, image embeddings 108 are extracted from the medical image 102 using an embedding module (vision model) 106. However, any other suitable method may be used to extract image embeddings from one or more training medical images.

[0026] The machine learning-based feature extraction model can be a pre-trained vision model, such as an autoencoder. The machine learning-based feature extraction network receives one or more training medical images as input and generates image embeddings as output. Image embeddings represent one or more training medical images as low-level latent features, represented by numerical feature vectors. Image embeddings are similar to tokens in text embedding.

[0027] exist Figure 2 At step 206, one or more instructions are generated based on the text-based report using a language model. In one embodiment, the language model is an LLM (Large Language Model). In one example, Figure 1 As shown in the workflow 100, the one or more instructions are instructions 112 generated by the LLM 110 based on the report 104. However, the language model may be any other suitable machine learning-based language model (eg, a small language model).

[0028] The one or more generated instructions are guidelines or instructions that are provided to direct the behavior and output of the visual language model. The one or more generated instructions provide anatomical cues to the visual language model so that the visual language model learns to associate the anatomical cues with information from one or more training medical images. The one or more generated instructions may include, for example, commands, questions, constraints, requirements, contextual information, or any other guidelines or instructions that direct the behavior and output of the visual language model. In one example, the one or more generated instructions include: "Segment the most calcified artery from the given image," which involves an understanding of arteries and calcifications and a quantitative estimation of calcifications. In another example, the one or more generated instructions instruct the language model to segment a specific artery (e.g., "Segment the circumflex branch of the left coronary artery in the given image") or detect disease (e.g., "Is there plaque on the main branch of the left anterior descending artery in the given image"), which involves an understanding of anatomical segments and plaque classification. In another example, the one or more generated instructions may inquire about treatment planning: "Will a 20 mm stent completely cover the proximal left anterior descending artery lesion?", which involves an understanding of anatomical segments, lesion lengths, and quantitative comparisons. The one or more generated instructions are generated as instruction embeddings. Instruction embedding is a low-level latent feature representation of one or more generated instructions as a numerical feature vector.

[0029] In one embodiment, template technology is applied to generate one or more instructions. In this embodiment, a language model is queried based on multiple predefined templates. The multiple predefined templates include different initial instructions for extracting specific information from a text-based report and / or for generating one or more instructions. In one embodiment, the multiple predefined templates may include a template containing a question and multiple possible answer options. For example, the question may be "Which of the following multiple possible answer options is the calcified portion of the artery", and the multiple possible answer options may be (1) left diagonal, (2) left circumflex, (3) left main branch, and (4) distal circumflex. In another example, the multiple predefined templates can help identify whether there is a pathology. The language model will identify which artery has calcification from the text-based report and generate an instruction that means "segment the artery with calcification". In another example, given more context (such as knowing that the artery is a side branch of the right coronary artery), the multiple predefined templates may include a template containing the instruction "only segment the diagonal branch of the right coronary artery with calcification". Using this information, the visual language model can be trained to know what is diagonal, which is the right coronary artery, etc. The plurality of predefined templates may be generated by a user or according to any other suitable method. The language model receives as input one or more prompts including a text-based report and a plurality of predefined templates, and generates as output an instruction embedding representing one or more generated instructions. The prompt is an input to the language model for generating a response. The prompt may be received from the computer system via one or more APIs (application programming interfaces), or from a user interacting with the computer system.

[0030] In one embodiment, one or more generated instructions are organized into a hierarchical curriculum, where instructions become increasingly complex. In one example, in the context of coronary arteries, a first set of instructions may associate anatomical features depicted in one or more input medical images with textual anatomical descriptors embedded within a first language model. For example, the first set of instructions may include: 1) "Segment the left anterior descending aorta in the given image," which will generate a segmentation mask of the left anterior descending aorta; 2) "Segment the aortic root in the given image," which will generate a segmentation mask of the aortic root; and so on. Like textual anatomical descriptors, associating different anatomical features depicted in one or more input medical images with each other may also be relevant. Thus, a second set of instructions may associate anatomical features of the entire cardiac structure (e.g., four chambers, valves, myocardium, etc.). Once anatomical features are associated, a third set of instructions may focus on image findings, for example, associating textual anatomical descriptors with image findings (e.g., lesions, calcifications, occlusions, etc.) of the one or more input medical images. For example, the third set of instructions may include: "Find calcifications in the left anterior descending aorta in the given image," which will generate a mask for the calcifications. The anatomical features may include features associated with any anatomical object, such as, for example, an organ, bone, blood vessel, tumor, or abnormality or pathology, or any other suitable anatomical object(s) of interest in the patient. In some embodiments, the one or more instructions may include a binary instruction, such as: "Is there calcification in the left coronary artery," which will generate a binary "yes" or "no" response.

[0031] In one embodiment, the image findings may include quantitative image findings such that one or more instructions associate textual anatomical descriptors with quantitative image findings (e.g., size of a lesion, lesion volume, calcification volume fraction, etc.) of one or more input medical images. In this embodiment, the first language model may be prompted to, for example, determine one or more measurements on a segmentation mask and generate one or more instructions associated with the measurements. An example of such an instruction is: "In the given image, what is the size of the lesion on the right coronary artery branch?", which would generate a segmentation of the lesion and report the size of the lesion based on the segmentation.

[0032] In the case where the language model is an LLM, the LLM can be any suitable pre-trained deep learning based LLM. For example, the LLM can be based on a transformer architecture that uses a self-attention mechanism to capture long-term dependencies in text. An example of a transformer-based architecture is GPT (Generative Pre-trained Transformer), which has a multi-layer transformer decoder architecture that can be pre-trained to optimize the next word prediction task and then fine-tuned for various downstream tasks using labeled data. GPT-based LLMs can be trained and / or fine-tuned using reinforcement learning with human feedback for performing various natural language processing tasks. Other exemplary transformer-based architectures include BLOOM (BigScience Large Open Science Open Access Multilingual Language Model) and BERT (Bidirectional Encoder Representations from Transformers). In some embodiments, in addition to text-based reports and multiple predefined templates, the LLM can be a multimodal LLM that receives, for example, imaging data (e.g., one or more input medical images).

[0033] exist Figure 2 At step 208, based on the image embedding and the one or more generated instructions, the visual language model is trained to perform one or more medical imaging analysis tasks. The one or more medical imaging analysis tasks may include (one or more) any suitable medical imaging analysis tasks, such as, for example, segmentation, classification, detection, quantification, etc. In one example, Figure 1 As shown in the workflow 100 of FIG, the visual language model is VLM 114, which is used to perform medical imaging analysis tasks of segmentation, landmark detection, and report generation. VLM 114 generates segmentation masks 116-A, landmarks 116-B, and text 116-C from image embeddings 108 and instructions 112.

[0034] A visual language model is a multimodal model that receives both text-based and image-based data as input. The visual language model can be any suitable pre-trained deep learning-based visual language model. Examples of visual language models include CLIP (Contrastive Language Image Pretraining), Flamingo, and VL-BERT (Visual BERT).

[0035] To train a visual language model to perform a medical imaging analysis task, an image embedding and an instruction embedding are first combined (e.g., concatenated). The visual language model receives as input one or more prompts comprising the combined image and instruction embeddings, and generates as output the results of the medical imaging analysis task. The visual language model thus performs the medical imaging analysis task according to the one or more generated instructions represented by the instruction embeddings. The prompts may be received from the computer system via one or more APIs (application programming interfaces), or from a user interacting with the computer system. Based on the results of the medical imaging analysis task, the visual language model is trained using any suitable loss function.

[0036] exist Figure 2 At step 210, the trained visual language model is output. For example, the trained visual language model can be stored in a memory or storage device (e.g., Figure 7 710 or storage device 712 of the computer 702), or by transmitting the trained visual language model to a remote computer system (e.g., Figure 7 702) to output the trained visual language model.

[0037] Figure 3 A method 300 for performing a medical imaging analysis task using a trained visual language model according to one or more embodiments is shown. The steps of the method 300 may be performed by one or more suitable computing devices (such as, for example, Figure 7 computer 702) to execute. Figure 3 The steps of method 300 are performed during an online or inference phase using a trained visual language model. Figure 1 Workflow 100 or Figure 2 The method 200 of , the trained visual language model is trained during a previous offline or training phase.

[0038] exist Figure 3 At step 302, one or more input medical images are received. The one or more input medical images may depict one or more anatomical objects of interest. The one or more input medical images may be of any suitable imaging modality, such as, for example, CT, MRI, US, X-ray, PET, SPECT, etc., and may be 2D images and / or 3D volumes.

[0039] One or more input medical images may be received, for example, by directly obtaining the medical images from an image acquisition device (e.g., Figure 7 The image acquisition device 714 receives one or more input medical images by reading from a storage device or memory (e.g., Figure 7702 memory 710 or storage device 712) of the computer 702, or by loading one or more input medical images from a remote computer system (e.g., Figure 7 The computer system 702 receives one or more input medical images. Such a computer system or a remote computer system may include one or more patient databases.

[0040] exist Figure 3 At step 304, image embeddings are extracted from the one or more input medical images. In one embodiment, the image embeddings are extracted from the one or more input medical images using a machine learning-based feature extraction model. For example, the machine learning-based feature extraction model can be a pre-trained visual model, such as, for example, an autoencoder. However, any other suitable method can be used to extract image embeddings from the one or more training medical images. The machine learning-based feature extraction network receives one or more input medical images as input and generates image embeddings as output. Image embeddings are low-level latent features that represent the one or more input medical images as digital feature vectors.

[0041] exist Figure 6 At step 306, one or more medical imaging analysis tasks are performed based on the image embedding using the trained visual language model. The medical imaging analysis tasks may include (one or more) any suitable medical imaging analysis tasks, such as, for example, segmentation, classification, detection, quantification, etc. The trained visual language model receives the image embedding as input and generates the results of the one or more medical imaging analysis tasks as output. The visual language model may be any suitable pre-trained deep learning based visual language model, such as, for example, CLIP, Flamingo, and VL-BERT. In one embodiment, according to Figure 1 Workflow 100 or Figure 2 Method 200 to train the trained visual language model.

[0042] exist Figure 3 At step 308, the results of one or more medical imaging analysis tasks are output. For example, the results may be displayed on a display device (e.g., Figure 7 The results of one or more medical imaging analysis tasks are displayed on the I / O 708 of the computer 702, and the results of one or more medical imaging analysis tasks are displayed on the memory or storage device of the computer system (e.g., Figure 7 The results of one or more medical imaging analysis tasks are stored on a memory 710 or storage device 712 of the computer 702, or by transmitting the results of one or more medical imaging analysis tasks to a remote computer system (e.g., Figure 7 702) to output the results of one or more medical imaging analysis tasks.

[0043] In one embodiment, where the trained visual language model has been trained with extensive textual knowledge of, for example, cardiovascular disease, it can be used during inference. Figure 3 Step 306 of the process queries the trained visual language model in a broader context. For example, the visual language model may be queried: "Is a patient with a given cardiac exam image and a troponin-I value of 40 pg / mL likely to have MACE within the next year?", where pg / ml refers to picograms per milliliter and MACE refers to "major adverse cardiovascular events."

[0044] The embodiments described herein are described with respect to the claimed systems and claimed methods. Features, advantages, or alternative embodiments herein may be assigned to other claimed objects, and vice versa. In other words, claims and embodiments for systems may be modified using features described or claimed in the context of the corresponding methods. In this case, the functional features of the methods are implemented by the physical units of the system.

[0045] In addition, certain embodiments described herein are described with respect to methods and systems for utilizing trained machine learning models, and with respect to methods and systems for providing trained machine learning models. Features, advantages, or alternative embodiments herein may be assigned to other claimed objects, and vice versa. In other words, claims and embodiments for providing trained machine learning models may be improved with features described or claimed in the context of utilizing trained machine learning models, and vice versa. In particular, the datasets used in the methods and systems for utilizing trained machine learning models may have the same properties and characteristics as the corresponding datasets used in the methods and systems for providing trained machine learning models, and the trained machine learning models provided by the corresponding methods and systems may be used in the methods and systems for utilizing trained machine learning models.

[0046] Generally speaking, a trained machine learning model mimics the cognitive functions that humans associate with other human minds. In particular, by training on training data, a machine learning model can adapt to new situations and detect and infer patterns. Another term for a "trained machine learning model" is a "trained function."

[0047] In general, the parameters of a machine learning model can be adapted by means of training. In particular, supervised training, semi-supervised training, unsupervised training, reinforcement learning, and / or active learning can be used. In addition, representation learning (an alternative term is "feature learning") can be used. In particular, the parameters of a machine learning model can be iteratively adapted over several training steps. In particular, within training, a certain cost function can be minimized. In particular, within the training of a neural network, a backpropagation algorithm can be used.

[0048] In particular, the machine learning models disclosed herein, such as e.g. Figure 1 The embedding module 106, LLM 110 and VLM 114, the feature extraction network used at step 204, the language model used at step 206, and the second language model used at step 208, and Figure 2 The visual language model utilized at step 208, and Figure 3 The feature extraction network used in step 304 and Figure 3 The trained visual language model utilized at step 306 may include, for example, a neural network. In particular, the neural network may be, for example, a deep neural network, a convolutional neural network, or a convolutional deep neural network. Furthermore, the neural network may be, for example, an adversarial network, a deep adversarial network, and / or a generative adversarial network.

[0049] Figure 4 An embodiment of an artificial neural network 400 is shown that can be used to implement one or more machine learning models described herein. Alternative terms for "artificial neural network" are "neural network," "artificial neural net," or "neural net."

[0050] The artificial neural network 400 includes nodes 420, ..., 432 and edges 440, ..., 442, wherein each edge 440, ..., 442 is a directed connection from a first node 420, ..., 432 to a second node 420, ..., 432. In general, the first node 420, ..., 432 and the second node 420, ..., 432 are different nodes 420, ..., 432. The first node 420, ..., 432 and the second node 420, ..., 432 may also be the same. For example, Figure 4 , edge 440 is a directed connection from node 420 to node 423, and edge 442 is a directed connection from node 430 to node 432. The edges 440, ..., 442 from the first nodes 420, ..., 432 to the second nodes 420, ..., 432 are also labeled as "incoming edges" of the second nodes 420, ..., 432 and "outgoing edges" of the first nodes 420, ..., 432.

[0051] In this embodiment, the nodes 420, ..., 432 of the artificial neural network 400 can be arranged in layers 410, ..., 413, wherein the layers can include an inherent order introduced by edges 440, ..., 442 between the nodes 420, ..., 432. In particular, edges 440, ..., 442 can only exist between adjacent layers of nodes. In the illustrated embodiment, there is an input layer 410 that includes only nodes 420, ..., 422 and no incoming edges, an output layer 413 that includes only nodes 431, 432 and no outgoing edges, and hidden layers 411, 412 between the input layer 410 and the output layer 413. In general, the number of hidden layers 411, 412 can be arbitrarily selected. The number of nodes 420, ..., 422 within the input layer 410 is generally related to the number of input values of the neural network, and the number of nodes 431, 432 within the output layer 413 is generally related to the number of output values of the neural network.

[0052] In particular, a (real) number may be assigned as a value to each node 420, ..., 432 of the neural network 400. Here, x (n) i denotes the value of the i-th node 420, ..., 432 of the n-th layer 410, ..., 413. The values of the nodes 420, ..., 422 of the input layer 410 correspond to the input values of the neural network 400, and the values of the nodes 431, 432 of the output layer 413 correspond to the output values of the neural network 400. In addition, each edge 440, ..., 442 may include a weight as a real number, in particular, the weight is a real number in the interval [-1, 1] or in the interval [0, 1]. Here, w (m,n) i,j The weight of the edge between the i-th node 420, ..., 432 of the m-th layer 410, ..., 413 and the j-th node 420, ..., 432 of the n-th layer 410, ..., 413 is indicated. In addition, the weight w (n,n+1) i,j Defines the abbreviation w (n) i,j .

[0053] Specifically, to calculate the output value of the neural network 400, the input value is propagated through the neural network. Specifically, the values of the nodes 420, ..., 432 of the (n+1)th layer 410, ..., 413 can be calculated based on the values of the nodes 420, ..., 432 of the nth layer 410, ..., 413 by

[0054]

[0055] In this article, function f is a transfer function (another term is "activation function"). Known transfer functions are step functions, sigmoid functions (e.g., logistic function, generalized logistic function, hyperbolic tangent function, inverse tangent function, error function, smoothed step function), or rectifier functions. Transfer functions are mainly used for normalization purposes.

[0056] In particular, these values are propagated layer by layer through the neural network, where the value of the input layer 410 is given by the input of the neural network 400, where the value of the first hidden layer 411 can be calculated based on the value of the input layer 410 of the neural network, where the value of the second hidden layer 412 can be calculated based on the value of the first hidden layer 411, and so on.

[0057] To set the edge value Training data must be used to train the neural network 400. Specifically, the training data includes training input data and training output data (denoted as t i ). For the training step, the neural network 400 is applied to the training input data to generate the calculated output data. In particular, the training data and the calculated output data include a number of values that is equal to the number of nodes of the output layer.

[0058] In particular, the comparison between the calculated output data and the training data is used to recursively adapt the weights within the neural network 400 (back propagation algorithm). In particular, the weights are changed according to

[0059]

[0060] Where γ is the learning rate, and if the (n+1)th layer is not the output layer, it can be based on δ (n+1) j The quantity δ (n) j Recursively computed as

[0061]

[0062] And if the (n+1)th layer is the output layer 413, then

[0063]

[0064] where f' is the first derivative of the activation function, and t (n+1) j is the comparative training value of the j-th node of the output layer 413.

[0065] A convolutional neural network is a neural network that uses convolution operations instead of general matrix multiplications in at least one of its layers (the so-called "convolutional layer"). In particular, a convolutional layer performs a dot product of one or more convolution kernels with the input data / image of the convolutional layer, where the entries of the one or more convolution kernels are parameters or weights that are adapted through training. In particular, one can use Frobenius inner products and ReLU activation functions. A convolutional neural network can include additional layers such as pooling layers, fully connected layers, and normalization layers.

[0066] By using convolutional neural networks, input images can be processed in a very efficient manner because convolution operations based on different kernels can extract various image features, so that by adapting the weights of the convolution kernels, relevant image features can be found during training. In addition, based on the weight sharing in the convolution kernels, fewer parameters need to be trained, which prevents overfitting during the training phase and allows for faster training or more layers in the network, thereby improving the performance of the network.

[0067] Figure 5 An embodiment of a convolutional neural network 500 that can be used to implement one or more machine learning models described herein is shown. In the embodiment shown, the convolutional neural network 500 includes an input node layer 510, a convolutional layer 511, a pooling layer 513, a fully connected layer 514, and an output node layer 516, as well as hidden node layers 512 and 514. Alternatively, the convolutional neural network 500 can include several convolutional layers 511, several pooling layers 513, and several fully connected layers 515, as well as other types of layers. The order of the layers can be chosen arbitrarily, and typically the fully connected layers 515 are used as the last layers before the output layer 516.

[0068] In particular, within the convolutional neural network 500, the nodes 520, 522, 524 of the node layers 510, 512, 514 can be viewed as being arranged as a d-dimensional matrix or a d-dimensional image. In particular, in the two-dimensional case, the values of the nodes 520, 522, 524 indexed by i and j in the n-th node layer 510, 512, 514 can be denoted as x(n)[i,j]. However, the arrangement of the nodes 520, 522, 524 of a node layer 510, 512, 514 has no effect on the computations performed thereby within the convolutional neural network 500, as these are given solely by the structure and weights of the edges.

[0069] The convolution layer 511 is a connection layer between the previous node layer 510 (having node values x(n-1)) and the subsequent node layer 512 (having node values x(n)). In particular, the convolution layer 511 is characterized by the structure and weight of the input edges of the convolution operation based on a certain number of kernels. In particular, the structure and weight of the edges of the convolution layer 511 are selected so that the value x(n) of the node 522 of the subsequent node layer 512 is calculated as the convolution x(n) = K*x(n-1) based on the value x(n-1) of the node 520 of the previous node layer 510, where the convolution* is defined in the two-dimensional case as

[0070]

[0071] Here, the kernel K is a d-dimensional matrix (in this embodiment, a two-dimensional matrix), which is typically small compared to the number of nodes 520, 522 (e.g., a 3×3 matrix or a 5×5 matrix). In particular, this means that the weights of the edges in the convolution layer 511 are not independent, but are chosen so that they produce the convolution equation. In particular, for a kernel that is a 3×3 matrix, there are only 9 independent weights (each entry of the kernel matrix corresponds to an independent weight), regardless of the number of nodes 520, 522 in the preceding node layer 510 and the following node layer 512.

[0072] In general, the convolutional neural network 500 uses node layers 510, 512, 514 with multiple channels, especially due to the use of multiple kernels in the convolution layer 511. In those cases, the node layer can be thought of as a (d+1)-dimensional matrix (the first dimension indexing the channel). The action of the convolution layer 511 is then a two-dimensional example, which is defined as follows

[0073]

[0074] in Corresponding to the ath channel of the front node layer 510, corresponds to the bth channel of the back node layer 512, and K a,b Corresponds to one of the kernels. If the convolutional layer 511 acts on a front node layer 510 with A channels and outputs a back node layer 512 with B channels, then there are A·B independent d-dimensional kernels K a,b .

[0075] Generally speaking, an activation function is used in the convolutional neural network 500. In this embodiment, re ReLU (the acronym for "rectified linear unit") is used, where R(z) = max(0,z), so that the action of the convolution layer 511 in the two-dimensional example is

[0076]

[0077] It is also possible to use other activation functions, such as ELU (acronym for "Exponential Linear Unit"), LeakyReLU, Sigmoid, Tanh or Softmax.

[0078] In the embodiment shown, the input layer 510 includes 36 nodes 520 arranged in a two-dimensional 6×6 matrix. The first hidden node layer 512 includes 72 nodes 522 arranged in two two-dimensional 6×6 matrices, each of which is the result of convolving the values of the input layer with a 3×3 kernel within the convolution layer 511. Equivalently, the nodes 522 of the first hidden node layer 512 can be interpreted as being arranged in a three-dimensional 2×6×6 matrix, where the first dimension corresponds to the channel dimension.

[0079] The advantage of using convolutional layers 511 is that the spatial local correlation of the input data can be exploited by enforcing local connectivity patterns between nodes in adjacent layers, in particular by having each node connected only to a small region of nodes in the previous layer.

[0080] The pooling layer 513 is a connection layer between the front node layer 512 (with node value x(n-1)) and the back node layer 514 (with node value x(n)). In particular, the pooling layer 513 can be characterized by the structure and weight of the edges and the activation function that form the pooling operation based on the nonlinear pooling function f. For example, in the two-dimensional case, the value x(n) of the node 524 of the back node layer 514 can be calculated based on the value x(n-1) of the node 522 of the front node layer 512 as follows

[0081]

[0082] In other words, by using the pooling layer 513, the number of nodes 522, 524 can be reduced by relocating d1·d2 number of adjacent nodes 522 in the preceding node layer 512, where the single node 522 in the subsequent node layer 514 is calculated as a function of the values of the adjacent nodes. In particular, the pooling function f can be a maximum function, an average value, or an L2 norm. In particular, for the pooling layer 513, the weights of the input edges are fixed and not modified by training.

[0083] The advantage of using the pooling layer 513 is that the number of nodes 522, 524 and the number of parameters are reduced. This results in a reduction in the amount of computation in the network and in control of overfitting.

[0084] In the embodiment shown, pooling layer 513 is a max pooling layer, which replaces four adjacent nodes with only one node whose value is the maximum of the values of the four adjacent nodes. Max pooling is applied to each d-dimensional matrix of the previous layer; in this embodiment, max pooling is applied to each of the two 2D matrices, thereby reducing the number of nodes from 72 to 18.

[0085] In general, the last few layers of the convolutional neural network 500 are fully connected layers 515. The fully connected layer 515 is a connection layer between the previous node layer 514 and the next node layer 516. The fully connected layer 513 may be characterized by the fact that most edges (particularly all edges) exist between the nodes 514 of the previous node layer 514 and the nodes 516 of the next node layer, and wherein the weight of each of these edges can be adjusted individually.

[0086] In this embodiment, nodes 524 of the preceding node layer 514 of the fully connected layer 515 are displayed both as a two-dimensional matrix and as non-correlated nodes (indicated as a row of nodes, where the number of nodes is reduced for better presentation). This operation is also referred to as "flattening." In this embodiment, the number of nodes 526 in the subsequent node layer 516 of the fully connected layer 515 is less than the number of nodes 524 in the preceding node layer 514. Alternatively, the number of nodes 526 may be equal or greater.

[0087] Furthermore, in this embodiment, a Softmax activation function is used within the fully connected layer 515. By applying the Softmax function, the sum of the values of all nodes 526 of the output layer 516 is 1, and all values of all nodes 526 of the output layer 516 are real numbers between 0 and 1. In particular, if the convolutional neural network 500 is used to classify input data, the value of the output layer 516 can be interpreted as the probability that the input data falls into one of the different categories.

[0088] In particular, the convolutional neural network 500 can be trained based on a back-propagation algorithm. To prevent overfitting, regularization methods can be used, such as dropping nodes 520, ..., 524, random pooling, using artificial data, weight decay based on L1 or L2 norm, or maximum norm constraint.

[0089] According to one aspect, the machine learning model may include one or more residual networks (ResNets). In particular, a ResNet is an artificial neural network comprising at least one skip or jump connection for skipping at least one layer of the artificial neural network. In particular, a ResNet may be a convolutional neural network comprising one or more jump connections that respectively skip one or more convolutional layers. According to some examples, a ResNet may be represented as an m-layer ResNet, where m is the number of layers in the corresponding architecture, and according to some examples, may take values of 34, 50, 101, or 152. According to some examples, such an m-layer ResNet may each include (m-2) / 2 jump connections.

[0090] A skip connection can be thought of as a bypass that feeds the output of the previous layer above one or more bypass layers directly to the subsequent layer of one or more bypass layers. The bypass layer will then have to fit a residual map that "balances" the directly fed output, rather than having to fit the desired map directly.

[0091] Fitting a residual mapping is computationally easier to optimize than a directed mapping. More importantly, this alleviates the problem of vanishing / exploding gradients during optimization when training machine learning models: if the bypass layer encounters such a problem, its contribution can be reduced by regularizing the output of the direct feed. Therefore, the advantage of using ResNet is that much deeper networks can be trained.

[0092] In particular, a recurrent machine learning model is a machine learning model whose output depends not only on the input values and the parameters of the machine learning model adapted through a training process, but also on a hidden state vector, where the hidden state vector is based on previous inputs to the recurrent machine learning model. In particular, the recurrent machine learning model may include additional memory states or additional structures that incorporate time delays or include feedback loops.

[0093] In particular, the underlying structure of the recurrent machine learning model can be a neural network, which can be denoted as a recurrent neural network. Such a recurrent neural network can be described as an artificial neural network in which the connections between nodes form a directed graph along a time series. In particular, the recurrent neural network can be interpreted as a directed acyclic graph. In particular, the recurrent neural network can be a finite-spur recurrent neural network or an infinite-spur recurrent neural network (wherein the finite-spur network can be expanded and replaced by a strictly feedforward neural network, and the infinite-spur network cannot be expanded and replaced by a strictly feedforward neural network).

[0094] In particular, training a recurrent neural network may be based on a BPTT algorithm (acronym for “Back Propagation Through Time”), an RTRL algorithm (acronym for “Real-Time Recurrent Learning”) and / or a genetic algorithm.

[0095] By using a recurrent machine learning model, input data consisting of sequences of variable length can be used. In particular, this means that the method cannot be used only for a fixed number of input datasets (and would need to be trained differently for each additional number of input datasets used as input), but can be used for any number of input datasets. This means that, independent of the number of input datasets contained in the different sequences, the entire training dataset can be used within the training, and the training data is not reduced to training data corresponding to a certain number of consecutive input datasets.

[0096] Figure 6The schematic structure of a recurrent machine learning model F is shown, both in a recurrent representation 602 and in an expanded representation 604, which can be used to implement one or more machine learning models described herein. The recurrent machine learning model takes several input data sets x, x1, ..., x N 606 as input and create a corresponding set of output data sets y, y1, ..., y N 608. Furthermore, the output depends on the so-called hidden vectors h, h1, ..., h N 610, which implicitly includes information about the input dataset previously used as input to the recurrent machine learning model F 612. By using these latent vectors h, h1, ..., h N 610, the sequential nature of the input data set can be exploited.

[0097] In a single step of processing, the recurrent machine learning model F 612 converts the hidden vector h created in the previous step n-1 and the input dataset x n As input. In this step, the recurrent machine learning model F generates an updated hidden vector h n and the output dataset y n As output. In other words, a processing step computes (y n ,h n )=F(x n ,h n-1 ), or by dividing the recurrent machine learning model F 612 into a part F(y) for calculating the output data and a part F(h) for calculating the hidden vector, one processing step calculates y n =F (y) (x n ,h n-1 ) and h n =F (h) (x n ,h n-1 ). For the first processing step, h0 can be randomly selected or filled with all zero entries. The parameters of the recurrent machine learning model F 612 previously trained based on the training dataset remain unchanged between different processing steps.

[0098] In particular, the output data and hidden vectors of a processing step depend on all previous input datasets used in previous steps. n =F (y) (x n ,F (h) (x n-1 ,h n-2 )) and h n =F (h) (x n ,F (h) (xn-1 ,h n-2 )).

[0099] The systems, devices, and methods described herein can be implemented using digital circuitry or using one or more computers utilizing known computer processors, memory units, storage devices, computer software, and other components. Typically, a computer includes a processor for executing instructions and one or more memories for storing instructions and data. A computer may also include or be coupled to one or more mass storage devices, such as one or more magnetic disks, internal hard disks and removable disks, magneto-optical disks, optical disks, and the like.

[0100] The systems, devices, and methods described herein can be implemented using computers operating in a client-server relationship. Typically, in such a system, the client computer is remotely located from the server computer and interacts via a network. The client-server relationship can be defined and controlled by computer programs running on the respective client and server computers.

[0101] The systems, devices, and methods described herein can be implemented within a network-based cloud computing system. In such a network-based cloud computing system, a server or another processor connected to the network communicates with one or more client computers via the network. For example, a client computer can communicate with a server via a web browser application that resides and operates on the client computer. The client computer can store data on the server and access the data via the network. The client computer can transmit a request for data or a request for an online service to the server via the network. The server can perform the requested service and provide the data to the client computer(s). The server can also transmit data adapted to cause the client computer to perform a specified function (e.g., perform a calculation, display specified data on a screen, etc.). For example, the server can transmit data adapted to cause the client computer to perform one or more of the steps or functions of the methods and workflows described herein (including Figure 1-Figure 3 Certain steps or functions of the methods and workflows described herein (including Figure 1-Figure 3 One or more of the steps or functions of the method and workflow described herein may be performed by a server or another processor in a network-based cloud computing system. Figure 1-Figure 3 One or more of the steps of the method and workflow described herein may be performed by a client computer in a network-based cloud computing system. Figure 1-Figure 3 One or more of the steps of ) may be performed by a server and / or a client computer in a network-based cloud computing system in any combination.

[0102] The systems, apparatus, and methods described herein may be implemented using a computer program product tangibly embodied in an information carrier (e.g., tangibly embodied in a non-transitory machine-readable storage device and executed by a programmable processor); and the methods and workflow steps described herein (including Figure 1-Figure 3 One or more of the steps or functions of a program may be implemented using one or more computer programs executable by such a processor. A computer program is a set of computer program instructions that can be used, directly or indirectly, in a computer to perform a specific activity or produce a specific result. A computer program may be written in any form of programming language (including compiled or interpreted languages), and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0103] A high-level block diagram of an example computer 702 that can be used to implement the systems, apparatus, and methods described herein is provided in Figure 7 . The computer 702 includes a processor 704 operatively coupled to a data storage device 712 and a memory 710. The processor 704 controls the overall operation of the computer 702 by executing computer program instructions that define such operation. The computer program instructions may be stored in the data storage device 712 or other computer readable medium and loaded into the memory 710 when execution of the computer program instructions is desired. Thus, Figure 1-Figure 3 The methods and workflow steps or functions may be defined by computer program instructions stored in the memory 710 and / or the data storage device 712 and controlled by the processor 704 executing the computer program instructions. For example, the computer program instructions may be implemented as computer executable code programmed by a person skilled in the art to perform Figure 1-Figure 3 Thus, by executing the computer program instructions, the processor 704 performs Figure 1-Figure 3 The computer 702 may also include one or more network interfaces 706 for communicating with other devices via a network. The computer 702 may also include one or more input / output devices 708 that enable a user to interact with the computer 702 (e.g., a display, keyboard, mouse, speakers, buttons, etc.).

[0104] The processor 704 may include both general-purpose and special-purpose microprocessors and may be the sole processor or one of multiple processors of the computer 702. For example, the processor 704 may include one or more central processing units (CPUs). The processor 704, the data storage device 712, and / or the memory 710 may include, be supplemented by, or incorporate one or more application-specific integrated circuits (ASICs) and / or one or more field-programmable gate arrays (FPGAs).

[0105] The data storage device 712 and the memory 710 each include a tangible, non-transitory computer-readable storage medium. The data storage device 712 and the memory 710 may each include a high-speed random access memory, such as a dynamic random access memory (DRAM), a static random access memory (SRAM), a double data rate synchronous dynamic random access memory (DDR RAM), or other random access solid-state storage devices, and may include non-volatile memory, such as one or more magnetic disk storage devices, such as internal hard disks and removable disks, magneto-optical disk storage devices, optical disk storage devices, flash memory devices, semiconductor memory devices, such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), compact disk read-only memory (CD-ROM), digital versatile disk read-only memory (DVD-ROM) disks, or other non-volatile solid-state storage devices.

[0106] Input / output devices 708 may include peripheral devices such as printers, scanners, display screens, etc. For example, input / output devices 708 may include a display device such as a cathode ray tube (CRT) or liquid crystal display (LCD) monitor for displaying information to a user, a keyboard, and a pointing device such as a mouse or trackball through which a user can provide input to the computer 702.

[0107] Image acquisition device 714 can be connected to computer 702 to input image data (e.g., medical images) into computer 702. It is possible that image acquisition device 714 and computer 702 are implemented as a single device. Image acquisition device 714 and computer 702 can also communicate wirelessly via a network. In a possible embodiment, computer 702 can be remotely located relative to image acquisition device 714.

[0108] One or more computers, such as computer 702 , may be used to implement any or all of the systems, apparatus, and methods discussed herein.

[0109] Those skilled in the art will recognize that actual computer or computer system implementations may have other structures and may include other components, and Figure 7is a high-level representation of some components of such a computer for illustrative purposes.

[0110] Independent of the grammatical use of the term, individuals with both male and female identities are included within the term.

[0111] The foregoing detailed description should be understood to be illustrative and exemplary in every respect, rather than restrictive, and the scope of the invention disclosed herein should not be determined from the detailed description, but rather from the claims interpreted in accordance with the full breadth permitted by the patent laws. It should be understood that the embodiments shown and described herein are merely illustrative of the principles of the invention, and that various modifications may be implemented by those skilled in the art without departing from the scope and spirit of the invention. Various other feature combinations may be implemented by those skilled in the art without departing from the scope and spirit of the invention.

[0112] The following is a list of non-limiting illustrative embodiments disclosed herein:

[0113] Illustrative embodiments 1. A computer-implemented method comprising: receiving one or more input medical images; extracting image embeddings from the one or more input medical images; using a trained visual language model to perform one or more medical imaging analysis tasks based on the image embeddings extracted from the one or more input medical images; and outputting results of the one or more medical imaging analysis tasks, wherein the trained visual language model is trained by: receiving one or more training medical images and text-based reports associated with the one or more training medical images, extracting image embeddings from the one or more training medical images, generating one or more instructions based on the text-based reports using a language model, and training the visual language model to perform the one or more medical imaging analysis tasks based on the image embeddings extracted from the one or more training medical images and the one or more generated instructions.

[0114] Illustrative embodiment 2. A computer-implemented method according to illustrative embodiment 1, wherein generating one or more instructions based on a text-based report using a language model includes: further generating the one or more instructions based on multiple predefined templates, the multiple predefined templates including different initial instructions for extracting information from the text-based report and generating the one or more instructions.

[0115] Illustrative embodiment 3. A computer-implemented method according to any one of illustrative embodiments 1-2, wherein generating one or more instructions based on a text-based report using a language model includes: generating the one or more instructions for associating anatomical features depicted in the one or more training medical images with textual anatomical descriptors.

[0116] Illustrative embodiment 4. A computer-implemented method according to any one of illustrative embodiments 1-3, wherein generating one or more instructions based on a text-based report using a language model includes: generating the one or more instructions for associating anatomical features depicted in the one or more training medical images with each other.

[0117] Illustrative embodiment 5. A computer-implemented method according to any one of illustrative embodiments 1-4, wherein generating one or more instructions based on a text-based report using a language model includes: generating the one or more instructions for associating textual anatomical descriptors with image findings of the one or more training medical images.

[0118] Illustrative embodiment 6. The computer-implemented method of illustrative embodiment 5, wherein the image findings comprise quantitative image findings of the one or more input medical images.

[0119] Illustrative embodiment 7. A computer-implemented method according to any one of illustrative embodiments 1-6, wherein: generating one or more instructions based on a text-based report using a first language model includes: generating instruction embeddings representing the one or more instructions; and training a visual language model based on image embeddings extracted from the one or more training medical images and the one or more generated instructions to perform the one or more medical imaging analysis tasks includes: combining the image embeddings and instruction embeddings extracted from the one or more training medical images, and generating results of the one or more medical imaging analysis tasks based on the combined image embeddings and instruction embeddings.

[0120] Illustrative embodiment 8. An apparatus comprising: a component for receiving one or more input medical images; a component for extracting image embeddings from the one or more input medical images; a component for performing one or more medical imaging analysis tasks based on the image embeddings extracted from the one or more input medical images using a trained visual language model; and a component for outputting results of the one or more medical imaging analysis tasks, wherein the trained visual language model is trained by: receiving one or more training medical images and text-based reports associated with the one or more training medical images, extracting image embeddings from the one or more training medical images, generating one or more instructions based on the text-based reports using a language model, and training the visual language model to perform the one or more medical imaging analysis tasks based on the image embeddings extracted from the one or more training medical images and the one or more generated instructions.

[0121] Illustrative embodiment 9. An apparatus according to illustrative embodiment 8, wherein generating one or more instructions based on a text-based report using a language model includes: further generating the one or more instructions based on multiple predefined templates, the multiple predefined templates including different initial instructions for extracting information from the text-based report and generating the one or more instructions.

[0122] Illustrative embodiment 10. An apparatus according to any one of illustrative embodiments 8-9, wherein generating one or more instructions based on a text-based report using a language model includes: generating the one or more instructions for associating anatomical features depicted in the one or more training medical images with textual anatomical descriptors.

[0123] Illustrative embodiment 11. An apparatus according to any one of illustrative embodiments 8-10, wherein the component for generating one or more instructions based on a text-based report using a language model includes: generating the one or more instructions for associating anatomical features depicted in the one or more training medical images with each other.

[0124] Illustrative embodiment 12. A non-transitory computer-readable storage medium comprising instructions that, when executed by a computer, cause the computer to perform operations comprising: receiving one or more input medical images; extracting image embeddings from the one or more input medical images; using a trained visual language model to perform one or more medical imaging analysis tasks based on the image embeddings extracted from the one or more input medical images; and outputting results of the one or more medical imaging analysis tasks, wherein the trained visual language model is trained by: receiving one or more training medical images and text-based reports associated with the one or more training medical images, extracting image embeddings from the one or more training medical images, generating one or more instructions based on the text-based reports using a language model, and training the visual language model to perform the one or more medical imaging analysis tasks based on the image embeddings extracted from the one or more training medical images and the one or more generated instructions.

[0125] Illustrative embodiment 13. A non-transitory computer-readable storage medium according to illustrative embodiment 12, wherein generating one or more instructions based on a text-based report using a language model includes: further generating the one or more instructions based on multiple predefined templates, the multiple predefined templates including different initial instructions for extracting information from the text-based report and generating the one or more instructions.

[0126] Illustrative embodiment 14. A non-transitory computer-readable storage medium according to any one of illustrative embodiments 12-13, wherein generating one or more instructions based on a text-based report using a language model includes: generating the one or more instructions for associating textual anatomical descriptors with image findings of the one or more training medical images.

[0127] Illustrative embodiment 15. The non-transitory computer-readable storage medium of illustrative embodiments 12-14, wherein the image findings include quantitative image findings of the one or more input medical images.

[0128] Illustrative embodiment 16. A non-transitory computer-readable storage medium according to any one of illustrative embodiments 12-15, wherein: generating one or more instructions based on a text-based report using a first language model includes: generating instruction embeddings representing the one or more instructions; and training a visual language model based on image embeddings extracted from the one or more training medical images and the one or more generated instructions to perform the one or more medical imaging analysis tasks includes: combining the image embeddings and instruction embeddings extracted from the one or more training medical images, and generating results of the one or more medical imaging analysis tasks based on the combined image embeddings and instruction embeddings.

[0129] Illustrative embodiment 17. A computer-implemented method comprising: receiving one or more training medical images and text-based reports associated with the one or more training medical images; extracting image embeddings from the one or more training medical images; generating one or more instructions based on the text-based reports using a language model; and training a visual language model to perform one or more medical imaging analysis tasks based on the image embeddings extracted from the one or more training medical images and the one or more generated instructions.

[0130] Illustrative embodiment 18. A computer-implemented method according to illustrative embodiment 17, wherein generating one or more instructions based on a text-based report using a language model includes: further generating the one or more instructions based on multiple predefined templates, the multiple predefined templates including different initial instructions for extracting information from the text-based report and generating the one or more instructions.

[0131] Illustrative embodiment 19. A computer-implemented method according to any one of illustrative embodiments 17-18, wherein generating one or more instructions based on a text-based report using a language model includes: generating the one or more instructions for associating anatomical features depicted in the one or more training medical images with textual anatomical descriptors.

[0132] Illustrative embodiment 20. A computer-implemented method according to any one of illustrative embodiments 17-19, wherein generating one or more instructions based on a text-based report using a language model includes: generating the one or more instructions for associating anatomical features depicted in the one or more training medical images with each other.

Claims

1. A computer-implemented method comprising: receiving one or more input medical images; extracting an image embedding from the one or more input medical images; performing one or more medical imaging analysis tasks based on image embeddings extracted from the one or more input medical images using the trained visual-language model; and outputting the results of the one or more medical imaging analysis tasks, The trained visual language model is trained by: receiving one or more training medical images and a text-based report associated with the one or more training medical images, extracting image embeddings from the one or more training medical images, generating one or more instructions based on the text-based report using the language model, and A visual language model is trained based on the image embeddings extracted from the one or more training medical images and the one or more generated instructions to perform the one or more medical imaging analysis tasks.

2. The computer-implemented method of claim 1 , wherein generating one or more instructions based on the text-based report using a language model comprises: One or more instructions are further generated based on a plurality of predefined templates, the plurality of predefined templates including different initial instructions for extracting information from the text-based report and generating the one or more instructions.

3. The computer-implemented method of claim 1 , wherein generating one or more instructions based on a text-based report using a language model comprises: The one or more instructions are generated for associating anatomical features depicted in the one or more training medical images with textual anatomical descriptors.

4. The computer-implemented method of claim 1 , wherein generating one or more instructions based on a text-based report using a language model comprises: The one or more instructions are generated for associating anatomical features depicted in the one or more training medical images with each other.

5. The computer-implemented method of claim 1 , wherein generating one or more instructions based on a text-based report using a language model comprises: The one or more instructions for associating textual anatomical descriptors with image findings of the one or more training medical images are generated. The computer-implemented method of claim 5 , wherein the image findings comprise quantitative image findings of the one or more input medical images.

7. The computer-implemented method of claim 1 , wherein: Generating one or more instructions based on a text-based report using a language model includes: generating an instruction embedding representing the one or more instructions; and Training a visual language model based on image embeddings extracted from the one or more training medical images and the one or more generated instructions to perform the one or more medical imaging analysis tasks comprises: combining image embeddings and instruction embeddings extracted from the one or more training medical images, and Results of the one or more medical imaging analysis tasks are generated based on the combined image embedding and instruction embedding.

8. A device comprising: means for receiving one or more input medical images; means for extracting an image embedding from the one or more input medical images; means for performing one or more medical imaging analysis tasks based on image embeddings extracted from the one or more input medical images using the trained visual-language model; and means for outputting results of said one or more medical imaging analysis tasks, The trained visual language model is trained by: receiving one or more training medical images and a text-based report associated with the one or more training medical images, extracting image embeddings from the one or more training medical images, generating one or more instructions based on the text-based report using the language model, and A visual language model is trained based on the image embeddings extracted from the one or more training medical images and the one or more generated instructions to perform the one or more medical imaging analysis tasks.

9. The apparatus of claim 8, wherein generating one or more instructions based on the text-based report using the language model comprises: One or more instructions are further generated based on a plurality of predefined templates, the plurality of predefined templates including different initial instructions for extracting information from the text-based report and generating the one or more instructions.

10. The apparatus of claim 8, wherein generating one or more instructions based on the text-based report using the language model comprises: The one or more instructions are generated for associating anatomical features depicted in the one or more input medical images with textual anatomical descriptors.

11. The apparatus of claim 8, wherein generating one or more instructions based on a text-based report using a language model comprises: The one or more instructions are generated for associating anatomical features depicted in the one or more input medical images with one another.

12. A non-transitory computer-readable storage medium comprising instructions that, when executed by a computer, cause the computer to perform operations comprising: receiving one or more input medical images; extracting an image embedding from the one or more input medical images; performing one or more medical imaging analysis tasks based on image embeddings extracted from the one or more input medical images using the trained visual-language model; and outputting the results of the one or more medical imaging analysis tasks, The trained visual language model is trained by: receiving one or more training medical images and a text-based report associated with the one or more training medical images, extracting image embeddings from the one or more training medical images, generating one or more instructions based on the text-based report using the language model, and A visual language model is trained based on the image embeddings extracted from the one or more training medical images and the one or more generated instructions to perform the one or more medical imaging analysis tasks.

13. The non-transitory computer-readable storage medium of claim 12, wherein generating one or more instructions based on a text-based report using a language model comprises: One or more instructions are further generated based on a plurality of predefined templates, the plurality of predefined templates including different initial instructions for extracting information from the text-based report and generating the one or more instructions.

14. The non-transitory computer-readable storage medium of claim 12, wherein generating one or more instructions based on a text-based report using a language model comprises: The one or more instructions for associating textual anatomical descriptors with image findings of the one or more training medical images are generated. 15 . The non-transitory computer-readable storage medium of claim 14 , wherein the image findings comprise quantitative image findings of the one or more input medical images.

16. The non-transitory computer-readable storage medium of claim 12, wherein: Generating one or more instructions based on a text-based report using a language model includes: generating an instruction embedding representing the one or more instructions; and Training a visual language model based on image embeddings extracted from the one or more training medical images and the one or more generated instructions to perform the one or more medical imaging analysis tasks comprises: combining image embeddings and instruction embeddings extracted from the one or more training medical images, and Results of the one or more medical imaging analysis tasks are generated based on the combined image embedding and instruction embedding.

17. A computer-implemented method comprising: receiving one or more training medical images and a text-based report associated with the one or more training medical images; extracting image embeddings from the one or more training medical images; generating one or more instructions based on the text-based report using the language model; and A visual language model is trained based on the image embeddings extracted from the one or more training medical images and the one or more generated instructions to perform one or more medical imaging analysis tasks.

18. The computer-implemented method of claim 17, wherein generating one or more instructions based on the text-based report using the language model comprises: One or more instructions are further generated based on a plurality of predefined templates, the plurality of predefined templates including different initial instructions for extracting information from the text-based report and generating the one or more instructions.

19. The computer-implemented method of claim 17, wherein generating one or more instructions based on a text-based report using a language model comprises: The one or more instructions are generated for associating anatomical features depicted in the one or more training medical images with textual anatomical descriptors.

20. The computer-implemented method of claim 17, wherein generating one or more instructions based on a text-based report using a language model comprises: The one or more instructions are generated for associating anatomical features depicted in the one or more training medical images with each other.