Training Method, Device, Medium and Equipment for Multimodal Question-Answering Model of Medical Images

By integrating CLIP model, SAM, visual language connector and LLM in the question and answer model, and using a large amount of medical data to train the model, the problem of insufficient visual perception and lack of medical knowledge in the medical field is solved, and the professional ability of the model and the accuracy of the answer is improved.

CN119170254BActive Publication Date: 2025-06-27BEIJING UCAP INTERNET TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411206273.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2025-06-27
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

The application of question-and-answer models in the medical field has problems such as insufficient visual perception and lack of medical knowledge, which leads to low accuracy of answers.

Method used

By creating a multimodal question and answer model, including a contrasting language-image pre-trained CLIP model, segmentation model SAM, visual language connector and large language model LLM, use a large number of medical images and text training models to inject medical knowledge and provide fine-grained visual perception through pixel-level SAM.

Benefits of technology

The multimodal question and answer model's professional ability in understanding and processing medical images and text has been improved, and its ability to process fine and complex medical images has been enhanced, and the accuracy and precision of the answers have been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119170254B_ABST
    Figure CN119170254B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, medium, and device for training a multi-modal question answering model for medical images, belonging to the technical field of deep learning. Create a multi-modal question answering model; obtain the first and second training sets; fix the weights of the CLIP model, SAM, and LLM, and use the first training samples to train the weights of the first visual-language connector and the second language-visual connector; fix the weights of the CLIP model and SAM, and use the first training samples to train the weights of the first visual-language connector, the second language-visual connector, and the LLM; fix the weights of the CLIP model and SAM, and use the second training samples to train the weights of the first visual-language connector, the second language-visual connector, and the LLM; load the trained weights into the question answering model to obtain a trained multi-modal question answering model. The multi-modal question answering model can process fine and complex medical images, improving the accuracy and fineness of the answers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning technology, and particularly relates to a method, device, medium, and equipment for training a multi-modal question-answering model for medical images. Background Art

[0002] The input of the question-answering model is a medical image and a question about the medical image, and the output is an answer to the question based on the medical image. The question-answering model can help patients obtain feedback on their conditions and can also help doctors generate diagnostic opinions.

[0003] The question-answering model in the related technology is a multi-modal model, which includes a Contrastive Language-Image Pre-Training (CLIP) model, a visual-language connector, and a Large Language Model (LLM). The CLIP model is responsible for converting the medical image into a feature vector, the visual-language connector is responsible for performing visual-language space conversion on the feature vector, and the LLM is responsible for generating an answer based on the converted feature vector and the text features of the question.

[0004] Although the question-answering model can achieve multi-modal combination of visual and language understanding, its application in the medical field has the disadvantages of insufficient visual perception and lack of medical knowledge, resulting in low accuracy of the answers generated by the question-answering model. Summary of the Invention

[0005] This application provides a method, device, medium, and equipment for training a multi-modal question-answering model for medical images, which is used to solve the problems of insufficient visual perception and lack of medical knowledge in the application of the question-answering model in the medical field. The technical solutions are as follows:

[0006] According to the first aspect of this application, a method for training a multi-modal question-answering model for medical images is provided. The method includes:

[0007] Create a multi-modal question-answering model, which includes a Contrastive Language-Image Pre-Training (CLIP) model, a Segment Anything Model (SAM), a first visual-language connector, a second visual-language connector, and a Large Language Model (LLM). The CLIP model is connected to the LLM through the first visual-language connector, and the SAM is connected to the LLM through the second visual-language connector;

[0008] Obtain a first training set and a second training set. The first training samples in the first training set include medical images, image annotations, simple Q&A texts with image annotations, and simple Q&A texts without image annotations. The second training samples in the second training set include medical images, image annotations, detailed Q&A texts with image annotations, and detailed Q&A texts without image annotations;

[0009] Fix the weights of the CLIP model, the SAM, and the LLM, and use the first training samples to train the weights of the first visual-language connector and the second language-visual connector;

[0010] Fix the weights of the CLIP model and the SAM, and use the first training samples to train the weights of the first visual-language connector, the second language-visual connector, and the LLM;

[0011] Fix the weights of the CLIP model and the SAM, and use the second training samples to train the weights of the first visual-language connector, the second language-visual connector, and the LLM;

[0012] Load the trained weights into the multi-modal Q&A model to obtain a trained multi-modal Q&A model.

[0013] In one possible implementation, the method further includes:

[0014] Use the multi-modal Q&A model to process test samples to obtain answers output by the multi-modal Q&A model;

[0015] Obtain feedback information from the user on the answers;

[0016] Use the Direct Preference Optimization (DPO) algorithm and the feedback information to optimize the multi-modal Q&A model.

[0017] In one possible implementation, the step of fixing the weights of the CLIP model, the SAM, and the LLM, and using the first training samples to train the weights of the first visual-language connector and the second language-visual connector includes:

[0018] Use the CLIP model with fixed weights to extract features from the medical images in the first training samples to obtain first image features;

[0019] Use the first visual-language connector to encode the first image features to obtain first image encodings;

[0020] When training using the simple Q&A text without image annotations in the first training sample, use the SAM with fixed weights to extract features from the medical image to obtain a second image feature; use the second vision-language connector to encode the second image feature to obtain a second image encoding; use tokenization and an embedding layer to randomly mask and then encode the simple Q&A text without image annotations to obtain a first text encoding;

[0021] When training using the image annotations and the simple Q&A text with image annotations in the first training sample, use the SAM with fixed weights to extract features from the medical image and the image annotations to obtain a second image feature; use the second vision-language connector to encode the second image feature to obtain a second image encoding; use tokenization and an embedding layer to randomly mask and then encode the simple Q&A text with image annotations to obtain a first text encoding;

[0022] Use the LLM with fixed weights to process the first image encoding, the second image encoding, and the first text encoding to obtain the predicted text of the simple Q&A text;

[0023] Update the weights of the first vision-language connector and the second vision-language connector according to the simple Q&A text and the predicted text.

[0024] In a possible implementation, fixing the weights of the CLIP model and the SAM and training the weights of the first vision-language connector, the second vision-language connector, and the LLM using the first training sample includes:

[0025] Use the CLIP model with fixed weights to extract features from the medical image in the first training sample to obtain a third image feature;

[0026] Use the first vision-language connector to encode the third image feature to obtain a third image encoding;

[0027] When training using the simple Q&A text without image annotations in the first training sample, use the SAM with fixed weights to extract features from the medical image to obtain a fourth image feature; use the second vision-language connector to encode the fourth image feature to obtain a fourth image encoding; use tokenization and an embedding layer to randomly mask and then encode the simple Q&A text without image annotations to obtain a second text encoding;

[0028] When training using the image annotations in the first training sample and the simple Q&A text with image annotations, use the SAM with fixed weights to extract features from the medical image and the image annotations to obtain the fourth image feature; use the second vision-language connector to encode the fourth image feature to obtain the fourth image encoding; use tokenization and the embedding layer to randomly mask and then encode the simple Q&A text with image annotations to obtain the second text encoding;

[0029] Use the LLM to process the third image encoding, the fourth image encoding, and the second text encoding to obtain the predicted text of the simple Q&A text;

[0030] Update the weights of the first vision-language connector, the second vision-language connector, and the LLM according to the simple Q&A text and the predicted text.

[0031] In a possible implementation, fixing the weights of the CLIP model and the SAM, and training the weights of the first vision-language connector, the second vision-language connector, and the LLM using the second training sample includes:

[0032] Use the CLIP model with fixed weights to extract features from the medical images in the second training sample to obtain the fifth image feature;

[0033] Use the first vision-language connector to encode the fifth image feature to obtain the fifth image encoding;

[0034] When training using the detailed Q&A text without image annotations in the second training sample, use the SAM with fixed weights to extract features from the medical image to obtain the sixth image feature; use the second vision-language connector to encode the sixth image feature to obtain the sixth image encoding; use tokenization and the embedding layer to mask and then encode the answer in the detailed Q&A text without image annotations to obtain the third text encoding;

[0035] When training using the image annotations in the second training sample and the detailed Q&A text with image annotations, use the SAM with fixed weights to extract features from the medical image and the image annotations to obtain the sixth image feature; use the second vision-language connector to encode the sixth image feature to obtain the sixth image encoding; use tokenization and the embedding layer to mask and then encode the answer in the detailed Q&A text with image annotations to obtain the third text encoding;

[0036] Use the LLM to process the fifth image encoding, the sixth image encoding, and the third text encoding to obtain the predicted text of the answer;

[0037] Update the weights of the first vision-language connector, the second language-vision connector, and the LLM according to the said answer and the said predicted text.

[0038] In one possible implementation, creating the multi-modal question-answering model includes:

[0039] Obtain a multi-modal model, and pre-train the multi-modal model using a third training set, where the multi-modal model includes a CLIP model, a first vision-language connector, and an LLM;

[0040] Obtain SAM, and fine-tune SAM using a fourth training set;

[0041] Create the multi-modal question-answering model according to the pre-trained multi-modal model and the fine-tuned SAM model.

[0042] In one possible implementation, the method further includes:

[0043] Use the CLIP model to extract features from the medical image to be processed, obtaining a seventh image feature;

[0044] Use the first vision-language connector to encode the seventh image feature, obtaining a seventh image encoding;

[0045] When the medical image has no image annotation, use SAM to extract features from the medical image, obtaining an eighth image feature; use the second vision-language connector to encode the eighth image feature, obtaining an eighth image encoding; use a tokenizer and an embedding layer to encode the question to be processed, obtaining a fourth text encoding;

[0046] When the medical image has an image annotation, use SAM to extract features from the medical image and the image annotation, obtaining an eighth image feature; use the second vision-language connector to encode the eighth image feature, obtaining an eighth image encoding; use a tokenizer and an embedding layer to encode the question to be processed, obtaining a fourth text encoding;

[0047] Use the LLM to process the seventh image encoding, the eighth image encoding, and the fourth text encoding, obtaining the answer output by the multi-modal question-answering model.

[0048] According to a second aspect of the present application, there is provided a training device for a multi-modal question-answering model of medical images, the device includes:

[0049] A creation module for creating a multimodal question-answering model, the multimodal question-answering model including a contrastive language-image pre-training CLIP model, a segmentation model SAM, a first vision-language connector, a second vision-language connector, and a large language model LLM. The CLIP model is connected to the LLM through the first vision-language connector, and the SAM is connected to the LLM through the second vision-language connector;

[0050] An acquisition module for acquiring a first training set and a second training set. The first training samples in the first training set include medical images, image annotations, simple question-and-answer texts with image annotations, and simple question-and-answer texts without image annotations. The second training samples in the second training set include medical images, image annotations, detailed question-and-answer texts with image annotations, and detailed question-and-answer texts without image annotations;

[0051] A training module for fixing the weights of the CLIP model, the SAM, and the LLM, and training the weights of the first vision-language connector and the second vision-language connector using the first training samples;

[0052] The training module is also used for fixing the weights of the CLIP model and the SAM, and training the weights of the first vision-language connector, the second vision-language connector, and the LLM using the first training samples;

[0053] The training module is also used for fixing the weights of the CLIP model and the SAM, and training the weights of the first vision-language connector, the second vision-language connector, and the LLM using the second training samples;

[0054] A loading module for loading the trained weights into the multimodal question-answering model to obtain a trained multimodal question-answering model.

[0055] According to the third aspect of the present application, there is provided a computer-readable storage medium storing at least one instruction, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned training method for a multimodal question-answering model of medical images.

[0056] According to the fourth aspect of the present application, there is provided a computer device including the above-mentioned training device for a multimodal question-answering model of medical images.

[0057] The beneficial effects of the technical solution provided by the present application at least include:

[0058] Training the multimodal question-answering model with a large number of first training sets and second training sets can inject medical knowledge into the multimodal question-answering model and improve the professional ability of the multimodal question-answering model to understand and process medical images and texts.

[0059] Processing medical images through pixel - level SAM can provide fine - grained and dense visual perception capabilities for multi - modal question - answering models, improving their ability to process fine and complex medical images.

[0060] By adding image annotations of regions of interest in medical images, the attention location of the questions consulted by users can be narrowed down from the whole image to the regions of interest, improving the accuracy and fineness of the answers.

[0061] Through the DPO algorithm, feedback information on the answers output by the multi - modal question - answering model can be collected, and value alignment, preference alignment, and professional alignment can be performed according to this feedback information, improving the safety, accuracy, and professionalism of the answers output by the multi - modal question - answering model. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0063] Figure 1 It is a schematic structural diagram of a multi - modal question - answering model provided by an embodiment of the present application;

[0064] Figure 2 It is a flowchart of a method for training a multi - modal question - answering model for medical images provided by an embodiment of the present application;

[0065] Figure 3 It is a flowchart of a method for training a multi - modal question - answering model for medical images provided by an embodiment of the present application;

[0066] Figure 4 It is a flowchart of a method for using a multi - modal question - answering model for medical images provided by an embodiment of the present application;

[0067] Figure 5 It is a block diagram of the structure of a device for training a multi - modal question - answering model for medical images provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the drawings.

[0069] As Figure 1As shown in the figure, the multimodal question-answering model includes a CLIP model, a SAM, a first visual-language connector (VL-Connector1), a second visual-language connector (VL-Connector2), and an LLM. The CLIP model is connected to the LLM through the first visual-language connector (VL-Connector1), and the SAM is connected to the LLM through the second visual-language connector (VL-Connector2). The multimodal question-answering model also includes a tokenizer and an embedding layer (tokenizer&embedding), and the tokenizer and embedding layer are connected to the LLM. Among them, the CLIP model and the SAM are used to process medical images and output the image encodings to the LLM through their respective connectors. The tokenizer and embedding layer are used to encode the questions and output the text encodings to the LLM. The LLM processes the image encodings and text encodings and outputs answers.

[0070] As Figure 2 shown in the figure, it shows a flowchart of a method for training a multimodal question-answering model for medical images provided by an embodiment of the present application. The method for training a multimodal question-answering model for medical images can be applied to a computer device. The method for training a multimodal question-answering model for medical images may include:

[0071] Step 201, create a multimodal question-answering model.

[0072] Among them, the structure of the multimodal question-answering model is as Figure 1 shown in the figure.

[0073] The multimodal question-answering model is a model formed by adding a visual-language connector and a SAM to a multimodal model. We can pre-train the multimodal model and fine-tune the SAM so that both the multimodal model and the SAM have their respective initial weights.

[0074] Step 202, obtain a first training set and a second training set. The first training samples in the first training set include medical images, image annotations, simple question-and-answer texts with image annotations, and simple question-and-answer texts without image annotations. The second training samples in the second training set include medical images, image annotations, detailed question-and-answer texts with image annotations, and detailed question-and-answer texts without image annotations.

[0075] Medical images are images output after examining body parts using medical devices, including but not limited to X-ray films, B-ultrasound images, computed tomography (CT) scan images, and nuclear magnetic resonance imaging (MRI) images.

[0076] Image annotations are used to represent the regions of interest selected by the user in medical images. Image annotations can be points, boxes, contours, masks, etc.

[0077] According to the complexity and refinement of the Q&A text, we divide the Q&A text into two categories: simple Q&A text and detailed Q&A text. The classification rules can be set according to content, word count, etc., which are not limited here. Among them, simple Q&A text is used in the pre-training stage to enable the multi-modal Q&A model to simply describe the general and rough content of medical images; detailed Q&A text is used in the fine-tuning stage to enable the multi-modal Q&A model to more precisely describe the content of medical images, or to generate answers with reasoning properties.

[0078] In this embodiment, the first training set and the second training set are used to supplement professional medical knowledge for the multi-modal Q&A model.

[0079] Step 203, fix the weights of the CLIP model, SAM, and LLM, and use the first training samples to train the weights of the first visual-language connector and the second language-visual connector.

[0080] At the beginning of training, the medical images and simple Q&A text without image annotations in the first training samples can be used for training; after the first visual-language connector and the second language-visual connector gradually stabilize, the medical images, image annotations, and simple Q&A text with image annotations in the first training samples can be used for training with a certain probability, so that the multi-modal Q&A model supports both Q&A without image annotations and Q&A with image annotations.

[0081] Step 204, fix the weights of the CLIP model and SAM, and use the first training samples to train the weights of the first visual-language connector, the second language-visual connector, and the LLM.

[0082] Similarly, at the beginning of training, the medical images and simple Q&A text without image annotations in the first training samples can be used for training; subsequently, the medical images, image annotations, and simple Q&A text with image annotations in the first training samples can be used for training with a certain probability.

[0083] Step 205, fix the weights of the CLIP model and SAM, and use the second training samples to train the weights of the first visual-language connector, the second language-visual connector, and the LLM.

[0084] Similarly, at the beginning of training, the medical images and detailed Q&A text without image annotations in the second training samples can be used for training; subsequently, the medical images, image annotations, and detailed Q&A text with image annotations in the second training samples can be used for training with a certain probability.

[0085] Step 206: Load the trained weights into the multi-modal question answering model to obtain the trained multi-modal question answering model.

[0086] In summary, the method for training a multi-modal question answering model for medical images provided by the embodiments of the present application trains the multi-modal question answering model through a large number of first training sets and second training sets, which can inject medical knowledge into the multi-modal question answering model and improve the professional ability of the multi-modal question answering model to understand and process medical images and texts.

[0087] By processing medical images through pixel-level SAM, it can provide the multi-modal question answering model with fine-grained and dense visual perception capabilities, and improve its ability to process fine and complex medical images.

[0088] By adding image annotations of regions of interest in medical images, the attention localization of the questions consulted by users can be narrowed down from the entire image to the regions of interest, improving the accuracy and fineness of the answers.

[0089] As Figure 3 shown, it shows a flowchart of the method for training a multi-modal question answering model for medical images provided by an embodiment of the present application. The method for training a multi-modal question answering model for medical images can be applied to a computer device. The method for training a multi-modal question answering model for medical images can include:

[0090] Step 301: Create a multi-modal question answering model.

[0091] Specifically, creating a multi-modal question answering model can include:

[0092] (1) Obtain a multi-modal model and pre-train the multi-modal model using a third training set. The multi-modal model includes a CLIP model, a first vision-language connector, and an LLM.

[0093] The multi-modal model includes a CLIP model, a first vision-language connector, and an LLM.

[0094] The training samples in the third training set include medical images and the question-and-answer texts for at least one round of question-and-answer for the medical images.

[0095] After pre-training the multi-modal model based on the image-text pairs in the third training set, the multi-modal model can learn the correlation between images and texts, and can effectively extract and fuse the information of these two modalities, enabling the multi-modal model to have strong cross-modal understanding capabilities. It can perform well in a variety of tasks, including but not limited to image caption generation, visual question answering, multi-modal dialogue, etc.

[0096] To enable multi-modal models to better serve the medical industry, it is necessary to further fine-tune them using specialized medical data. This fine-tuning can not only enhance the multi-modal model's ability to understand medical images but also help the multi-modal model master more professional medical knowledge. Specifically, the fine-tuned multi-modal model can: improve the understanding ability of professional vocabulary, better understand and generate medical terms; improve the accuracy of disease recognition, more accurately identify and classify diseases; enhance medical decision support, and provide more accurate diagnostic suggestions or treatment plans. While retaining its general capabilities, the multi-modal model can acquire medical-industry-specific professional knowledge, thus playing a greater role in practical applications. For example, it can demonstrate higher efficiency in aspects such as auxiliary diagnosis, case analysis, and medical education.

[0097] (2) Obtain SAM and fine-tune SAM using the fourth training set.

[0098] The CLIP model is a powerful image encoder that performs well in processing a wide range of image types. However, when dealing with complex images in specific fields, such as medical images, its performance may be limited. Because medical images often contain a large amount of detail and specific anatomical structures, which pose higher requirements for the model's visual understanding ability. For example, the subtle differences in X-ray, CT scan, or MRI images may be crucial for disease diagnosis, but these details may not be easily captured by the CLIP model. Additionally, medical images often require precise semantic segmentation to distinguish different tissue types or lesion areas. Although the CLIP model can handle general object recognition tasks well, it may be insufficient in the fine analysis of medical images. Therefore, relying solely on the CLIP model may not be sufficient to meet the high-precision requirements of medical image processing. To overcome the above limitations, an advanced semantic segmentation model such as SAM is introduced in this application to enhance the visual perception ability of the multi-modal model.

[0099] SAM can automatically segment the regions of interest from the image and assign a class label to each pixel. This ability is particularly useful for medical images because it can help the multi-modal model more accurately identify and locate important anatomical structures or lesion sites. By integrating SAM into the multi-modal model, we can achieve the following goals:

[0100] ① Enhance image understanding ability: SAM can provide more detailed image segmentation results, which helps the multi-modal model to more deeply understand the structures and details in medical images;

[0101] ② Improve the accuracy of lesion detection: When processing medical images, SAM can highlight key regions, such as tumors, fractures, or other abnormalities, thus improving the detection accuracy of these regions by the multi-modal model;

[0102] ③Improve the overall performance of the model: By combining the segmentation ability of SAM and the general image understanding ability of CLIP, the multimodal model can exhibit better performance in medical image processing tasks.

[0103] In this embodiment, a fourth training set can be used to fine-tune SAM. The training samples in the fourth training set include medical images and masks, where the mask represents pixel-level annotations of the objects to be recognized in the medical images.

[0104] (3) Create a multimodal question-answering model based on the pre-trained multimodal model and the fine-tuned SAM model.

[0105] Specifically, the output of SAM can be used as additional image features and input into the multimodal model, or the structure of the multimodal model can be directly modified to integrate the functions of SAM. In this way, the multimodal model can utilize the additional information provided by SAM to improve its understanding and processing of medical images, and obtain the final multimodal question-answering model.

[0106] In an alternative embodiment, SAM is connected to the multimodal large model using a second vision-language connector to obtain a multimodal question-answering model.

[0107] After pre-training the multimodal model and fine-tuning SAM, both the multimodal model and SAM have their respective initial weights.

[0108] Step 302: Obtain a first training set and a second training set. The first training samples in the first training set include medical images, image annotations, simple question-and-answer texts with image annotations, and simple question-and-answer texts without image annotations. The second training samples in the second training set include medical images, image annotations, detailed question-and-answer texts with image annotations, and detailed question-and-answer texts without image annotations.

[0109] The first training set and the second training set are crucial for constructing a multimodal question-answering model that can understand and process medical images and texts. These training sets usually contain a large number of medical images and related clinical reports, diagnostic descriptions, or surgical guidelines, providing rich training materials for the multimodal question-answering model to learn the unique image and text features in the medical field. For example, the multimodal question-answering model can identify specific types of fractures or lesions from X-ray images and understand the corresponding medical terms.

[0110] Step 303: Fix the weights of the CLIP model, SAM, and LLM, and use the first training samples to train the weights of the first vision-language connector and the second language-vision connector.

[0111] Specifically, by fixing the weights of the CLIP model, SAM, and LLM, and training the weights of the first visual-language connector and the second language-visual connector using the first training sample, it may include:

[0112] (1) Extract features from the medical images in the first training sample using the CLIP model with fixed weights to obtain first image features.

[0113] (2) Encode the first image features using the first visual-language connector to obtain first image encodings.

[0114] (3) When training using the simple Q&A text without image annotations in the first training sample, extract features from the medical images using the SAM with fixed weights to obtain second image features; encode the second image features using the second visual-language connector to obtain second image encodings; randomly mask and then encode the simple Q&A text without image annotations using the tokenizer and embedding layer to obtain first text encodings.

[0115] Randomly masking the simple Q&A text means randomly selecting some words from the simple Q&A text for occlusion.

[0116] (4) When training using the image annotations and the simple Q&A text with image annotations in the first training sample, extract features from the medical images and the image annotations using the SAM with fixed weights to obtain second image features; encode the second image features using the second visual-language connector to obtain second image encodings; randomly mask and then encode the simple Q&A text with image annotations using the tokenizer and embedding layer to obtain first text encodings.

[0117] (5) Process the first image encodings, second image encodings, and first text encodings using the LLM with fixed weights to obtain the predicted text for the simple Q&A text.

[0118] In this embodiment, the first image encodings, second image encodings, and first text encodings can be concatenated, and then the concatenated encodings are processed using the LLM to obtain the predicted text.

[0119] The predicted text is the text obtained after restoring the randomly masked simple Q&A text.

[0120] (6) Update the weights of the first visual-language connector and the second language-visual connector according to the simple Q&A text and the predicted text.

[0121] In this embodiment, the simple Q&A text can be compared with the predicted text to determine whether the randomly masked words are the same as the restored words, and then the weights of the first visual-language connector and the second language-visual connector are updated based on the comparison result.

[0122] Step 304: Fix the weights of the CLIP model and SAM, and use the first training sample to train the weights of the first visual-language connector, the second language-visual connector, and the LLM.

[0123] Specifically, fixing the weights of the CLIP model and SAM and using the first training sample to train the weights of the first visual-language connector, the second language-visual connector, and the LLM may include:

[0124] (1) Use the CLIP model with fixed weights to extract features from the medical images in the first training sample to obtain the third image features.

[0125] (2) Use the first visual-language connector to encode the third image features to obtain the third image encoding.

[0126] (3) When training with the simple Q&A text without image annotations in the first training sample, use the SAM with fixed weights to extract features from the medical images to obtain the fourth image features; use the second visual-language connector to encode the fourth image features to obtain the fourth image encoding; use the tokenizer and embedding layer to randomly mask and then encode the simple Q&A text without image annotations to obtain the second text encoding.

[0127] (4) When training with the image annotations and the simple Q&A text with image annotations in the first training sample, use the SAM with fixed weights to extract features from the medical images and the image annotations to obtain the fourth image features; use the second visual-language connector to encode the fourth image features to obtain the fourth image encoding; use the tokenizer and embedding layer to randomly mask and then encode the simple Q&A text with image annotations to obtain the second text encoding.

[0128] (5) Use the LLM to process the third image encoding, the fourth image encoding, and the second text encoding to obtain the predicted text of the simple Q&A text.

[0129] (6) Update the weights of the first visual-language connector, the second language-visual connector, and the LLM according to the simple Q&A text and the predicted text.

[0130] Step 305: Fix the weights of the CLIP model and SAM, and use the second training sample to train the weights of the first visual-language connector, the second language-visual connector, and the LLM.

[0131] Specifically, fixing the weights of the CLIP model and SAM and using the second training sample to train the weights of the first visual-language connector, the second language-visual connector, and the LLM may include:

[0132] (1) Extract the features of the medical images in the second training sample using the CLIP model with fixed weights to obtain the fifth image features.

[0133] (2) Encode the fifth image features using the first vision-language connector to obtain the fifth image encoding.

[0134] (3) When training using the detailed Q&A text without image annotations in the second training sample, extract the features of the medical images using the SAM with fixed weights to obtain the sixth image features; encode the sixth image features using the second vision-language connector to obtain the sixth image encoding; mask and encode the answers in the detailed Q&A text without image annotations using the tokenizer and embedding layer to obtain the third text encoding.

[0135] Masking the answers in the detailed Q&A text means completely obscuring the answers and only retaining the questions.

[0136] (4) When training using the image annotations and the detailed Q&A text with image annotations in the second training sample, extract the features of the medical images and the image annotations using the SAM with fixed weights to obtain the sixth image features; encode the sixth image features using the second vision-language connector to obtain the sixth image encoding; mask and encode the answers in the detailed Q&A text with image annotations using the tokenizer and embedding layer to obtain the third text encoding.

[0137] (5) Process the fifth image encoding, the sixth image encoding, and the third text encoding using the LLM to obtain the predicted text of the answer.

[0138] The predicted text is the restoration of the completely masked answer.

[0139] (6) Update the weights of the first vision-language connector, the second vision-language connector, and the LLM according to the answer and the predicted text.

[0140] In this embodiment, the answer can be compared with the predicted text to determine whether the completely masked answer is the same as the restored answer, and then update the weights of the first vision-language connector, the second vision-language connector, and the LLM based on the comparison result.

[0141] Step 306: Load the trained weights into the multi-modal Q&A model to obtain the trained multi-modal Q&A model.

[0142] Step 307: Process the test sample using the multi-modal Q&A model to obtain the answer output by the multi-modal Q&A model; obtain the feedback information of the user on the answer; optimize the multi-modal Q&A model using the DPO algorithm and the feedback information.

[0143] In this embodiment, the Direct Preference Optimization (DPO) algorithm is used to optimize the multi-modal question-answering model. The DPO algorithm can learn human preference habits in the medical scenario. Traditional training methods are not sufficient to achieve this goal because they mainly rely on minimizing the loss function to optimize the model, which does not always guarantee that the output of the model is consistent with human preference habits. The DPO algorithm is an effective alignment optimization technique that enables the model to learn how to adjust its behavior according to user feedback to meet user expectations while minimizing the gap between the model behavior and the target behavior. For the multi-modal question-answering model, this means that the multi-modal question-answering model not only needs to be able to understand text information but also be able to process image data and synthesize this information to generate higher-quality answers.

[0144] 1. Value alignment: The output of the multi-modal question-answering model should conform to human values and social ethical standards. In the medical field, the multi-modal question-answering model needs to understand and abide by medical ethical guidelines, such as patient privacy protection, the principle of non-maleficence, etc. By using the DPO algorithm, the multi-modal question-answering model can be trained to learn and imitate the interaction process between doctors and patients, as well as the ethics and values reflected in these processes. For example, the multi-modal question-answering model can be taught not to disclose sensitive information without explicit consent.

[0145] 2. Preference alignment: What the DPO algorithm needs to focus on is how the multi-modal question-answering model adapts to the specific preferences of different users. In the medical consultation scenario, different patients may have different preferences. For example, some people may prefer to get concise answers, while others may hope to understand more detailed explanations. By using the DPO algorithm, we can collect user feedback on the output of the multi-modal question-answering model, which can include positive evaluations, negative evaluations, or specific improvement suggestions. Based on this feedback, the multi-modal question-answering model can gradually learn how to better meet the personalized needs of users.

[0146] 3. Professional alignment: The multi-modal question-answering model not only needs to provide accurate information but also needs to ensure that its output conforms to the professional knowledge system in the medical field. This requires that the multi-modal question-answering model not only be able to understand complex medical terms and concepts but also be able to give evidence-based suggestions. Through the DPO algorithm, the multi-modal question-answering model can continuously learn from professionals, such as by simulating the conversation between doctors and patients to improve the quality of its answers. In addition, medical experts can be introduced as supervisors, and they can provide professional feedback to help the multi-modal question-answering model further improve its professionalism.

[0147] Specifically, when testing the multi-modal question answering model with test samples, feedback information from a certain number of users can be collected, including the acceptance level of the results by users (for example, 0 means completely unacceptable, 5 means average, and 10 means very willing to accept) and the modification of the results by users; then, using the DPO algorithm, align the output results of the multi-modal question answering model on the feedback information set of user feedback to improve the generation performance of the multi-modal question answering model and make it meet user expectations and application scenarios.

[0148] In summary, the training method for the multi-modal question answering model of medical images provided by the embodiments of the present application trains the multi-modal question answering model through a large number of first training sets and second training sets, which can inject medical knowledge into the multi-modal question answering model and improve the professional ability of the multi-modal question answering model to understand and process medical images and texts.

[0149] Processing the medical image through pixel-level SAM can provide the multi-modal question answering model with fine-grained and dense visual perception ability and improve its ability to process fine and complex medical images.

[0150] By adding image annotations of regions of interest in the medical image, the attention positioning of the questions consulted by users can be reduced from the whole image to the regions of interest, improving the accuracy and fineness of the answers.

[0151] Through the DPO algorithm, feedback information from users on the answers output by the multi-modal question answering model can be collected, and value alignment, preference alignment, and professional alignment can be performed according to this feedback information to improve the safety, accuracy, and professionalism of the answers output by the multi-modal question answering model.

[0152] After obtaining the trained multi-modal question answering model, users can use this multi-modal question answering model to consult medical questions. Taking the example of a user consulting about a CT image of the lungs, the usage method of the multi-modal question answering model includes:

[0153] Step 401, use the CLIP model to extract features from the medical image to be processed to obtain the seventh image feature.

[0154] The user uploads a CT image of the lungs and optionally selects an interested region (if not, the whole image is noted), and inputs a consulting question (for example, what problems appear in the lung region in the picture?).

[0155] The CLIP model extracts features from the CT image of the lungs to obtain the first image feature.

[0156] Step 402, use the first vision-language connector to encode the seventh image feature to obtain the seventh image encoding.

[0157] VL-Connector1 encodes the seventh image feature to obtain the seventh image encoding (image tokens1).

[0158] Step 403, when there is no image annotation for the medical image, use SAM to extract features from the medical image to obtain the eighth image feature; use the second visual language connector to encode the eighth image feature to obtain the eighth image encoding; use the tokenizer and embedding layer to encode the problem to be processed to obtain the fourth text encoding.

[0159] SAM extracts features from the CT image of the lung to obtain the eighth image feature.

[0160] VL-Connector2 encodes the eighth image feature to obtain the eighth image encoding (image tokens2).

[0161] The tokenizer&embedding encodes the consultation question to obtain the fourth text encoding (text tokens).

[0162] Step 404, when there is an image annotation for the medical image, use SAM to extract features from the medical image and the image annotation to obtain the eighth image feature; use the second visual language connector to encode the eighth image feature to obtain the eighth image encoding; use the tokenizer and embedding layer to encode the problem to be processed to obtain the fourth text encoding.

[0163] SAM extracts features from the CT image of the lung and the region of interest to obtain the eighth image feature.

[0164] VL-Connector2 encodes the eighth image feature to obtain the eighth image encoding (image tokens2).

[0165] The tokenizer&embedding encodes the consultation question to obtain the fourth text encoding (text tokens).

[0166] Step 405, use the LLM to process the seventh image encoding, the eighth image encoding, and the fourth text encoding to obtain the answer output by the multi-modal Q&A model.

[0167] The seventh image encoding (image tokens1), the eighth image encoding (image tokens2), and the fourth text encoding (text tokens) are concatenated and fed into the LLM at the time step, and a professional answer is generated through analysis and reasoning.

[0168] For example, the answer generated by the multi-modal question answering model is: There is an air-fluid level visible in the right thoracic cavity, and the adjacent lung tissue is compressed with reduced volume. There are a few hazy shadows in both lungs, without consolidation shadows.

[0169] As Figure 5 shown, it shows a structural block diagram of a training device for a multi-modal question answering model of medical images provided by an embodiment of the present application. The training device for the multi-modal question answering model of medical images can be applied to a computer device. The training device for the multi-modal question answering model of medical images may include:

[0170] A creation module 510, configured to create a multi-modal question answering model. The multi-modal question answering model includes a CLIP model, a SAM, a first vision-language connector, a second vision-language connector, and an LLM. The CLIP model is connected to the LLM through the first vision-language connector, and the SAM is connected to the LLM through the second vision-language connector;

[0171] An acquisition module 520, configured to acquire a first training set and a second training set. The first training samples in the first training set include medical images, image annotations, simple question-answering texts with image annotations, and simple question-answering texts without image annotations. The second training samples in the second training set include medical images, image annotations, detailed question-answering texts with image annotations, and detailed question-answering texts without image annotations;

[0172] A training module 530, configured to fix the weights of the CLIP model, the SAM, and the LLM, and use the first training samples to train the weights of the first vision-language connector and the second vision-language connector;

[0173] The training module 530 is further configured to fix the weights of the CLIP model and the SAM, and use the first training samples to train the weights of the first vision-language connector, the second vision-language connector, and the LLM;

[0174] The training module 530 is further configured to fix the weights of the CLIP model and the SAM, and use the second training samples to train the weights of the first vision-language connector, the second vision-language connector, and the LLM;

[0175] A loading module 540, configured to load the trained weights into the multi-modal question answering model to obtain a trained multi-modal question answering model.

[0176] In an optional embodiment, the training module 530 is further configured to:

[0177] Use the multi-modal question answering model to process test samples to obtain the answers output by the multi-modal question answering model;

[0178] Obtain the feedback information of the user on the answer;

[0179] Optimize the multi-modal question-answering model using the Direct Preference Optimization (DPO) algorithm and feedback information.

[0180] In an alternative embodiment, the training module 530 is further configured to:

[0181] Extract features from the medical images in the first training sample using a CLIP model with fixed weights to obtain first image features;

[0182] Encode the first image features using the first vision-language connector to obtain first image encodings;

[0183] When training using the simple question-and-answer text without image annotations in the first training sample, extract features from the medical images using a fixed-weight SAM to obtain second image features; encode the second image features using the second vision-language connector to obtain second image encodings; randomly mask and encode the simple question-and-answer text without image annotations using a tokenization and embedding layer to obtain first text encodings;

[0184] When training using the image annotations and the simple question-and-answer text with image annotations in the first training sample, extract features from the medical images and the image annotations using a fixed-weight SAM to obtain second image features; encode the second image features using the second vision-language connector to obtain second image encodings; randomly mask and encode the simple question-and-answer text with image annotations using a tokenization and embedding layer to obtain first text encodings;

[0185] Process the first image encodings, second image encodings, and first text encodings using a fixed-weight LLM to obtain predicted texts for the simple question-and-answer texts;

[0186] Update the weights of the first vision-language connector and the second vision-language connector according to the simple question-and-answer texts and the predicted texts.

[0187] In an alternative embodiment, the training module 530 is further configured to:

[0188] Extract features from the medical images in the first training sample using a CLIP model with fixed weights to obtain third image features;

[0189] Encode the third image features using the first vision-language connector to obtain third image encodings;

[0190] When training using the simple Q&A text without image annotations in the first training sample, use the SAM with fixed weights to extract features from the medical image to obtain the fourth image feature; use the second vision-language connector to encode the fourth image feature to obtain the fourth image encoding; use the tokenization and embedding layer to randomly mask and then encode the simple Q&A text without image annotations to obtain the second text encoding;

[0191] When training using the image annotations and the simple Q&A text with image annotations in the first training sample, use the SAM with fixed weights to extract features from the medical image and the image annotations to obtain the fourth image feature; use the second vision-language connector to encode the fourth image feature to obtain the fourth image encoding; use the tokenization and embedding layer to randomly mask and then encode the simple Q&A text with image annotations to obtain the second text encoding;

[0192] Use the LLM to process the third image encoding, the fourth image encoding, and the second text encoding to obtain the predicted text of the simple Q&A text;

[0193] Update the weights of the first vision-language connector, the second vision-language connector, and the LLM according to the simple Q&A text and the predicted text.

[0194] In an optional embodiment, the training module 530 is further configured to:

[0195] Use the CLIP model with fixed weights to extract features from the medical images in the second training sample to obtain the fifth image feature;

[0196] Use the first vision-language connector to encode the fifth image feature to obtain the fifth image encoding;

[0197] When training using the detailed Q&A text without image annotations in the second training sample, use the SAM with fixed weights to extract features from the medical image to obtain the sixth image feature; use the second vision-language connector to encode the sixth image feature to obtain the sixth image encoding; use the tokenization and embedding layer to mask and then encode the answers in the detailed Q&A text without image annotations to obtain the third text encoding;

[0198] When training using the image annotations and the detailed Q&A text with image annotations in the second training sample, use the SAM with fixed weights to extract features from the medical image and the image annotations to obtain the sixth image feature; use the second vision-language connector to encode the sixth image feature to obtain the sixth image encoding; use the tokenization and embedding layer to mask and then encode the answers in the detailed Q&A text with image annotations to obtain the third text encoding;

[0199] Use the LLM to process the fifth image encoding, the sixth image encoding, and the third text encoding to obtain the predicted text of the answer;

[0200] Update the weights of the first vision-language connector, the second language-vision connector, and the LLM according to the answer and the predicted text.

[0201] In an optional embodiment, the creation module 510 is further configured to:

[0202] Obtain a multimodal model, and pre-train the multimodal model using a third training set. The multimodal model includes a CLIP model, a first vision-language connector, and an LLM;

[0203] Obtain SAM, and fine-tune SAM using a fourth training set;

[0204] Create a multimodal question-answering model according to the pre-trained multimodal model and the fine-tuned SAM model.

[0205] In an optional embodiment, the device further includes:

[0206] An extraction module, configured to use the CLIP model to extract features from the medical image to be processed, to obtain the seventh image feature;

[0207] An encoding module, configured to use the first vision-language connector to encode the seventh image feature to obtain the seventh image encoding;

[0208] The encoding module is further configured to, when the medical image has no image annotation, use SAM to extract features from the medical image to obtain the eighth image feature; use the second vision-language connector to encode the eighth image feature to obtain the eighth image encoding; use a tokenizer and an embedding layer to encode the question to be processed to obtain the fourth text encoding;

[0209] The encoding module is further configured to, when the medical image has an image annotation, use SAM to extract features from the medical image and the image annotation to obtain the eighth image feature; use the second vision-language connector to encode the eighth image feature to obtain the eighth image encoding; use a tokenizer and an embedding layer to encode the question to be processed to obtain the fourth text encoding;

[0210] A processing module, configured to use the LLM to process the seventh image encoding, the eighth image encoding, and the fourth text encoding to obtain the answer output by the multimodal question-answering model.

[0211] In summary, the multimodal question-answering model training device for medical images provided by the embodiments of the present application trains the multimodal question-answering model through a large number of first training sets and second training sets, can inject medical knowledge into the multimodal question-answering model, and improve the professional ability of the multimodal question-answering model to understand and process medical images and texts.

[0212] Processing medical images through pixel-level SAM can provide fine-grained and dense visual perception capabilities for multi-modal question-answering models, improving their ability to process fine and complex medical images.

[0213] By adding image annotations of regions of interest in medical images, the attention localization of the questions consulted by users can be narrowed down from the whole image to the regions of interest, improving the accuracy and fineness of the answers.

[0214] Through the DPO algorithm, feedback information on the answers output by the multi-modal question-answering model can be collected, and value alignment, preference alignment, and professional alignment can be performed according to this feedback information, improving the safety, accuracy, and professionalism of the answers output by the multi-modal question-answering model.

[0215] An embodiment of the present application provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned training method of the multi-modal question-answering model for medical images.

[0216] An embodiment of the present application provides a computer device, which includes the above-mentioned training device of the multi-modal question-answering model for any medical image.

[0217] It should be noted that when the above-mentioned training device of the multi-modal question-answering model for medical images trains the multi-modal question-answering model for medical images, only the above-mentioned division of each functional module is used as an example for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the training device of the multi-modal question-answering model for medical images is divided into different functional modules to complete all or part of the functions described above. In addition, the above-mentioned training device of the multi-modal question-answering model for medical images and the embodiment of the training method of the multi-modal question-answering model for medical images belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be repeated here.

[0218] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a disk, or an optical disc, etc.

[0219] The above does not intend to limit the embodiments of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present application shall be included in the protection scope of the embodiments of the present application.

Claims

1. A multimodal question-answering model training method for medical images, characterized in that: The method comprises: Creating a multimodal question-answering model, the multimodal question-answering model comprising a contrastive language-image pre-trained CLIP model, a segmentation model SAM, a first visual language connector, a second visual language connector, and a large language model LLM, the CLIP model being connected to the LLM via the first visual language connector, and the SAM being connected to the LLM via the second visual language connector; Obtaining a first training set and a second training set, wherein the first training samples in the first training set include medical images, image annotations, simple question-and-answer texts with image annotations, and simple question-and-answer texts without image annotations, and the second training samples in the second training set include medical images, image annotations, detailed question-and-answer texts with image annotations, and detailed question-and-answer texts without image annotations; Fixing the weights of the CLIP model, the SAM, and the LLM, and using the first training sample to train the weights of the first visual-linguistic connector and the second visual-linguistic connector; Fixing the weights of the CLIP model and the SAM, and using the first training sample to train the weights of the first visual-linguistic connector, the second language-visual connector, and the LLM; Fixing the weights of the CLIP model and the SAM, and using the second training sample to train the weights of the first visual-linguistic connector, the second language-visual connector, and the LLM; Loading the trained weights into the multimodal question-answering model to obtain a trained multimodal question-answering model; The method of creating a multimodal question-answering model includes: obtaining a multimodal model, pre-training the multimodal model using a third training set, wherein the multimodal model includes a CLIP model, a first visual language connector, and an LLM; obtaining a SAM, fine-tuning the SAM using a fourth training set; and creating the multimodal question-answering model based on the pre-trained multimodal model and the fine-tuned SAM model.

2. The multimodal question-answering model training method for medical images according to claim 1, characterized in that: The method further comprises: Processing the test sample using the multimodal question-answering model to obtain an answer output by the multimodal question-answering model; Obtaining user feedback information on the answer; The multimodal question-answering model is optimized using a direct preference optimization (DPO) algorithm and the feedback information.

3. The multimodal question-answering model training method for medical images according to claim 1, characterized in that: The step of fixing the weights of the CLIP model, the SAM, and the LLM, and using the first training sample to train the weights of the first visual language connector and the second language visual connector includes: Using a CLIP model with fixed weights to extract features from the medical images in the first training sample to obtain first image features; Encoding the first image feature using the first visual language connector to obtain a first image code; When the simple question-answer text without image annotations in the first training sample is used for training, the medical image is feature extracted using a SAM with a fixed weight to obtain a second image feature; the second image feature is encoded using the second visual language connector to obtain a second image encoding; the simple question-answer text without image annotations is randomly masked and encoded using a word segmentation and an embedding layer to obtain a first text encoding; When the image annotations and the simple question-answer text with image annotations in the first training sample are used for training, the medical image and the image annotation are extracted using a fixed-weight SAM to obtain a second image feature; the second image feature is encoded using the second visual language connector to obtain a second image encoding; the simple question-answer text with image annotations is encoded after random masking using a word segmentation and an embedding layer to obtain a first text encoding; Using a fixed-weight LLM to process the first image code, the second image code, and the first text code to obtain a predicted text of the simple question-and-answer text; The weights of the first visual-language connector and the second visual-language connector are updated according to the simple question-answer text and the predicted text.

4. The multimodal question-answering model training method for medical images according to claim 1, characterized in that: The step of fixing the weights of the CLIP model and the SAM, and using the first training sample to train the weights of the first visual language connector, the second language visual connector, and the LLM comprises: Using a CLIP model with fixed weights to extract features from the medical images in the first training sample, to obtain third image features; encoding the third image feature by using the first visual language connector to obtain a third image code; When the simple question-answer text without image annotations in the first training sample is used for training, the medical image is feature extracted by using the SAM with fixed weights to obtain a fourth image feature; the fourth image feature is encoded by using the second visual language connector to obtain a fourth image encoding; the simple question-answer text without image annotations is randomly masked and encoded by using the word segmentation and embedding layer to obtain a second text encoding; When the image annotations and the simple question-answer text with image annotations in the first training sample are used for training, the medical image and the image annotation are feature extracted using a fixed-weight SAM to obtain a fourth image feature; the fourth image feature is encoded using the second visual language connector to obtain a fourth image encoding; the simple question-answer text with image annotations is randomly masked and encoded using a word segmentation and an embedding layer to obtain a second text encoding; Using the LLM to process the third image code, the fourth image code and the second text code to obtain a predicted text of the simple question and answer text; The weights of the first visual-language connector, the second language-visual connector, and the LLM are updated according to the simple question-answer text and the predicted text.

5. The multimodal question-answering model training method for medical images according to claim 1, characterized in that: The step of fixing the weights of the CLIP model and the SAM, and using the second training sample to train the weights of the first visual language connector, the second language visual connector, and the LLM comprises: Using a CLIP model with fixed weights to extract features from the medical images in the second training sample, to obtain fifth image features; encoding the fifth image feature by using the first visual language connector to obtain a fifth image code; When the detailed question-and-answer text without image annotations in the second training sample is used for training, the medical image is feature extracted using a SAM with a fixed weight to obtain a sixth image feature; the sixth image feature is encoded using the second visual language connector to obtain a sixth image encoding; the answer in the detailed question-and-answer text without image annotations is masked and encoded using a word segmentation and an embedding layer to obtain a third text encoding; When the image annotations and the detailed question-and-answer text with the image annotations in the second training sample are used for training, the medical image and the image annotation are feature extracted using a SAM with fixed weights to obtain a sixth image feature; the sixth image feature is encoded using the second visual language connector to obtain a sixth image encoding; the answer in the detailed question-and-answer text with the image annotations is masked and encoded using a word segmentation and an embedding layer to obtain a third text encoding; Processing the fifth image code, the sixth image code and the third text code by using the LLM to obtain a predicted text of the answer; The weights of the first visual-linguistic connector, the second linguistic-visual connector, and the LLM are updated according to the answer and the predicted text.

6. The multimodal question-answering model training method for medical images according to claim 1, characterized in that: The method further comprises: Using the CLIP model to perform feature extraction on the medical image to be processed, to obtain a seventh image feature; encoding the seventh image feature by using the first visual language connector to obtain a seventh image code; When the medical image has no image annotation, extracting features from the medical image using the SAM to obtain an eighth image feature; encoding the eighth image feature using the second visual language connector to obtain an eighth image code; encoding the problem to be processed using a word segmentation and an embedding layer to obtain a fourth text code; When the medical image has an image annotation, the SAM is used to extract features from the medical image and the image annotation to obtain an eighth image feature; the second visual language connector is used to encode the eighth image feature to obtain an eighth image code; and the word segmentation and embedding layer are used to encode the problem to be processed to obtain a fourth text code; The seventh image code, the eighth image code and the fourth text code are processed by using the LLM to obtain an answer output by the multimodal question-answering model.

7. A multimodal question-answering model training device for medical images, characterized in that: The device comprises: A creation module, used to create a multimodal question-answering model, wherein the multimodal question-answering model includes a contrastive language-image pre-trained CLIP model, a segmentation model SAM, a first visual language connector, a second visual language connector, and a large language model LLM, wherein the CLIP model is connected to the LLM via the first visual language connector, and the SAM is connected to the LLM via the second visual language connector; an acquisition module, configured to acquire a first training set and a second training set, wherein the first training samples in the first training set include medical images, image annotations, simple question-and-answer texts with image annotations, and simple question-and-answer texts without image annotations, and the second training samples in the second training set include medical images, image annotations, detailed question-and-answer texts with image annotations, and detailed question-and-answer texts without image annotations; A training module, used for fixing the weights of the CLIP model, the SAM and the LLM, and training the weights of the first visual-linguistic connector and the second visual-linguistic connector using the first training sample; The training module is further used to fix the weights of the CLIP model and the SAM, and train the weights of the first visual language connector, the second language visual connector, and the LLM using the first training sample; The training module is further used to fix the weights of the CLIP model and the SAM, and train the weights of the first visual language connector, the second language visual connector, and the LLM using the second training sample; A loading module, used to load the trained weights into the multimodal question-answering model to obtain a trained multimodal question-answering model; The creation module is also used to: obtain a multimodal model, pre-train the multimodal model using a third training set, the multimodal model including a CLIP model, a first visual language connector and an LLM; obtain a SAM, fine-tune the SAM using a fourth training set; and create the multimodal question-answering model based on the pre-trained multimodal model and the fine-tuned SAM model.

8. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the multimodal question-answering model training method for medical images as described in any one of claims 1 to 6.

9. A computer device, characterized in that: The computer device includes: the multimodal question-answering model training device for medical images as described in claim 7.

Citation Information

Patent Citations

  • X-ray image diagnosis method and device based on dialogue function and electronic equipment

    CN118098558A

  • Automatic auxiliary diagnosis method based on multi-modal LLM and model construction method thereof

    CN118098564A