A medical visual question answering method and system based on multi-task modeling
By adopting multi-task modeling and cross-modal alignment technology in medical visual question-and-answer systems, the existing system's shortcomings in question-and-answer type distinction, site feature extraction and interactive flexibility are solved, and higher question-and-answer accuracy and user experience are achieved, especially in the recognition of fine lesion features in medical images.
Patent Information
- Application Number
- CN202411711397.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-11-27
AI Technical Summary
The existing medical visual question and answer system has shortcomings in question-and-answer type distinction, site feature extraction, detail capture ability and interaction flexibility, which makes it difficult for the question-and-answer accuracy and user interaction experience to meet medical needs.
Using a multi-task modeling method, using technologies such as feature learning, cross-modal alignment and multi-task learning, combined with SAM models and large models, we realize graphic category prediction, important mask prediction and multi-round question-answer analysis, and optimize the performance and interactive experience of the medical visual question-and-answer system.
Through automated importance mask generation and multi-layer cascaded weighted fusion model, the accuracy of image-problem alignment and multi-modal information expression capabilities are improved, the flexibility and robustness of the model in complex problem environments are enhanced, and the accuracy of identification of fine lesion features in medical images is improved.
Smart Images

Figure CN119202334B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and medical intelligence technology, and relates to a medical vision question-answering method and system based on multi-task modeling. Background Art
[0002] Medical visual question answering is a multimodal task that combines computer vision and natural language processing. It predicts the answer to a question by inputting a patient's medical image and related questions. The medical visual question answering model has great potential for clinical application. On the one hand, it can help inexperienced doctors provide a second opinion and improve diagnostic efficiency and accuracy. On the other hand, the medical visual question answering model can provide patients with fast and accurate answers, alleviating the problem of unbalanced medical resources.
[0003] However, existing medical visual question-answering systems still have many technical limitations in some practical application scenarios, such as obvious deficiencies in question-answering type distinction, part feature extraction, detail capture capabilities, and interactive flexibility, which make it difficult for their question-answering accuracy and user interaction experience to meet medical needs. Specifically:
[0004] 1) Current medical visual question-answering systems lack effective classification strategies when dealing with closed and open questions, and fail to fully consider the essential differences between these two types of questions. Closed questions are mostly simple binary choices, such as "yes / no" or "yes / no", which are suitable for quick and clear answers; while open questions often involve complex descriptive answers and require a more detailed semantic understanding of the question content. This defect makes it difficult for the system to accurately understand the question type, affecting its question-answering effect in different contexts.
[0005] 2) The visual encoder of the existing model has difficulty effectively distinguishing the features of pathological images of different parts (such as chest, abdominal and brain CT images). CT images of different parts have certain similarities in grayscale, texture and structure. In the absence of image part classification function, the existing model is prone to confuse CT images of different parts, resulting in misjudgment of lesions in different organs. This confusion normalization not only affects the accuracy of image classification, but also easily leads to misjudgment of the model in the pathological analysis link, which seriously restricts the application value of the system in multi-part image analysis.
[0006] 3) Medical images usually have low color contrast and high similarity features, and are easily disturbed by noise, artifacts and other factors, making it difficult for existing medical image encoders to effectively distinguish lesions from normal tissues. Especially in the process of identifying subtle lesions, existing visual question answering models are often unable to accurately extract the detailed features of lesions, resulting in insufficient accuracy in distinguishing lesions from normal tissues. This problem is mainly attributed to the model's lack of ability to capture high-precision information and the inability to achieve precise alignment with text, which limits the model's performance in high-resolution medical images.
[0007] 4) Existing medical visual question-answering systems mostly adopt a fixed multi-classification output format of single-round question-answering, which lacks the ability to interact deeply with users. Such models can often only provide fixed answers and cannot be dynamically adjusted according to the context, resulting in users' needs not being met in complex diagnosis or multi-round questioning situations. In addition, the limited length of single-round question-answering makes it difficult for the system to handle multi-round, context-sensitive questions, significantly reducing its practicality and scalability in actual medical applications.
[0008] In response to the above problems, a new medical visual question answering solution is urgently needed to optimize the performance of the medical visual question answering system, improve its answering ability and interactive experience in diverse medical images, and better meet the application needs in the medical field. Summary of the invention
[0009] The purpose of the present invention is to address the deficiencies in the prior art and provide a medical visual question-answering method and system based on multi-task modeling. The method utilizes feature learning, cross-modal alignment, multi-task learning and other technologies, combined with the SAM model and a large model, to achieve multi-task modeling of image and text category prediction, important mask prediction and multi-round question-answering analysis, thereby optimizing the performance and interactive experience of the medical visual question-answering system.
[0010] In order to achieve the purpose of the above invention, the present invention specifically adopts the following technical solutions:
[0011] In a first aspect, a medical visual question answering method based on multi-task modeling of the present invention comprises:
[0012] Step s1: Load the user's historical question and answer data, and simultaneously obtain the medical image to be analyzed and the preliminary question instructions of the current user;
[0013] Step s2: extracting features from the medical image to be analyzed and the preliminary question instruction, that is, extracting features from the medical image using a visual encoder to obtain image features, inputting the historical question and answer data and the preliminary question instruction into the dialogue model for semantic analysis, identifying and outputting question instructions suitable for image analysis, and extracting question features through a text encoder to obtain text features;
[0014] Step s3: The obtained text features and image features are processed by self-attention image importance weighting to highlight the image area with stronger relevance to the question instruction, and the weighted image features are obtained. After the image and text are aligned, the image and text fusion representation is obtained;
[0015] Step s4: Input the acquired image-text fusion representation into the multi-target output projection layer for multi-task prediction. The predicted output includes question answers, image categories, and important area masks. The question answers and image categories are input into the large dialogue model, combined with the dialogue context and multi-round interaction data, combined with the important area masks to finally generate a diagnostic opinion with detailed description.
[0016] In the above technical solution, further, in step s2, the visual encoder adopts the ResNet architecture and is pre-trained on medical image data to enable it to have the ability to extract medical image features. It uses the convolutional neural network CNN layer to hierarchically encode image features, and each convolution layer can gradually extract features of different levels, and finally form a high-dimensional image feature. ;
[0017] The text encoder uses an LSTM network to embed the question instructions and input them into the LSTM network, gradually extracting the keywords and semantic association information in the question instructions, and generating text features with time sequence information. .
[0018] Furthermore, the self-attention image importance weighting module includes a mapping layer, a self-attention layer and a feedforward output layer connected in sequence, wherein:
[0019] The mapping layer transforms text features and image features Projection to high-dimensional space, projecting the original input to a unified feature dimension through matrix mapping;
[0020] The self-attention layer calculates the similarity scores of the mapped text features and image features and determines the important areas in the image features, where the text features and image features Similarity score The calculation formula is as follows:
[0021]
[0022] in, and are the weight matrices of text features and image features respectively, represents the scaling factor of the feature dimension; the attention score is normalized by the softmax function so that the sum of the importance weights of each region is 1; the weighted image feature is expressed as:
[0023]
[0024] Weighted image features After being processed by the feedforward output layer, the final weighted features are output. The feedforward output layer contains multiple nonlinear activation layers and regularization layers. The output weighted image features for:
[0025]
[0026] in, and are the weight matrix and bias term of the feedforward output layer, respectively. represents a non-linear activation function.
[0027] Furthermore, the image and text alignment described in step s3 is specifically as follows:
[0028] Weighted image features and text features Input cascade image-text alignment weighted module, which includes feature cascade layer, cross-modal attention fusion layer, and hierarchical feedforward network;
[0029] The feature cascade layer is used to cascade the weighted image features with the text features to obtain the cascaded feature representation. ; The cross-modal attention fusion layer generates a cross-modal weight matrix through the self-attention mechanism :
[0030] in, and is the projection weight matrix, and the output cross-modal weighted features are , in obtaining cross-modal weighted features After that, the cascaded image-text alignment weighted module inputs it into the hierarchical feedforward network; the hierarchical feedforward network consists of multiple layers of fully connected layers and nonlinear activation layers. Its structural design is to refine cross-modal features layer by layer and generate a more information-dense fusion representation. The calculation formula of the hierarchical feedforward network is: ,in and are the weight matrix and bias term of the feedforward network, and the final output of the hierarchical feedforward network After normalization, a unified image-text fusion representation is generated , the representation integrates deep image and text interaction information.
[0031] Furthermore, the multi-target output projection layer in step s4 includes an answer prediction model, an image category prediction model and an important difference predictor; the answer prediction model is used to generate answer categories related to user questions, and is composed of multiple layers of fully connected layers. By projecting the input features layer by layer, the final output is a probability distribution vector of the answer category; the image category prediction model is tasked with classifying the anatomical parts of medical images, and is composed of multiple layers of fully connected layers, and the final output is a probability distribution vector of the image category; the important area predictor is used to locate key areas related to diagnosis in medical images, and adopts a decoder network structure, which is composed of multiple convolutional layers Conv2D and upsampling layers UpSampling2D, and achieves accurate area positioning by gradually restoring the spatial resolution.
[0032] Furthermore, when the multi-target output projection layer performs multi-target training, the method for constructing the data label training data of the importance mask is as follows:
[0033] 1) Construct an indicator point matrix, where each indicator point acts as a local indicator to help the medical image segmentation model capture areas in the image that may contain important information;
[0034] 2) Input the constructed cue point matrix and the medical image into the Med-SAM model, prompting the model to identify the salient areas in the entire image and perform segmentation operations, generate multiple segmentation regions, and select the top N segmentation regions with the highest confidence. These regions are regarded as the most informative parts of the image;
[0035] 3) Based on the association problem of medical images, the segmented regions are further screened; among the first N high-confidence segmented regions obtained, the regions most relevant to the question are manually screened and identified based on the semantic information provided by the medical question, ensuring that the generated mask data focuses on the image content required to answer the medical question;
[0036] 4) The selected areas are labeled as "importance masks". The importance masks are not only the areas of interest to the model, but also the core data used as labels in medical visual question answering tasks. After the labeling is completed, these importance masks are saved in the training set so that they can be used as label data for model learning during multi-target training.
[0037] In a second aspect, a medical visual question answering system based on multi-task modeling of the present invention is used to implement the medical visual question answering method based on multi-task modeling as described in any one of the above items, and the system includes:
[0038] A user interaction module, used to interact with the user, the interaction includes receiving medical images uploaded by the user, preliminary question instructions input, and displaying answers and diagnosis results generated by the system, including text replies, image labels and mask displays, and supporting multi-round interaction modes;
[0039] The inference prediction module is used to process the received medical images and preliminary question instructions, including a medical image encoder, a medical dialogue model, a question text encoder, a hierarchical image weighter, a picture-text feature fusion device and a multi-target predictor, wherein the medical image encoder is used to extract the features of the medical image and generate image features; the medical dialogue model integrates the knowledge in the medical field and is used to integrate the user's historical question and answer data and the current preliminary question instructions to generate question instructions suitable for image analysis; the question text encoder is used to perform semantic analysis and feature extraction on the question instructions to generate text features; the hierarchical image weighter is used to weight the image features layer by layer so that the model can focus on the key areas related to the current question, and the weighting of the important areas of the image is achieved through the self-attention mechanism; the picture-text feature fusion device fuses the image features and the text features at multiple levels, ensures the correlation between the two modal features through the cascade network structure, and achieves semantic alignment; the multi-target predictor realizes multi-task output, including answer prediction, image category prediction and important area labeling, the answer prediction generates a diagnostic suggestion matching the question, the image category predictor is used to determine the type of image part, and the important area predictor generates an important area mask to label the key diagnostic area for the user;
[0040] The memory storage module is used to save medical images, interactive dialogue data and historical diagnosis reports uploaded by users, and allows viewing of historical question and answer data.
[0041] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the medical visual question-answering method based on multi-task modeling as described in any one of the above items.
[0042] In a fourth aspect, the present invention provides an electronic device, the device comprising:
[0043] one or more processors;
[0044] A memory for storing one or more programs;
[0045] When the one or more programs are executed by the one or more processors, the one or more processors implement the medical visual question answering method based on multi-task modeling as described in any one of the above items.
[0046] Compared with the prior art, the present invention has the following beneficial effects:
[0047] (1) The present invention uses the Med-SAM model to automatically segment and generate importance masks on existing medical image datasets to accurately mark key areas related to questions. By filtering and retaining mask data that is highly relevant to questions, the model can learn image features that are more diagnostically significant, thereby achieving higher-precision image-question alignment. This automated mask generation method not only reduces the workload of manual annotation, but also improves the expression of the importance of image regions, enabling the model to more effectively capture disease features and improve the accuracy and reliability of medical question answering.
[0048] (2) The present invention adopts a multi-layer cascade weighted fusion model for feature fusion to enhance the ability to express multimodal features. By adjusting the weights layer by layer, the features of different modalities are fully integrated and emphasized at each layer. The model architecture can iteratively optimize image and text features at different levels, thereby achieving richer multimodal information expression. This hierarchical fusion strategy avoids the information loss and over-simplification problems caused by single-layer fusion, making the model more flexible and robust in complex problem environments, while ensuring the integrity of multimodal information and improving the accuracy of identifying fine lesion features in medical images.
[0049] (3) The present invention models multi-task objectives at the prediction layer. Compared with the traditional single-objective predictor, the multi-objective loss optimization enables the model to simultaneously distinguish the features of multiple medical image parts (such as chest, abdomen, brain, etc.) and perform refined recognition. The multi-task objective optimization strategy effectively avoids the confusion of different modal features, helps the model accurately identify the characteristics of specific organs and parts, and thus reduces the risk of prediction errors. In addition, the introduction of important area prediction objectives enables the model to focus on question-related areas and accurately align image and text features, thereby improving the accuracy of prediction and the diagnostic value of the question-answering process. The optimization strategy of multi-task learning also enhances the model's perception of details in specific areas, which helps to make more accurate medical diagnoses in multimodal complex information.
[0050] (4) Based on the medical dialogue big model, the present invention constructs a question-and-answer process with rich question extraction and answering functions, which has higher flexibility and interactivity compared with the traditional medical question-and-answer model based on simple classifiers. The medical big model can input the user's historical conversation records and current consultation needs together to form multiple rounds of interactive questions and answers, effectively solving the limitation that the original model can only perform single-round questions and answers. Through multiple rounds of consideration of the conversation context, the model can provide more targeted answers to the user's diverse needs, support in-depth discussions of complex medical issues, and provide users with more instructive medical advice. In addition, this process enriches the content of diagnostic feedback, allowing the model to provide dynamic responses that are consistent with the user context, greatly improving the user experience and the actual diagnostic value of the question and answer. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a schematic diagram of the process of the present invention;
[0052] Figure 2 It is a structural schematic diagram of the feature fusion module in the method of the present invention;
[0053] Figure 3 A flow chart for constructing training data for important area prediction in one embodiment of the present invention;
[0054] Figure 4 It is a structural block diagram of a medical visual question answering system in one embodiment of the present invention;
[0055] Figure 5 Schematic diagram of the overall structure of a medical visual question answering system based on multi-task modeling in one embodiment of the present invention.
[0056] Figure 6 This is a comparison chart of the medical image question-answering effects of a system using the present invention and a baseline model system that does not use multi-task learning in an embodiment of the present invention. DETAILED DESCRIPTION
[0057] The technical solution of the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0058] According to a specific embodiment of the present invention, the specific process of the medical visual question answering method based on multi-task modeling is as follows: Figure 1 As shown. It includes steps s1 to s4:
[0059] Step s1: Load the user's historical data, obtain the medical images uploaded by the user and the preliminary question instructions.
[0060] The system first loads the user's historical question-and-answer data from the memory module, including the medical images, interaction records, and corresponding question-and-answer information previously uploaded by the user. Next, the user interacts with the medical image question-and-answer system in natural language through the chat box, inputting medical images to be diagnosed or asking questions. The user uploads new medical images and question instructions. These images and instructions will be used as input to the model for subsequent processing by the system to generate targeted diagnoses. This step realizes the contextual association of multiple rounds of questions and answers by providing the user's historical information, ensuring that the system fully considers the user's background and historical data support when generating answers.
[0061] Step s2: Extract features from the medical images and preliminary question instructions uploaded by the user. The medical images are input into the visual encoder of the medical model for feature extraction to obtain visual features. The user's historical conversation data and preliminary question instructions are input into the conversation model for understanding, and the question instructions are output. The question features are extracted through the text encoder to extract features such as keywords in the question instructions to form a multimodal input of images and text.
[0062] The visual encoder extracts features from the medical images uploaded by users. The visual encoder uses the ResNet architecture and is pre-trained on a large amount of medical imaging data, enabling it to capture the unique lesion features in medical images while effectively suppressing interference information such as noise and artifacts. The ResNet network model uses convolutional neural network (CNN) layers to hierarchically encode image features and extract multi-scale feature vectors. In each convolutional layer, features of different levels can be gradually extracted to form a high-dimensional visual feature representation. , that is, extracting multi-level visual information from the original image, including detailed features of the lesion area and various anatomical structure features. These features are used to associate with question instructions in subsequent steps to improve the model's question-answering accuracy and diagnostic quality.
[0063] The user's initial question instructions and historical conversation data are input into the large medical conversation model (LLaVA model) for semantic understanding and question parsing. The LLaVA model has learned rich medical expertise and long text comprehension capabilities by fine-tuning SFT on the medical question-answering dataset. It can identify the main instruction content of the current question and generate question instructions that are highly relevant to the current question and suitable for image analysis.
[0064] The question instructions generated by the dialogue model are then input into the text encoder for further feature extraction. The text encoder uses an LSTM network. By embedding the question instructions and inputting them into the LSTM network, the keywords, semantic associations and other information in the question are gradually extracted. The LSTM memory unit can retain the context information of the question and effectively model the keyword features in the input sequence, thereby generating a text feature representation with time series information. (Also referred to as question feature in the present invention), this text feature provides semantic support for the subsequent image-text alignment process.
[0065] Step s3: The text features and image features are first passed through the self-attention image importance weighting module to obtain weighted image features. In this module, the importance of image features is weighted based on text features, highlighting the image areas that are more relevant to the problem, while weakening the background information that is irrelevant to the diagnosis. After that, they are input into the image-text alignment module for further deep fusion to obtain the image-text fusion representation, which models high-dimensional visual and semantic fusion information.
[0066] Text features and image features Input to the self-attention image importance weighting module, where n and m represent the sequence lengths of image and question features respectively, and Represents the respective feature dimensions. This module calculates the attention weights of image features through the self-attention mechanism to identify and highlight the key image areas related to the problem and generate weighted image features. The self-attention mechanism dynamically adjusts the weight of each feature by learning the weight matrix, ensuring that the model can focus on the image area that best supports question answering. According to a specific embodiment of the present invention, specifically:
[0067] The self-attention image importance weighting module consists of a mapping layer, a self-attention layer, and a feedforward output layer. It uses the semantic information of the question to weight the image features to achieve adaptive screening of image regions, thereby focusing on the image regions most relevant to the question and weakening the influence of irrelevant background.
[0068] The mapping layer transforms the problem features and image features Projecting to a high-dimensional space, the original input is projected to a unified feature dimension through matrix mapping. That is, the image features and question features are respectively projected through the weight matrix and Mapped to a unified high-dimensional space, where Represents the dimension after mapping. The function of this layer is to convert data of different modalities into a consistent feature space, providing a basis for subsequent similarity calculations.
[0069] The self-attention layer calculates the similarity scores between the question features and the image features, determines the important areas in the image features, and generates weighted features that focus on specific image regions. Specifically, the question features and image features The weighted similarity score calculation formula is as follows:
[0070]
[0071] in, and are the weight matrices of question features and image features respectively, Represents the scaling factor of the feature dimension. The attention score is normalized by the softmax function so that the sum of the importance weights of each region is 1, thereby ensuring the standardization of the weighting. The weighted image feature can be expressed as:
[0072]
[0073] It represents the image information that is focused on under the guidance of the problem features. Weighted image features After being processed by the feedforward output layer, the final weighted features are output. The feedforward output layer contains multiple nonlinear activation layers and regularization layers, and the weight matrix and bias are used to ensure that the output features have stronger discriminative ability. Output weighted features for:
[0074]
[0075] in, and are the weight matrix and bias term of the feedforward layer, respectively. Represents a nonlinear activation function, which is used to increase the nonlinear representation of features. This output is used as the input of the subsequent image-text alignment module to provide filtered image information for the image-text alignment process.
[0076] like Figure 2 , the weighted image features and text features Input the image-text alignment module, and generate the final image-text fusion representation through further fusion operations , which is used as the final output of medical question answering. This module further associates the question features with the image features, enabling the model to more accurately understand the semantic relevance between the lesion features in the image and the question instructions. The output formula of the alignment operation is:
[0077]
[0078] in, Represents the parameter function of the image-text alignment model. This module aims to establish a highly matched association between image features and text features, thereby improving the question-answering system's ability to focus on image features and the accuracy of question responses.
[0079] According to a specific embodiment of the present invention, Figure 2 The image-text alignment module adopts a cascaded image-text alignment weighted module to further deeply fuse weighted image features with text features to generate a unified image-text fusion representation. This module achieves complementary and collaborative representation of visual and semantic information through a multi-layer cascade fusion structure. The module structure includes a feature cascade layer, a cross-modal attention fusion layer, and a hierarchical feedforward network:
[0080] Weighted Image Features With text features First, they are connected in series through a feature cascade layer to form a joint feature vector. This cascade layer uses a multi-layer neural network to align the dimensions of features of different modalities. Through the cascade operation, the image and text information are integrated into a high-dimensional feature vector , the formula is as follows:
[0081]
[0082] in, Represents a feature concatenation operation. The concatenated features can be further compressed into a unified feature space through dimensionality reduction operations to eliminate the dimensionality differences between different modalities.
[0083] The cross-modal attention fusion layer is the core part of the module, which is used to capture the relationship between image and text features. The cross-modal attention fusion layer generates a fusion weight matrix through the self-attention mechanism. The calculation formula is as follows:
[0084]
[0085] in, and is the projection weight matrix of different modal features, is the feature scaling factor. The weight matrix generated The cross-modal features are weighted, that is, the mutual weighting between image features and text features (i.e., question features) is achieved, aiming to capture the interactive information of features of different modalities and to control the importance of features of different modalities in the fusion process.
[0086] The cross-modal weighted features can be expressed as:
[0087]
[0088] Weighted cross-modal features Input into the hierarchical feedforward network. The network consists of multiple fully connected layers and nonlinear activation layers. Its structural design is to refine cross-modal features layer by layer and generate a more information-dense fusion representation. The processing formula of the hierarchical feedforward network is as follows:
[0089]
[0090] in, and is the weight matrix and bias term of the feedforward network, Represents an activation function (e.g. ReLU). A sparse and efficient image-text fusion representation is generated through layer-by-layer refinement of a multi-layer structure.
[0091] Output of a layered feedforward network The final image-text fusion representation contains the deep interaction information between the image and the text. The representation combines visual and semantic information in a high-dimensional way. The final representation formula can be expressed as:
[0092]
[0093] in, Represents a feature normalization operation to ensure the numerical stability of the representation when it is input to downstream tasks.
[0094] It can be seen that the self-attention image importance weighting module achieves weight distribution and dynamic adjustment of image and text features by focusing on the relationship between keywords in the question instructions and image features, highlighting important image areas and suppressing background noise. The image-text alignment module further performs a joint analysis of the weighted features to ensure semantic consistency between the image and text.
[0095] Step s4: Fusion features of acquired image and question Input the multi-target output projection layer to perform multi-task prediction. The prediction output includes question answers, image categories, and important area masks. The question answers and image categories are input into the dialogue model, combined with the dialogue context and multi-round interaction data, and finally generate a diagnostic opinion with detailed description.
[0096] Input fusion features The dimension of is d. After the dimensionality reduction processing of multiple layers of fully connected layers, the final output shape is (1, ), each element represents the model's probability prediction value for a certain answer category.
[0097] The answer prediction model is used to generate answer categories related to user questions. This module consists of multiple fully connected layers. By projecting the input features layer by layer, the final output is the probability distribution vector of the answer category. The specific output dimension is ,in Represents the total number of possible answer categories.
[0098] The task of the image category prediction model is to classify the anatomical parts of medical images, such as chest, abdomen, brain and other CT image categories. This module is also composed of multiple fully connected layers, and finally outputs the probability distribution vector of the image category, which is used to identify the category of the medical image uploaded by the user. The output of the image category prediction is ,in Represents the total number of image categories. The output shape obtained by projecting the fusion features layer by layer through multiple fully connected layers is (1, ), each value represents the probability of a specific class.
[0099] The important region predictor is used to locate the key diagnosis-related regions in medical images, helping users to intuitively understand the location of lesions that need to be focused on. This module adopts a decoder network structure, which is mainly composed of multiple convolutional layers Conv2D and upsampling layers UpSampling2D. It achieves accurate region positioning by gradually restoring the spatial resolution; the input of the decoder is the image-question fusion feature , whose dimension is d. After multiple layers of convolution and upsampling, the decoder gradually generates a 2D mask output that matches the size of the original image. Specifically, the output mask shape is , where H and W are the height and width of the image, respectively, representing the probability value of each pixel being the lesion area. The last convolution layer is 1 The convolution kernel of 1 is combined with the Sigmoid activation function to generate a binary prediction for each pixel: ,in, Represents the Sigmoid activation function, ensuring that the output value is in the range of [0,1]. This mask image can identify the location of the lesion area for users and assist them in understanding image diagnosis.
[0100] According to a specific embodiment of the present invention, the multi-target output projection layer in the present invention is implemented by a multi-target prediction module, including an answer predictor, an image category predictor, and an image decoder. Each output module is specially designed to cope with different prediction targets:
[0101] (1) The first type of answer predictor module is used to generate the most relevant answer to the user's question. The structure is a multi-layer fully connected layer, activation function and output layer. The answer predictor receives the image and text fusion representation as input and classifies the input features through a multi-layer fully connected structure to generate the most likely answer.
[0102] In the training loss function, the cross entropy loss function is used To measure the difference between the predicted answer and the true answer. The specific calculation formula is as follows:
[0103]
[0104] Where C is the total number of answer categories, is the one-hot encoding of the true label, is the probability value of the c-th answer predicted by the model. By minimizing , the system can gradually improve the recognition accuracy of the correct answer.
[0105] (2) The second type of image category predictor is used to classify the anatomical parts of medical images uploaded by users, including CT image categories such as lungs, chest, and brain. This module consists of a fully connected layer, an activation function, and an output layer, and infers the image category by classifying image features.
[0106] The loss function of the image category predictor is the cross entropy loss, and the specific calculation formula is as follows:
[0107]
[0108] Among them, K is the total number of image categories, is the one-hot encoding of the true label of the image category, is the probability value of the k-th image category predicted by the model. By minimizing , the model can improve the classification ability of medical images of different anatomical parts.
[0109] (3) The third type of image decoder is used to generate image masks related to the problem, showing the key lesion areas related to diagnosis. By generating a binary mask, users can intuitively identify the areas in the image that need to be focused on. The image decoder structure consists of multiple convolutional layers and upsampling layers. By gradually restoring the spatial resolution of the image, it outputs a mask that is the same size as the original image.
[0110] The decoder network takes as input the weighted image features and generates the final mask through the following steps:
[0111] First, the input feature map passes through a convolutional layer with 256 3x3 filters, followed by Batch Normalization and ReLU activation function to enhance the attention to details.
[0112] Then, the spatial resolution is gradually restored by upsampling by a factor of 2. This process is repeated three times, and the number of filters in the convolutional layer is reduced to 128, 64, and 32, respectively.
[0113] Finally, a single-channel 1x1 convolutional layer uses a Sigmoid activation function to generate a mask ranging from 0 to 1, representing the probability distribution of the key area.
[0114] During training, the loss function of the image decoder output mask uses binary cross entropy loss , to measure the difference between the predicted mask and the actual important mask. Its formula is:
[0115]
[0116] Where N is the total number of pixels in the mask, represents the true mask value of the i-th pixel, Represents the predicted mask value of the i-th pixel.
[0117] In summary, the final multi-task loss function To answer the weighted sum of prediction loss, image category prediction loss and mask prediction loss, the model is guided to achieve balanced optimization on the three tasks. The combination of loss functions is as follows:
[0118]
[0119] in, , and is the weight factor for each loss, which is used to balance the importance of the three tasks during training.
[0120] When performing multi-target training in the output layer of step s4 above, the data labels of the importance mask are given by Figure 3 The data structure process shown is generated. Figure 3 The flowchart of training data construction of the medical visual question answering method in a specific embodiment of the present invention includes the following:
[0121] s1, using the pre-trained Med-SAM model, before automatically segmenting the medical image, a cue point matrix is constructed, and a multi-point cue strategy is adopted to guide the model to complete the segmentation of the image area through a series of cue points. Specifically, the cue points are evenly distributed on the image in the form of a matrix, forming a 6×6 rectangular grid structure. Each cue point acts as a local indicator to help the model capture areas in the image that may contain important information.
[0122] s2, the constructed prompt point matrix and the medical image are input into the Med-SAM model, prompting the model to identify the salient areas in the entire image and perform segmentation operations. When the model is processing, it analyzes the image content layer by layer through the internal segmentation algorithm to generate multiple segmentation regions. In order to optimize the quality of the segmentation results and the validity of the data, the system selects the top 5 segmentation regions with the highest confidence, which are regarded as the most informative parts of the image.
[0123] s3, based on the association question of the medical image, further screen the segmented regions. Among the first 5 high-confidence segmented regions obtained, identify the region most relevant to the question based on the semantic information provided by the medical question. Ensure that the generated mask data focuses on the image content required to answer the medical question.
[0124] s4, annotate the selected areas as "importance masks". This importance mask is not only the area of interest for the model, but also the core data used as labels in medical visual question answering tasks. After the annotation is completed, these importance masks are saved in the training set so that they can be used as label data for model learning during multi-objective training.
[0125] like Figure 4 As shown, the medical visual question-answering device proposed in the embodiment of the present invention includes: a user interaction module, a reasoning prediction module and a memory storage module.
[0126] The user interaction module is used to interact directly with users. Through this module, users can upload medical images, enter diagnostic questions or instructions, and view the answers and diagnostic results generated by the model, as well as mark important areas of the image for intuitive display to users. The main functions of the module include data uploading, instruction reception, and diagnostic feedback: receiving medical images and text question instructions uploaded by users, and passing them to subsequent modules after formatting; displaying the model's answers and analysis results to users in an intuitive way, including text replies, image labels, and mask displays to ensure that the diagnostic information is clear and easy to understand; in addition, the module can support multi-round interaction modes, allowing users to ask further questions after the initial diagnosis, and the system will generate continuous diagnostic answers based on historical context.
[0127] The inference prediction module takes the user's question instructions and medical image data as input, and generates the final diagnosis results and key area annotations through the medical visual question answering model. The main components of the module include medical image encoder, medical dialogue model, question text encoder, hierarchical image weighter, image and text feature fusion and multi-target predictor.
[0128] The medical image encoder extracts the features of medical images through the convolutional neural network structure, can identify the disease features in the image, filter out irrelevant information such as noise and artifacts, and thus generate a high-dimensional image feature representation; the medical dialogue model integrates knowledge in the medical field, and through fine-tuning of a large number of medical question-and-answer data sets, it has rich medical professional understanding capabilities and long text understanding capabilities. The model can integrate the user's historical questions and current instructions to generate answers that meet the medical context, support multiple rounds of question-and-answer and complex instruction parsing; the question text encoder performs semantic analysis and feature extraction on the question instructions input by the user, generates structured text features, and ensures that the model can understand the context and keywords of the question, thereby achieving accurate question-and-answer matching; level The image weighter is used to weight image features layer by layer so that the model can focus on key areas related to the current problem. The self-attention mechanism is used to weight important areas of the image and improve the accuracy of diagnosis. The image-text feature fusion device fuses image features with text features at multiple levels, ensures the correlation between the two modal features through a cascade network structure, achieves semantic alignment, and generates a more comprehensive information representation. The multi-target predictor realizes multi-task output, including answer prediction, image category prediction, and important area labeling. The answer prediction generates diagnostic suggestions that match the question. The image category predictor is used to determine the type of image part (such as chest, abdomen, brain, etc.). The important area predictor generates lesion masks to mark key diagnostic areas for users.
[0129] The memory storage module is used to save the medical images, interactive dialogue data and historical diagnosis reports uploaded by users. On the one hand, the stored data is convenient for users to review at any time. On the other hand, the results of historical diagnosis reports will be considered in the subsequent medical question and answer diagnosis, and the consistency of diagnosis will be maintained. Specifically, a patient history medical record library can be established to save images and question and answer records at different time points. The medical images uploaded by users and the result data generated by each diagnosis will be stored, including text records of multiple rounds of questions and answers, diagnosis conclusions, and lesion annotations, etc., to ensure the integrity and security of the data. It also allows users to view previous consultation records at any time, which is convenient for reviewing and tracking past diagnosis results. The memory module provides a data call interface so that historical data can be taken into consideration in subsequent diagnosis. Records such as historical question and answer data can be used as a reference for subsequent diagnosis, providing users with a coherent diagnostic experience. Historical data can be called in new questions and answers as input to the large diagnostic model to ensure the continuity of diagnosis. By introducing historical report results as context information, the model can give more accurate judgments and diagnostic suggestions based on the user's past conditions.
[0130] Figure 5 The figure is a schematic diagram of the overall structure of a medical visual question answering system based on multi-task modeling according to a specific embodiment of the present invention.
[0131] After the user uploads the medical image, the system extracts the multi-level features of the image, including the details of the lesion area and anatomical structure, through the visual encoder. At the same time, the user enters the diagnosis question in the chat interface, and the system uses the medical dialogue model to parse the semantics of the question, extract the core instructions, and then convert the question into a feature vector through the text encoder.
[0132] In the feature fusion layer, the system aligns and fuses text features with image features. First, the self-attention module weights the image features based on the text features, highlighting the image area related to the question and weakening the background interference. Then, after further fusion by the image-text alignment weighted module, a unified image-text representation is generated, modeling the deep interactive information between image and text.
[0133] In the output prediction layer, the fused representation is input into the multi-task prediction module, including the image decoder, image category predictor, and answer predictor. The image decoder is used to reconstruct and highlight the lesion area, the image category predictor determines the part category of the image, and the answer predictor generates diagnostic suggestions. Finally, the system optimizes the answer and displays it in the chat interface to provide users with accurate diagnosis results and medical advice.
[0134] The specific steps can also be:
[0135] S1. First, the MedSAM model is used to automatically segment the medical image dataset to obtain multiple image segments. Then, based on the questions raised by the user, the segments related to the questions are screened out, and the areas containing important information are retained to highlight the key lesion sites as labels for the training dataset.
[0136] S2, after the user uploads the medical image, the system will use the visual encoder to extract the image features and obtain the image features. The feature extraction layer extracts multi-level visual information from the original image, including the detailed features of the lesion area and various anatomical structure features.
[0137] S3: Users interact with the medical image question-answering system in natural language through the chat box and input diagnosis or inquiry questions. The system uses the medical dialogue model, combines user questions and historical chat records for semantic analysis, identifies the main command content of the current question, and extracts question commands suitable for image analysis.
[0138] S4, the system inputs the question instruction into the text encoder, extracts text features, and obtains the feature vector of the question.
[0139] S5, in the feature fusion layer, the extracted text features and image features are fused and aligned. The system's self-attention image importance weighting module first weights the image features based on the text features, highlighting the image areas that are more relevant to the problem, while weakening the background information that is not related to the diagnosis.
[0140] S6, the image features and text features after importance weighting are further deeply fused in the image-text alignment weighted module to generate a unified image-text fusion representation, which models high-dimensional visual and semantic fusion information.
[0141] S7, in the output prediction layer, the system inputs the image-text fusion representation into multiple different output layers to complete the task prediction. It includes an image decoder, an image category predictor, and an answer predictor, which respectively obtain important areas, image part categories, and answers to questions. The image decoder is used to reconstruct and highlight important areas to assist in lesion localization. The image category predictor ensures diagnostic consistency in different anatomical locations by predicting the part category of the image. The answer predictor generates text answers and gives explanations or diagnostic suggestions based on the question and image information. The multi-task prediction mechanism provides predictions from multiple dimensions.
[0142] S8, the diagnosis results generated by the answer predictor are further input into the medical dialogue model for semantic optimization, and finally a complete diagnosis opinion is output in the system chat box. The diagnosis opinion includes both accurate answers to user questions and suggestions with medical guidance value formed based on multiple rounds of question and answer analysis.
[0143] The user can conduct the next round of medical inquiry in the chat box and repeat the process of S1-S8.
[0144] like Figure 6 As shown, it is a comparison diagram of the medical image question-answering effect of the system of the present invention and the baseline model system that does not adopt multi-task learning. It can be seen from the figure that the method of the present invention significantly improves the model's recognition accuracy of fine lesion features in medical images, and can help the model better capture the features of important areas in the image. From the results in the heat map, it is clearly observed that, thanks to the fusion model and multi-task training paradigm, the system of the present invention can focus more on the key areas in the image, and the activation output is concentrated in the areas closely related to the problem. In contrast, the activation pattern of the baseline model that does not adopt multi-task learning is more scattered, and it is difficult to effectively highlight the areas related to the problem. This shows that the method of the present invention significantly improves the model's ability to focus on problem-related areas.
[0145] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0146] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0147] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0148] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0149] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention are included in the protection scope of the present invention.
Claims
1. A medical visual question answering method based on multi-task modeling, characterized in that: include: Step s1: Load the user's historical question and answer data, and simultaneously obtain the medical image to be analyzed and the preliminary question instructions of the current user; Step s2: extracting features from the medical image to be analyzed and the preliminary question instruction, that is, extracting features from the medical image using a visual encoder to obtain image features, inputting the historical question and answer data and the preliminary question instruction into the dialogue model for semantic analysis, identifying and outputting question instructions suitable for image analysis, and extracting question features through a text encoder to obtain text features; Step s3: The obtained text features and image features are processed by self-attention image importance weighting to highlight the image area with stronger relevance to the question instruction, and the weighted image features are obtained. After the image and text are aligned, the image and text fusion representation is obtained; Step s4: Input the acquired image-text fusion representation into the multi-target output projection layer for multi-task prediction. The prediction output includes question answers, image categories, and important area masks. The question answers and image categories are input into the large dialogue model, and the dialogue context and multi-round interaction data are combined with the important area masks to finally generate a diagnostic opinion with detailed descriptions; the multi-target output projection layer includes an answer prediction model, an image category prediction model, and an important difference predictor; the answer prediction model is used to generate answer categories related to user questions, and is composed of multiple layers of fully connected layers. By projecting the input features layer by layer, the final output is a probability distribution vector of the answer category; the task of the image category prediction model is to classify the anatomical parts of the medical image, and is composed of multiple layers of fully connected layers, and the final output is a probability distribution vector of the image category; the important area predictor is used to locate the key areas related to diagnosis in the medical image, and adopts a decoder network structure, which is composed of multiple convolutional layers Conv2D and upsampling layers UpSampling2D, and achieves accurate area positioning by gradually restoring the spatial resolution.
2. The medical visual question answering method based on multi-task modeling according to claim 1, characterized in that: In step s2, the visual encoder adopts the ResNet architecture and is pre-trained on medical imaging data to enable it to have the ability to extract medical image features. It uses the convolutional neural network CNN layer to hierarchically encode image features. Each convolutional layer can gradually extract features of different levels, and finally form high-dimensional image features. ; The text encoder uses an LSTM network to embed the question instructions and input them into the LSTM network, gradually extracting the keywords and semantic association information in the question instructions, and generating text features with time sequence information. .
3. The medical visual question answering method based on multi-task modeling according to claim 2, characterized in that: The self-attention image importance weighting module comprises a mapping layer, a self-attention layer and a feedforward output layer connected in sequence, wherein: The mapping layer transforms text features and image features Projection to high-dimensional space, projecting the original input to a unified feature dimension through matrix mapping; The self-attention layer calculates the similarity scores of the mapped text features and image features and determines the important areas in the image features, where the text features and image features Similarity score The calculation formula is as follows: , in, and are the weight matrices of text features and image features respectively, represents the scaling factor of the feature dimension; the attention score is normalized by the softmax function so that the sum of the importance weights of each region is 1; the weighted image feature is expressed as: , Weighted image features After being processed by the feedforward output layer, the final weighted features are output. The feedforward output layer contains multiple nonlinear activation layers and regularization layers. The output weighted image features for: , in, and are the weight matrix and bias term of the feedforward output layer, respectively. represents a non-linear activation function.
4. The medical visual question answering method based on multi-task modeling according to claim 3, characterized in that: The image and text alignment described in step s3 is specifically as follows: Weighted image features and text features Input cascade image-text alignment weighted module, which includes feature cascade layer, cross-modal attention fusion layer, and hierarchical feedforward network; The feature cascade layer is used to cascade the weighted image features with the text features to obtain the cascaded feature representation. ; The cross-modal attention fusion layer generates a cross-modal weight matrix through the self-attention mechanism : , in, and is the projection weight matrix, and the output cross-modal weighted features are , in obtaining cross-modal weighted features Afterwards, the cascaded image-text alignment weighted module inputs it into the hierarchical feedforward network; The hierarchical feedforward network consists of multiple layers of fully connected layers and nonlinear activation layers. Its structural design aims to refine cross-modal features layer by layer and generate a more information-dense fusion representation. The calculation formula of the hierarchical feedforward network is: ,in and are the weight matrix and bias term of the feedforward network, and the final output of the hierarchical feedforward network After normalization, a unified image-text fusion representation is generated , the representation integrates deep image and text interaction information.
5. The medical visual question answering method based on multi-task modeling according to claim 1, characterized in that: When the multi-target output projection layer performs multi-target training, the data label training data of the importance mask is constructed as follows: 1) Construct an indicator point matrix, where each indicator point acts as a local indicator to help the medical image segmentation model capture the area containing important information in the image; 2) Input the constructed cue point matrix and the medical image into the Med-SAM model, prompting the model to identify the salient areas in the entire image and perform segmentation operations, generate multiple segmentation regions, and select the top N segmentation regions with the highest confidence. These regions are regarded as the most informative parts of the image; 3) Based on the association problem of medical images, the segmented regions are further screened; among the first N high-confidence segmented regions obtained, the regions most relevant to the question are manually screened and identified based on the semantic information provided by the medical question, ensuring that the generated mask data focuses on the image content required to answer the medical question; 4) The selected areas are marked as "importance masks". The importance masks are not only the areas of interest to the model, but also the core data used as labels in medical visual question answering tasks; After labeling, these importance masks are saved in the training set so that they can be used as label data for model learning during multi-objective training.
6. A medical visual question answering system based on multi-task modeling, characterized in that: For implementing the medical visual question answering method based on multi-task modeling as described in any one of claims 1 to 5, the system comprises: A user interaction module, used to interact with the user, the interaction includes receiving medical images uploaded by the user, preliminary question instructions input, and displaying answers and diagnosis results generated by the system, including text replies, image labels and mask displays, and supporting multi-round interaction modes; The inference prediction module is used to process the received medical images and preliminary question instructions, including a medical image encoder, a medical dialogue model, a question text encoder, a hierarchical image weighter, a picture-text feature fusion device and a multi-target predictor, wherein the medical image encoder is used to extract the features of the medical image and generate image features; the medical dialogue model integrates the knowledge in the medical field and is used to integrate the user's historical question and answer data and the current preliminary question instructions to generate question instructions suitable for image analysis; the question text encoder is used to perform semantic analysis and feature extraction on the question instructions to generate text features; the hierarchical image weighter is used to weight the image features layer by layer so that the model can focus on the key areas related to the current question, and the weighting of the important areas of the image is achieved through the self-attention mechanism; the picture-text feature fusion device fuses the image features and the text features at multiple levels, ensures the correlation between the two modal features through the cascade network structure, and achieves semantic alignment; the multi-target predictor realizes multi-task output, including answer prediction, image category prediction and important area labeling, the answer prediction generates a diagnostic suggestion matching the question, the image category predictor is used to determine the type of image part, and the important area predictor generates an important area mask to label the key diagnostic area for the user; The memory storage module is used to save medical images, interactive dialogue data and historical diagnosis reports uploaded by users, and allows viewing of historical question and answer data.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the medical visual question answering method based on multi-task modeling as described in any one of claims 1 to 5 is implemented.
8. An electronic device, characterized in that: The device comprises: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the medical visual question answering method based on multi-task modeling as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Pollutant high-precision target detection method and system based on river information guidance
CN118429622A
Medical visual question and answer method, device and equipment and storage medium
CN118467707A