Multi-modal medical AI auxiliary inquiry method fusing medical image and medical text
Through the combination of multimodal data augmentation and Combine-Former multimodal fusion device, the challenge of chat Q&A model in handling medical image and text data associations in the medical field is solved, achieving more efficient multimodal data understanding and more comprehensive medical-assisted consultation support.
Patent Information
- Application Number
- CN202510048533.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-06-13
AI Technical Summary
The difficulty in effectively handling the correlation between medical images and text data in the application of existing chat Q&A models in the healthcare field has led to challenges in understanding the differences and correlations of multimodal data.
Multimodal data augmentation technology is used to diversify medical data, and the feature fusion of medical images and medical text data is realized through Combine-Former multimodal fusion device, so that vision-language models can better understand and process multimodal data.
Through multimodal data fusion technology, AI-assisted consultation models can comprehensively consider patient text problems and medical images, thereby providing more comprehensive medical diagnosis and consulting support, improving the applicability and robustness of the model.
Smart Images

Figure CN120148909A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - technical fields of natural language processing, computer vision, multimodal fusion, large model fine - tuning, chat - based question - answering models, and AI - assisted medical consultation. Specifically, it relates to a multimodal medical AI - assisted medical consultation method that integrates medical images and medical texts. Background Art
[0002] A chat - based question - answering model is an intelligent conversation computer program that mimics the form of human natural conversation, which can process user input and generate an output. Generally, a chat - based question - answering model takes natural language text as input, and the output is the output most relevant to the user - input sentence. A chat - based question - answering model can also be defined as "an online human - machine dialogue system with natural language". Therefore, a chat - based question - answering model constitutes an automatic dialogue system that can talk to thousands of potential users simultaneously. In recent years, with the improvement of computing power, as well as the sharing of open - source technologies and frameworks, chat - based question - answering model programs have become increasingly common. The latest developments in artificial intelligence and natural language processing technologies have made chat - based question - answering models easier to implement, more flexible in terms of application and maintainability, and increasingly capable of mimicking human conversations. Chat - based question - answering models are currently applied in a variety of different fields and applications, ranging from education to e - commerce, covering healthcare.
[0003] Although chat - based question - answering models have made significant progress in the past few years, there are still many challenges in their practical applications in the healthcare field. On the one hand, there are challenges brought about by the diversification of medical data due to perspective changes. Doctors not only inquire about patients through text information, but patients' medical image information plays a key role in doctors' diagnosis. However, chat - based question - answering models can only receive text data and cannot understand medical image data. On the other hand, there are challenges brought about by multimodal data fusion. When the model understands image and text data, how the model can extract features related to the correlation between images and text. Simple merging of feature vectors will lead to a substantial increase in the computing resources required by the model, thereby reducing the efficiency of model training. Moreover, there are significant differences between medical images. For example, the difference between a fluoroscopy image with a lesion and a healthy fluoroscopy image may only be reflected in a local area of the image, and the difference between fluoroscopy images with different lesions may only be shown in more subtle areas. These factors make it challenging for the model to accurately understand the differences and correlations of multimodal data.
[0004] Currently, some pre-trained vision-language models have proposed BERT-like architectures to process cross-modal inputs, adopted contrastive learning to train vision-language models for medical images, used a large number of image-text pairs to train vision-language models, and performed end-to-end pre-training on large-scale image-text pair datasets. However, as the model scale increases, the computational cost of pre-training becomes extremely high. In addition, for end-to-end pre-trained models, it is not flexible to utilize existing single-modal pre-trained models such as BERT. Some researchers have also achieved target perception by jointly training image CNNs and text transformers. This method uses natural language prompts for pre-training, reducing the need for a large amount of labeled data. Through this method, the model can learn multiple tasks and demonstrate good performance in different downstream tasks. However, this method still has some limitations. It consumes a large amount of computational resources during joint training, and a large amount of medical imaging data and long text data will significantly reduce the model's performance. Summary of the Invention
[0005] The purpose of the present invention is to address the defects and deficiencies of the above-mentioned existing technologies, and provide a multi-modal medical AI-assisted consultation method that integrates medical images and medical texts. This method uses multi-modal data augmentation technology to diversify the existing medical data, fine-tunes the vision-language large model with the diversified medical data, and realizes feature fusion of image-text data during the fine-tuning process with the help of multi-modal fusion technology, enabling the vision-language model to have a full understanding of medical images and doctor-patient Q&A data.
[0006] The technical solution adopted by the present invention to solve its technical problems is: a multi-modal medical AI-assisted consultation method that integrates medical images and medical texts, comprising the following steps:
[0007] Step 1: Use the Medical-Diff-VQA dataset, and perform data augmentation on it using a negative sample generation enhancement algorithm. Divide the augmented dataset into a training set, a validation set, and a test set;
[0008] Step 2: For the chest fluoroscopy images in the dataset of Step 1, use the Vit algorithm model to divide the fluoroscopy images into blocks, add sequence encoding, pass through a 16-head self-attention encoding layer, and then process them into image feature vectors through a deep feed-forward network;
[0009] Step 3: Extract the conversation history in the same medical conversation in the dataset of Step 1, combine it into a conversation list, tokenize each group of conversation lists and the current patient's question, and use the tokenized text as the input of the generative bidirectional multi-head self-attention algorithm model to output text encoding vectors;
[0010] Step 4: Generate a learnable intermediate vector and input it, together with the image feature vector obtained in Step 2, into the Combine-Former multi-modal fusion device;
[0011] Step 5: The Combine-Former multi-modal fusion device first realizes the interactive fusion of the text encoding vector and the generated learnable intermediate vector through a self-attention layer with shared weights to determine the focus of each token when processing the image, and outputs a text feature vector.
[0012] Step 6: Input the text feature vector output in the above Step 5 into a cross-channel attention layer with randomly initialized weights to interact with the image features, fuse the image information and the text information, and output a fused feature vector;
[0013] Step 7: Use the set of fused feature vectors output in the above Step 6 as the input of the ChatGLM-6b algorithm model, and then fine-tune the algorithm model using a medical professional domain adapter to make it have more professional and accurate question-and-answer capabilities.
[0014] Furthermore, Step 1 of the present invention includes:
[0015] Step 11: Traverse the Medical-Diff-VQA dataset in a loop. The dataset includes question types, question contents, doctor answers, and other contents;
[0016] Step 12: For each row of data, extract the question type and the question content, generate new questions and answers according to different question types. For abnormal questions, modify the questions and answers to generate a situation opposite to the original question. Specifically, convert positive questions into negative questions. For existential questions, modify the questions and answers to generate the opposite situation and change the answer from affirmative to negative. For degree questions, extract keywords and convert the questions into degree-related questions. At the same time, according to the different answers, change the answers to affirmative or negative and append the original answer content.
[0017] Step 13: Merge it with the original dataset, and the data volume increases by about 20%.
[0018] Furthermore, Step 2 of the present invention includes:
[0019] Step 21: For the input chest fluoroscopy image I∈R H×W×D , where D represents the number of channels of the fluoroscopy image, R represents the set of chest fluoroscopy images, and H and W respectively represent the height and width of each fluoroscopy image;
[0020] Step 22: Divide the fluoroscopy image into fluoroscopy image blocks with a resolution of 14×14 where (p, p) represents the resolution of each fluoroscopy image block, C represents the number of channels, and N represents that one fluoroscopy image is divided into N blocks;
[0021] Step 23, after passing through the linear projection layer, the dimension of the perspective view block sequence is (N, p 2 , C), and then generate N rows of sequence encodings, which are added to the perspective view block sequence;
[0022] Step 24, use the 16-head self-attention mechanism to extract the feature vectors of the perspective view block sequence, and then obtain the required chest fluoroscopy feature vectors through a deep feed-forward network whose dimension is still (N, p 2 , C).
[0023] Furthermore, step 3 of the present invention includes:
[0024] Step 31, for the medical dialogue dataset D, combine the conversations C i,j of the same medical item into a dialogue history H i , that is, H i ={C i,1 , C i,2 ,.., C i,n}, where i is the dialogue history index and j is the sentence index.
[0025] Step 32, use the tokenizer of the generative unidirectional multi-head self-attention model to tokenize the sentences of the conversation C i,j to obtain the dialogue history word sequence, where w i,j,1 to w i,j,m is the set of tokenized words, and m is the word index:
[0026] Tokens i ={w i,j,1 , w i,j,2 ,...w i,j,m}
[0027] Step 33, use the word encoder based on the multi-head self-attention mechanism to encode these words to obtain the vector representation of the dialogue history:
[0028] HE i =TEB(Tokenns i )
[0029] where TEB represents the word encoding method, and HE i is the encoding representation of the entire dialogue history H i , and i is the dialogue history vector index.
[0030] Step 34, perform the same tokenization and encoding operations on the patient question Q i to obtain QE i .
[0031] Furthermore, step 5 of the present invention includes:
[0032] Step 51, first, the patient's question QE ∈ R l×d and the dialogue history HE ∈ R m×d are concatenated:
[0033] E concat = Concat(HE, QE)
[0034] where l is the length of the QE sequence, d is the dimension of the word embedding, Concat is the addition function, m is the length of the dialogue history sequence, and the generated intermediate vector is E concat .
[0035] Step 52, using the self-attention mechanism, fuse the text features of the patient's question and the dialogue history, and calculate the attention weight matrix:
[0036]
[0037] where W q is the weight matrix, I is the intermediate value of the attention calculation, A is the attention weight, i and j are the indices of I, and T is the matrix transpose symbol, which is used to convert E concat into the query matrix, indicating the attention degree of the model to the input sequence.
[0038] Step 53, multiply the attention weight at each position by the corresponding intermediate vector at that position, and sum the results of all positions to obtain the final context representation.
[0039] C = A · E concat
[0040] where C is the context representation after the attention mechanism. To provide a more informative input to the model, C and E concat are concatenated:
[0041] E interact = Concat(E concat , C)
[0042] Step 54, to increase the model's expression ability and adaptability to medical domain data, introduce a learnable linear transformation W intercat ∈ R( l+m)×(l+m) and a noise bias b intercat ∈ R (l+m)×1 act on E interact from step 53, obtain E final , and then perform activation. tanh is the activation function, and the final feature E activated is obtained:
[0043]
[0044] This interaction process describes in detail the complex relationship between the patient's problems, the doctor-patient dialogue history, and the generated intermediate vectors. At the same time, by introducing learnable parameters, the expression ability and adaptability of the model are increased.
[0045] Furthermore, step 6 of the present invention includes:
[0046] Step 61, taking the final feature E of the text information in step 54 activated ∈R (l+m)×(l+m) and the image feature vector IE ∈ R p×q , where p is the length of the image feature vector and q is the dimension of the image feature, and softmax is the activation function, and using them as the input to the cross-channel attention layer CA:
[0047]
[0048] where, W q , W k , W v are learnable weight matrices.
[0049] Step 62, concatenating the final feature of the text information with the output of the cross-channel attention layer to retain the original text information and the information interacting with the image information:
[0050] E cross = Concat(E activated , CA(E activated , IE))
[0051] Step 63, applying a linear transformation and the GELU activation function to obtain the fused feature vector:
[0052]
[0053] where, W fusion is the learned weight matrix and b fusion is the bias term.
[0054] Furthermore, step 7 of the present invention includes:
[0055] Step 71, using the fused feature vector obtained in step 63 as the input of the ChatGLM-6B language model to make it output the result based on general language understanding, where ChatGLM is the ChatGLM-6B language model:
[0056] O orig = ChatGLM(E fusion )
[0057] Step 72, to enable ChatGLM-6B to have the ability to answer questions in the medical professional field, we introduce a medical professional field adapter. An adapter is a small neural network that is added to the output of the original model to adjust the model to adapt to new tasks.
[0058] O med =O orig +Adapter med (GELU(W adapter O orig +b adapter ))
[0059] where Adapter med is the medical professional field adapter, W adapter is the adapter weight, b adapter is the adapter bias, and this adapter can include a linear transformation and a non-linear GELU activation function.
[0060] Step 73, fine-tune the adapter in Step 72. The goal of fine-tuning is to minimize the loss function L med of the new task, which focuses on the question-answering task in the medical professional field.
[0061]
[0062] where Y is the true label, log is the logarithmic function, O med is the prediction of the model on the fine-tuning task, and i is the label index. This objective function is used to measure the performance of the model on the medical field question-answering task. After that, the overall fine-tuning objective function is defined as the weighted sum of the pre-training task loss and the fine-tuning task loss:
[0063] L totnl =λL pretrain +(1 - λ)L med
[0064] where λ is a hyperparameter that balances the pre-training task and the fine-tuning task, and L pretrain is the loss of the ChatGLM-6B model on the pre-training task. Finally, use gradient descent to update the parameters W adapter and b adapter of the adapter:
[0065]
[0066] where μ is the learning rate.
[0067] Advantages of the present invention: The present invention introduces multi-modal data fusion technology, which mutually fuses medical images and medical text data, enabling the AI-assisted consultation model to comprehensively consider the text questions and medical images raised by patients, thereby providing more comprehensive support for medical diagnosis and consultation. At the same time, by adopting multi-modal data augmentation technology, the present invention realizes diversified processing of medical data, which not only improves the applicability of the model, but also increases the robustness of the model, enabling it to better cope with different patient situations and the diversity of medical data. In addition, by fine-tuning large-scale vision-language models, the present invention makes the model more suitable for the consultation tasks in the medical field, thereby improving the performance and accuracy of the model. Most importantly, the present invention also allows processing of the conversation history to provide more personalized and coherent Q&A support, which helps to improve the quality and efficiency of the communication between doctors and patients, thereby improving the interaction experience between patients and doctors.
[0068] Specifically, it includes:
[0069] 1. The multi-modal data augmentation technology of the present invention constructs a negative sample generation algorithm. Before model training, negative answer samples are generated on the basis of the original data set.
[0070] 2. The present invention fine-tunes the Vit image encoding model through a medical data set with negative samples, enabling the model to have a strong understanding ability of medical images, and then generating chest fluoroscopy feature vectors.
[0071] 3. The present invention uses the Combine-Former multi-modal data fusion technology to fuse chest fluoroscopy images and doctor-patient Q&A information in the medical field. The Q&A text features and conversation record features are fused through a generative bidirectional multi-head self-attention algorithm language model, and then the Q&A text features and chest fluoroscopy image features are fused through a cross-attention layer to obtain fused feature vectors.
[0072] 4. The present invention fine-tunes the GhatGLM model using medical image Q&A data, enabling it to generate medical diagnoses through fused feature vectors. Description of the Drawings
[0073] Figure 1 is a flowchart of the AI-assisted consultation method of the present invention.
[0074] Figure 2 is an architecture diagram of the multi-modal medical AI-assisted consultation model that fuses medical images and medical texts of the present invention. Detailed Embodiments
[0075] The embodiments of the present invention will be disclosed below with reference to the drawings. For the sake of clarity, many practical details will be described together in the following narrative. However, it should be understood that these practical details are not used to limit the present invention. That is to say, in some embodiments of the present invention, these practical details are unnecessary.
[0076] When addressing the challenges of chatbot models in the field of healthcare, the present invention first takes into account the complexity of multimodal data processing. Since medical information encompasses both text and image data, traditional chatbot models struggle to effectively handle the correlation between the two. To overcome this issue, we adopted an innovative approach by enhancing the generation of negative samples in the Medical-Diff-VQA dataset. By processing chest fluoroscopy images through the Vit algorithm model, we obtained more informative image feature vectors. Another key problem was how to enable the model to better understand and integrate text and image information. For this purpose, we introduced the Combine-Former multimodal fusion module, which achieved the interaction between the text encoding vector and the generated intermediate vector through the self-attention layer, and further fused the image information with the text information through the cross-channel attention layer. This multi-level fusion strategy aims to improve the model's comprehensive understanding of multimodal data, enabling it to respond to patient questions more accurately. To evaluate our method, we introduced the ChatGLM-6b algorithm model into the medical professional domain adapter for fine-tuning. Such a professional domain adapter not only enhances the model's professionalism in medical Q&A but also enables it to perform better in different medical scenarios. Through these innovative methods, we have successfully addressed the challenges of multimodal data fusion, enabling our model to achieve higher accuracy and professionalism in medical AI-assisted consultation.
[0077] As Figure 1 、 Figure 2 shown, the present invention provides a multimodal medical AI-assisted consultation method that integrates medical images and medical texts, including the following steps:
[0078] Step 1: Use the Medical-Diff-VQA dataset and apply the negative sample generation enhancement algorithm to enhance the data, and divide the enhanced dataset into a training set, a validation set, and a test set;
[0079] The negative sample generation enhancement algorithm includes the following steps:
[0080] Step 11: Traverse the Medical-Diff-VQA dataset in a loop. The dataset includes question types, question contents, doctor's answers, etc.;
[0081] Step 12: For each line of data, extract the question type and question content, generate new questions and answers according to different question types. For abnormal questions, modify the questions and answers to generate a situation opposite to the original question. Specifically, convert positive questions into negative questions. For existential questions, modify the questions and answers to generate the opposite situation and change the answer from affirmative to negative. For degree questions, extract the keywords and convert the questions into degree-related questions. At the same time, according to the different answers, change the answers to affirmative or negative and append the original answer content.
[0082] Step 13: Merge it with the original dataset, and the data volume increases by about 20%.
[0083] Step 2: For the chest fluoroscopy images in the dataset of Step 1, use the Vit algorithm model to divide the fluoroscopy images into blocks, add sequence encoding, pass through the 16-head self-attention encoding layer, and then process them into image feature vectors through a deep feed-forward network;
[0084] Among them, the Vit algorithm model encodes the chest fluoroscopy images into image feature vectors, including the following steps:
[0085] Step 21: For the input chest fluoroscopy image I∈R H×W×D , where D represents the number of channels of the fluoroscopy image, R represents the set of chest fluoroscopy images, and H and W represent the height and width of each fluoroscopy image respectively;
[0086] Step 22: Divide the fluoroscopy image into fluoroscopy image blocks with a resolution of 14×14 Among them, (p, p) represents the resolution of each fluoroscopy image block, C represents the number of channels, and N represents that a fluoroscopy image is divided into N blocks;
[0087] Step 23: After passing through the linear projection layer, the dimension of the fluoroscopy image block sequence is (N, p 2 , C), and then generate N rows of sequence encoding, which are added to the fluoroscopy image block sequence;
[0088] Step 24: Use the 16-head self-attention mechanism to extract the fluoroscopy image block sequence feature vectors, and then obtain the required chest fluoroscopy image feature vectors through a deep feed-forward network Its dimension is still (N, p 2 , C).
[0089] Step 3: Combine the conversation histories in the same medical conversation in the dataset into a conversation list, tokenize each group of conversation lists and the current patient's question, and use the tokenized text as the input of the generative bidirectional multi-head self-attention algorithm model to output text encoding vectors;
[0090] Specifically, the process of the text encoder based on the generative bidirectional multi-head self-attention algorithm to extract text encoding vectors includes the following steps:
[0091] Step 31, medical conversation dataset D, collects conversations C of the same medical project i,j Combined into a conversation history H i , both H i ={C i,1 , C i,2 , ..., C i,n}, where i is the conversation history index and j is the sentence index.
[0092] Step 32: Use the word segmenter of the generative unidirectional multi-head self-attention model to segment the dialogue C i,j Sentence segmentation, get the dialogue history word sequence, where w i,j,1 to w i,j,m is the word set after segmentation, and m is the word index:
[0093] Tokens i ={w i,j,1 , w i,j,2 , ...w i,j,m}
[0094] Step 33, use a text encoder based on a multi-head self-attention mechanism to encode these words and obtain a vector representation of the conversation history:
[0095] HE i =TEB(Tokens i )
[0096] Among them, TEB represents the text encoding method, HE i is the entire conversation history H i The encoding representation of , i is the dialogue history vector index.
[0097] Step 34, put the patient's question Q i Perform the same word segmentation and encoding operations to obtain QE i .
[0098] Step 4: Generate a learnable intermediate vector and input it into the Combine-Former multi-modal fusion device together with the image feature vector obtained in step 2);
[0099] In step 5, the Combine-Former multimodal fuser first realizes the interactive fusion of the text encoding vector and the generated learnable intermediate vector through a self-attention layer with shared weights to determine the focus of each token when processing the image and outputs the text feature vector.
[0100] The self-interaction method of the text encoding vector in the Combine-Former multimodal fuser includes the following steps:
[0101] Step 51: First, take the patient's question \(QE\in R\) l×d and the dialogue history \(HE\in R\) m×d and splice them together:
[0102] E concat = Concat(HE, QE)
[0103] where \(l\) is the length of the \(QE\) sequence, \(d\) is the dimension of the word embedding, Concat is the addition function, \(m\) is the length of the dialogue history sequence, and the generated intermediate vector is \(E\) concat .
[0104] Step 52: Use the self-attention mechanism to fuse the text features of the patient's question and the dialogue history, and calculate the attention weight matrix:
[0105]
[0106] where \(W\) q is the weight matrix, \(I\) is the intermediate value of the attention calculation, \(A\) is the attention weight, \(i\) and \(j\) are the indices of \(l\), and \(T\) is the matrix transpose symbol, which is used to convert \(E\) concat into a query matrix, representing the model's attention to the input sequence.
[0107] Step 53: Multiply the attention weight at each position by the corresponding intermediate vector and sum the results of all positions to obtain the final context representation.
[0108] C = A·E concat
[0109] where \(C\) is the context representation after the attention mechanism. To provide a more informative input to the model, \(C\) and \(E\) concat are spliced together:
[0110] E interact = Concat(E concat , C)
[0111] Step 54: To increase the model's expressive ability and adaptability to medical domain data, introduce a learnable linear transformation \(W\) intercat \(\in R\) (l+m)×(l+m) and a noise bias \(b\) intercat \(\in R\) (l+m)×1 and apply them to \(E\) in Step 53 interact to obtain \(E\) final , and then perform activation. tanh is the activation function, and the final feature \(E\) activated is obtained:
[0112]
[0113] This interaction process describes in detail the complex relationship between the patient's problems, the doctor-patient dialogue history, and the generated intermediate vectors. At the same time, by introducing learnable parameters, the expression ability and adaptability of the model are increased.
[0114] Step 6: Input the text feature vector output in Step 5 into the cross-channel attention layer with randomly initialized weights to interact with the image features, fuse the image information and the text information, and output the fused feature vector.
[0115] The specific method for fusing text information and image information includes the following steps:
[0116] Step 61: Take the final feature E of the text information in Step 54 activated ∈R (l+m)×(l+m) and the image feature vector IE ∈ R p×q , where p is the length of the image feature vector and q is the dimension of the image feature. Softmax is the activation function, and use them as the input of the cross-channel attention layer CA:
[0117]
[0118] where, W q , W k , W v are learnable weight matrices.
[0119] Step 62: Concatenate the final feature of the text information with the output of the cross-channel attention layer, retaining the original text information and the information interacting with the image information:
[0120] E cross = Concat(E activated , CA(E activated , IE))
[0121] Step 63: Apply a linear transformation and the GELU activation function to obtain the fused feature vector:
[0122]
[0123] where, W fusion is the learned weight matrix, and b fusion is the bias term.
[0124] Step 7: Take the set of fused feature vectors output in Step 6 as the input of the ChatGLM-6b algorithm model, and then use the medical professional domain adapter to fine-tune the algorithm model to make it have more professional and accurate question-answering capabilities.
[0125] The method of using the medical professional domain adapter to fine-tune the algorithm model includes the following steps:
[0126] Step 71: Use the fused feature vector obtained in Step 63 as the input of the ChatGLM-6B language model to make it output results based on general language understanding, where ChatGLM is the ChatGLM-6B language model:
[0127] O orig = ChatGLM(E fusion )
[0128] Step 72: To enable ChatGLM-6B to have the ability to answer questions in the medical professional field, we introduce a medical professional field adapter. The adapter is a small neural network that will be added to the output of the original model to adjust the model to adapt to the new task.
[0129] O med = O orig + Adapter med (GELU(W adapter O orig + b adapter ))
[0130] where Adapter med is the medical professional field adapter, W adapter is the adapter weight, b adapter is the adapter bias, and this adapter can include a linear transformation and a non-linear GELU activation function.
[0131] Step 73: Fine-tune the adapter in Step 72. The goal of fine-tuning is to minimize the loss function L med of the new task, and this loss function focuses on the question-answering task in the medical professional field.
[0132]
[0133] where Y is the true label, log is the logarithmic function, O med is the prediction of the model on the fine-tuning task, and i is the label index. This objective function is used to measure the performance of the model on the medical field question-answering task. Then, the overall fine-tuning objective function is defined as the weighted sum of the pre-training task loss and the fine-tuning task loss:
[0134] L total = λL pretrain + (1 - λ)L med
[0135] where λ is a hyperparameter that balances the pre-training task and the fine-tuning task, and L pretrain is the loss of the ChatGLM-6B model on the pre-training task. Finally, use gradient descent to update the parameters W adapter and b adapter of the adapter:
[0136]
[0137] Here, μ is the learning rate.
[0138] This paper adopts an innovative multimodal data processing strategy to address the challenges of chat question-answering models in the healthcare field. By enhancing the medical visual language question-answering dataset and applying the Vit algorithm model, the complexity of multimodal data processing is solved. The introduction of the Combine-Former multimodal fuser enables effective interaction between text and image information, and the professionalism of the model is improved through fine-tuning of professional field adapters. This integration method has achieved higher accuracy and comprehensive performance in medical AI-assisted consultation.
[0139] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A multimodal medical AI-assisted consultation method integrating medical images and medical texts, characterized by: The method comprises the following steps: Step 1: Use the Medical-Diff-VQA dataset and use the negative sample generation enhancement algorithm to enhance its data. The enhanced dataset is divided into a training set, a validation set, and a training set. Step 2: Use the Vit algorithm model to divide the chest fluorescence perspective image in the data set of step 1 into blocks, add sequence encoding, pass through 16 self-attention encoding layers, and then process it into an image feature vector through a deep feedforward network; Step 3: extract the conversation history of the same medical conversation in the data set of step 1 above, combine them into a conversation list, segment each group of conversation lists and the current patient question, use the segmented text as the input of the generative bidirectional multi-head self-attention algorithm model, and output the text encoding vector; Step 4: Generate a learnable intermediate vector and input it into the Combine-Former multimodal fusion device together with the image feature vector obtained in step 2 above; Step 5: Combine-Former multimodal fuser realizes interactive fusion of text encoding vector and generated learnable intermediate vector through self-attention layer with shared weights to determine the focus of each token when processing the image and outputs text feature vector; Step 6, input the text feature vector outputted from step 5 above into the cross-channel attention layer with randomly initialized weights to interact with the image features, fuse the image information with the text information, and output the fused feature vector; Step 7: Use the fused feature vector set output by step 6 above as the input of the ChatGLM-6b algorithm model, and then use the medical professional field adapter to fine-tune the algorithm model to make it more professional and accurate in question-answering capabilities.
2. According to claim 1, a multimodal medical AI-assisted consultation method integrating medical images and medical texts is characterized by: The negative sample generation enhancement algorithm of step 1 comprises the following steps: Step 11, loop through the Medical-Diff-VQA dataset, which includes question type, question content, and doctor's answer content; Step 12: for each row of data, extract the question type and question content, generate new questions and answers according to different question types, modify the questions and answers for abnormal questions to generate the opposite situation of the original questions, specifically, convert positive questions into negative questions, modify the questions and answers for existence questions to generate the opposite situation, and change the answers from positive to negative, and extract keywords for degree questions and convert the questions into degree-related questions, and change the answers to positive or negative according to the different answers, and attach the original answer content; Step 13, merging with the original data set, the data volume increases by about 20%.
3. According to claim 1, a multimodal medical AI-assisted consultation method integrating medical images and medical texts is characterized by: The Vit algorithm model described in step 2 encodes the chest fluorescence perspective image into an image feature vector, including the following steps: Step 21, input chest fluorescence perspective image I∈R H×W×D , where D represents the number of channels of the perspective image, R represents the set of chest fluorescence perspective images, and H and W represent the height and width of each perspective image, respectively; Step 22: Divide the perspective image into 14×14 resolution perspective blocks Where (p, p) represents the resolution of each perspective block, C represents the number of channels, and N represents that a perspective image is divided into N blocks; Step 23, after passing the linear projection layer, the perspective block sequence dimension is (N, p 2 , C), regenerate N lines of sequence code and add them to the perspective block sequence; Step 24: Use the 16-head self-attention mechanism to extract the perspective block sequence feature vector, and then use the deep feedforward network to obtain the required chest fluorescence perspective feature vector Its dimension is still (N, p 2 , C).
4. According to claim 1, a multimodal medical AI-assisted consultation method integrating medical images and medical texts is characterized by: The process of extracting the text encoding vector by the text encoder based on the generative bidirectional multi-head self-attention algorithm described in step 3 includes the following steps: Step 31, medical conversation dataset D, collects conversations C of the same medical project i,j Combined into a conversation history H i , both H i ={C i,1 , C i,2 , ..., C i,n }, where i is the conversation history index and j is the sentence index; Step 32: Use the word segmenter of the generative unidirectional multi-head self-attention model to segment the dialogue C i,j Sentence segmentation, get the dialogue history word sequence, where w i,j,1 to w i,j,m is the word set after segmentation, and m is the word index: Tokens i ={w i,j,1 ,w i,j,2 ,…w i,j,m } Step 33, use a text encoder based on a multi-head self-attention mechanism to encode these words and obtain a vector representation of the conversation history: HE i =TEB(Tokens i ) Among them, TEB represents the text encoding method, HE i is the entire conversation history H i The encoding representation of , i is the dialogue history vector index; Step 34, put the patient's question Q i Perform the same word segmentation and encoding operations to obtain QE i .
5. According to claim 1, a multimodal medical AI-assisted consultation method integrating medical images and medical texts is characterized by: The self-interaction method of the text encoding vector in the Combine-Former multimodal fuser described in step 5 includes the following steps: Step 51, firstly transform the patient problem QE∈R l×d and the dialogue history HE∈R m×d To splice: The concat =Concat(HE,QE) Where l is the length of the QE sequence, d is the dimension of the word embedding, Concat is the addition function, m is the length of the dialogue history sequence, and the generated intermediate vector is E concat ; Step 52, using the self-attention mechanism, fuses the text features of the patient's question and the conversation history, and calculates the attention weight matrix: Where W q is the weight matrix, I is the intermediate value of the attention calculation, A is the attention weight, i and j are the indices of I, and T is the matrix transpose operator used to convert E concat Converted into a query matrix, which represents the model's attention to the input sequence; Step 53, multiply the attention weight of each position by the intermediate vector of the corresponding position, and sum the results of all positions to obtain the final context representation: C=A·E concat Where C is the context representation after the attention mechanism. In order to provide the model with more informative input, C and E are combined. concat Splicing, E interact This is the concatenated matrix: It is interact =Concat(E concat ,C) Step 54: In order to increase the model's expressiveness and adaptability to medical data, a learnable linear transformation W is introduced. intercat ∈R (l+m)×(l+m) and noise offset b intercat ∈R (l+m)×1 Act on E in step 53 interact Go up and get E final , and then activate, tanh is the activation function, and get the final feature E activated : This interactive process describes in detail the complex relationship between patient questions, doctor-patient dialogue history, and the generated intermediate vectors, while increasing the expressiveness and adaptability of the model by introducing learnable parameters.
6. According to claim 1, a multimodal medical AI-assisted consultation method integrating medical images and medical texts is characterized by: The text information and image information fusion method described in step 6 includes the following steps: Step 61, the final feature E of the text information in step 54 activated ∈R (l+m)×(l+m) And the image feature vector IE∈R p ×q , where p is the length of the image feature vector, q is the dimension of the image feature, and softmax is the activation function, which is used as the input of the cross-channel attention layer CA: Among them, W q , W k , W v is a learnable weight matrix; Step 62, concatenate the final features of the text information with the output of the cross-channel attention layer, retaining the original text information and the information interacting with the image information: It is cross =Concat(E activated ,CA(E activated , IE)) Step 63, apply a linear transformation and GELU activation function to obtain the fused feature vector: Among them, W fusion is the learned weight matrix, b fusion is the bias term.
7. According to claim 1, a multimodal medical AI-assisted consultation method integrating medical images and medical texts is characterized by: The method of fine-tuning the algorithm model using a large language model decoder in step 7 comprises the following steps: Step 71, using the fused feature vector obtained in the above step 63 as the input of the ChatGLM-6B language model, so that it outputs a result based on universal language understanding, wherein ChatGLM is the ChatGLM-6B language model: ZERO orig =ChatGLM(E fusion ) Step 72, in order to make ChatGLM-6B have the ability to answer questions in the medical field, a medical field adapter is introduced. The adapter is a small neural network that is added to the output of the original model to adjust the model to adapt to the new task: O med =O orig +Adapter med (GELU(W adapter O orig +b adapter )) Among them, Adapter med It is a medical professional adapter. adapter is the adapter weight, b adapter It is the adapter bias, which can include linear transformation and nonlinear GELU activation function; Step 73, fine-tune the adapter in step 72 above. The goal of fine-tuning is to minimize the loss function L of the new task. med ,This loss function focuses on the question answering task in the medical professional field; Where Y is the true label, log is the logarithmic function, O med is the prediction of the model on the fine-tuning task, i is the label index, and this objective function is used to measure the performance of the model on the medical question-answering task. The objective function of the overall fine-tuning is then defined as the weighted sum of the pre-training task loss and the fine-tuning task loss: THE total =λL pretrain +(1-λ)L med Where λ is a hyperparameter that weighs the pre-training task and the fine-tuning task, L pretrain is the loss of the ChatGLM-6B model on the pre-training task, and finally the adapter parameter W is adjusted using gradient descent adapter and b adapter To update: Here, μ is the learning rate.
Citation Information
Cited By
Skin disease auxiliary evaluation system and device, storage medium and program product
CN120809176A
Multi-modal medical data synthesis method and device, electronic equipment and storage medium
CN122337676A