Medical image visual question and answer method based on multi-modal large language model

Through the dual vision encoder architecture and progressive feature fusion strategy, the visual perception ability of the medical multimodal large language model is enhanced, the fine-grained information loss and visual deviation problems are solved, and efficient and stable medical image visual Q&A performance is achieved.

CN120339796APending Publication Date: 2025-07-18CHONGQING UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510473383.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing medical multimodal large language model has problems with insufficient fine-grained information capture and visual representation bias in visual encoder, resulting in limited performance when processing medical images.

Method used

Using a dual vision encoder architecture, integrating CLIP-ViT and DinoV2, the visual perception ability is enhanced through the progressive feature fusion module, and combining the lightweight large language model Phi-3.5-mini and connector module to perform multi-level feature fusion and training optimization.

Benefits of technology

It significantly improves the model's fine-grained information capture ability of medical images, reduces visual bias, improves the training stability and performance of the model, and is better than the existing 7B-13B scale model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339796A_ABST
    Figure CN120339796A_ABST
Patent Text Reader

Abstract

The invention discloses a medical image visual question and answer method based on a multi-modal large language model, and relates to the technical field of artificial intelligence and medical images. The method has the advantages that the fine-grained information capturing capability is remarkably enhanced, detail information such as edges and textures of medical images is effectively reserved by fusing middle-level features (such as the sixteenth layer) and high-level features (such as the twenty-third layer) of a visual encoder, and the problem of fine-grained information loss caused by single high-level features is solved; compared with the prior art, visual representation comprehensiveness is improved, CLIP-ViT and DinoV2 double visual encoders are integrated, image-text consistency features and image inherent structure features are captured respectively, diversified semantic information is complementarily covered, and visual deviation of a single encoder is remarkably reduced; according to the method, a progressive fusion strategy is adopted, multi-level features of the double encoders are integrated in a staged mode, the influence of feature distribution differences on gradients is reduced through feature normalization and alignment operation, and efficient and stable convergence of the model is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and medical image technology, and particularly relates to a medical image visual question answering method based on a multimodal large language model. Background Art

[0002] In recent years, large language models (LLMs) have made significant progress in natural language processing tasks. On this basis, by integrating visual encoders and modality connectors, multimodal large language models (MLLMs) have further expanded their general capabilities and achieved breakthrough progress in multimodal tasks such as image caption generation and visual question answering (VQA). This trend has rapidly extended to the medical field, driving the development of a series of medical-specific multimodal large language models. These medical multimodal large language models generally adopt conventional multimodal large language model architectures, mainly consisting of visual encoders, modality connectors, and large language models. By training on medical data sets, these models have significantly improved their performance in medical tasks such as medical visual question answering, medical report generation, and disease diagnosis.

[0003] However, the current development of medical multimodal large language models focuses more on the optimization of large language model components and training data, such as using larger-sized large language models, more efficient large language models, and larger-scale high-quality medical data, but relatively less attention is paid to visual encoders. Currently, the mainstream method usually uses CLIP-ViT as the visual encoder and uses the output of its penultimate layer as the visual representation. Existing medical multimodal large language models have significant limitations in visual encoders, especially in terms of capturing fine-grained information and the comprehensiveness of visual representations.

[0004] The main problems are as follows:

[0005] (1) Problem of loss of fine-grained information: Although the high-level abstract features of the visual encoder can effectively capture the global information and overall concepts of images, they lack the capture of detailed information such as edges and textures. This limitation significantly reduces the ability of multimodal large language models to perceive fine-grained information, resulting in their inability to fully capture the complexity of medical images. Medical images usually contain a large amount of detailed information, such as the edges, textures, and shapes of lesions, which are crucial for accurate disease diagnosis and treatment. Existing visual encoders perform poorly in processing this detailed information, resulting in limited performance of the model in medical tasks.

[0006] (2) Visual Representation Bias Problem: Existing visual encoders are usually optimized based on specific pre-training objectives, which leads to an inherent bias in their visual focus. Therefore, a single visual encoder is prone to falling into a "perspective blind spot" and having difficulty comprehensively covering the diverse semantic information of images. For example, CLIP trained based on image-text contrast is good at capturing image-text consistency features and pays extra attention to object-level targets; while DinoV2 based on self-supervised learning focuses more on the internal structure of images and can effectively capture textures, contours, and the relationships between local regions. This bias makes it difficult for the model to comprehensively capture all relevant visual information when processing diverse medical images, thus affecting the generalization ability and robustness of the model.

[0007] Therefore, there is a need for a new visual coding method to solve the above problems and improve the performance of medical multi-modal large language models in complex medical tasks. Summary of the Invention

[0008] The purpose of the present invention is to provide a medical image visual question answering method based on a multi-modal large language model, which enhances the visual perception ability of the multi-modal large language model by progressively fusing multi-level features of multiple visual encoders, so as to solve the technical problems proposed in the background art.

[0009] To achieve the above purpose, the present invention provides the following technical solutions: A medical image visual question answering method based on a multi-modal large language model, at least including the following steps:

[0010] S1: Build a dual visual encoder architecture, integrating CLIP-ViT and DinoV2, and reduce the "visual bias" of the multi-modal large language model through feature complementarity to enhance the medical image representation ability;

[0011] S2: Design a progressive feature fusion module (DualProFusion), and fuse the multi-level features output by the dual visual encoder architecture to enhance the visual perception ability of the multi-modal large language model;

[0012] S3: Construct a lightweight multi-modal large language model, using lightweight Phi-3.5-mini as the large language model (LLM) to reduce the number of model parameters and the computational resource requirements;

[0013] S4: Design a connector module, the connector module follows the design of LLaVA, and uses a two-layer multi-layer perceptron (MLP), and the connector module includes two linear layers and a GELU activation function;

[0014] S5: Build the Agamotto model based on the dual visual encoder architecture, the progressive feature fusion module, the connector module, and the large language model module;

[0015] S6: Adopt a phased training strategy to further optimize and improve the performance of the Agamotto model;

[0016] S7: Evaluate the Agamotto model on multiple medical visual question - answering benchmarks, verify the model performance in multiple dimensions to ensure the effectiveness and reliability of the Agamotto model in actual medical applications. After passing the verification, use the Agamotto model for medical image visual question - answering.

[0017] Furthermore, the dual - vision encoder architecture includes two pre - trained vision encoders, aiming to process the input medical images and generate corresponding visual representations;

[0018] The application of the dual - vision encoder architecture at least includes the following steps:

[0019] Given an input image

[0020] The input image is first divided into several non - overlapping image patches by the CLIP - ViT image processor;

[0021] Subsequently, the divided image patches are respectively input into two vision encoders, CLIP - ViT and DinoV2. Each vision encoder will output corresponding visual feature representations Z at each level i,j , see the following formula:

[0022] Z i,j =V i,j (P x ), i ∈ {CLIP, Dinov2}, j ∈ {16, 23}

[0023] Where: P x represents the image patches; N represents the number of image patches; V i,j represents the j - th layer of the vision encoder i, and D i,j represents the feature dimension of the j - th layer of the vision encoder i.

[0024] Furthermore, the feature fusion of the progressive feature fusion module at least includes the following steps:

[0025] First, concatenate the middle - layer features of CLIP - ViT and the middle - layer features of DinoV2 along the channel dimension and perform normalization processing through LayerNorm to generate a preliminary fusion feature F1, see the following formula:

[0026] F1 = Norm1(Concat[Z DINOv2,16 , Z CLIP,16 , dim = channe

[0027] Among them, Concat means to splice two feature tensors along the channel dimension, i.e., dim = channel; Norm1 means to normalize the spliced features; Z DINOv2,16 is the middle-level feature of DinoV2; Z CLIP,16 is the middle-level feature of CLIP-ViT;

[0028] Subsequently, F1 is spliced with the high-level feature of DinoV2 and normalized again to obtain the intermediate feature F2, as shown in the following formula:

[0029] F2 = Norm2(Concat[F1, Z DINOv2,23 , dim = channel)

[0030] Among them, Z DINOv2,23 is the high-level feature of DinoV2;

[0031] Finally, F2 is spliced with the high-level feature of CLIP-ViT and finally normalized to output the fused feature F final , as shown in the following formula:

[0032] F final = Norm3(Concat[F2, Z CLIP,23 , dim = channel)

[0033] Among them, Z CLIP,23 is the high-level feature of CLIP-ViT;

[0034] Through the above phased progressive fusion strategy and layer-by-layer normalization operation, the differences in numerical scale and gradient distribution of different encoder features are effectively aligned, avoiding the problem of gradient explosion caused by direct fusion;

[0035] At the same time, on the basis of retaining the advantageous features of each encoder, the efficient fusion of global semantics and local details is realized, significantly improving the model's perception ability of medical images.

[0036] Furthermore, for the input text data, the large language model module first converts it into a text Token sequence through a tokenizer

[0037] Subsequently, the text Token sequence T x is spliced with the visual embedding E x and input into the large language model for joint processing;

[0038] Based on the input visual embedding and text Token sequence, the large language model generates corresponding multi-modal response outputs.

[0039] Furthermore, the connector module is used to reduce the dimension of the visual feature representation and map it to the word vector space. Specifically:

[0040] The first layer represents the visual feature F final The dimension of is adjusted to the hidden dimension D of the large language model (LLM) t Stay consistent;

[0041] The second layer maintains the dimension D t The visual feature representation remains unchanged and is further mapped to the word vector space so that it can be effectively understood by LLM;

[0042] After the connector processing, the final visual embedding is obtained

[0043] Further, the S6 at least includes the following steps:

[0044] First, in the pre-training stage, the dual visual encoder and large language model parameters are frozen, and only the connector module is trained to achieve visual-language modality alignment;

[0045] Secondly, in the stage of fine-tuning the instructions: unfreeze all parameters and optimize the responsiveness of medical instructions;

[0046] In the supervised fine-tuning stage, the performance for specific tasks is improved by fine-tuning on the VQA-RAD, SLAKE and Path-VQA datasets as needed.

[0047] Compared with the prior art, the present invention has the following beneficial effects:

[0048] 1. The ability to capture fine-grained information is significantly enhanced. By fusing the middle-level features (such as the 16th layer) and high-level features (such as the 23rd layer) of the visual encoder, the present invention effectively retains the edge, texture and other detail information of the medical image, solving the problem of fine-grained information loss caused by a single high-level feature;

[0049] 2. Comprehensive improvement of visual representation. The present invention integrates CLIP-ViT and DinoV2 dual visual encoders to capture image-text consistency features and image inherent structural features respectively, complementarily covering diverse semantic information, and significantly reducing the visual bias of a single encoder;

[0050] 3. Training stability optimization. The present invention adopts a progressive fusion strategy to integrate the multi-level features of the dual encoders in stages, and reduces the impact of feature distribution differences on the gradient through feature normalization and alignment operations, ensuring efficient and stable convergence of the model.

[0051] 4. Parameter efficiency and performance advantages: With only 4.6B parameters, the model (Agamotto) proposed in the present invention significantly outperforms existing models with a scale of 7B - 13B on average in multiple medical visual question - answering benchmark tests, verifying the efficiency and practicality of the technical solution. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0053] Figure 1 It is a training strategy diagram of the Agamotto model provided by the present invention;

[0054] Figure 2 It is a schematic structural diagram of the Agamotto model provided by the present invention;

[0055] Figure 3 It is a visualization example diagram of the feature maps of different layers of the existing CLIP - ViT and DinoV2 models provided by the present invention;

[0056] Figure 4 It is a schematic diagram of a more comprehensive and more fine - grained answer of the Agamotto model provided by the present invention compared with existing methods. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments.

[0058] Regarding the following technical problems existing in existing medical multi - modal large language models (MLLMs): (1) Fine - grained information loss: The single visual encoder relies on high - level abstract features, resulting in the loss of detailed information such as edges and textures in medical images; (2) Visual representation bias: Due to the bias of the pre - training objective of the single visual encoder, there is a "viewpoint blind spot" and it is difficult to comprehensively capture the diverse semantic information of medical images. The present invention proposes a technical solution based on dual visual encoders and progressive feature fusion, aiming to enhance the fine - grained information perception ability through multi - level feature fusion and alleviate visual bias through the complementarity of the dual encoders, thereby improving the performance and robustness of medical multi - modal large language models in tasks such as visual question - answering and disease diagnosis. The specific objectives are as follows:

[0059] (1)Propose to use multi-level feature fusion to make up for the fine-grained information loss problem existing in the existing methods; use a dual visual encoder to alleviate the "visual bias" problem of the existing methods;

[0060] (2)Fuse the multi-level features of the dual visual encoder in stages, and perform normalization and alignment at each fusion stage to enhance visual perception while stabilizing the training process;

[0061] Furthermore, the present invention proposes: A medical image visual question answering method based on a multi-modal large language model, at least including the following steps:

[0062] S1: Build a dual visual encoder architecture, integrate CLIP-ViT and DinoV2, and alleviate the "visual bias" of the multi-modal large language model through feature complementarity to enhance the medical image representation ability. The feature map is shown as Figure 3 ;

[0063] Relying solely on the high-level semantic features of a single visual encoder is likely to lead to the problems of fine-grained information loss and visual bias, thus severely restricting the visual perception ability of the medical multi-modal large language model. CLIP based on image-text contrast training is good at capturing image-text consistency features and pays extra attention to object-level targets; while DinoV2 based on self-supervised learning pays more attention to the self-structure of the image and can effectively capture the relationship between textures, contours and local regions (such as Figure 3 visualization results). Through comparative experiments, it is verified that for the same medical image, CLIP and DinoV2 have different visual attention points, and their complementarity covers the object-level semantics and local structure features in the medical image. The middle-level features (the 16th layer) reduce redundancy while retaining details;

[0064] The dual visual encoder architecture includes two pre-trained visual encoders, which are designed to process the input medical image and generate the corresponding visual representation;

[0065] The application of the dual visual encoder architecture at least includes the following steps:

[0066] Given the input image

[0067] The input image is first divided into several non-overlapping image patches (patches) by the CLIP-ViT image processor;

[0068] Subsequently, the divided image patches are respectively input into two visual encoders, CLIP-ViT and DinoV2. Each visual encoder will output the corresponding visual feature representation Z i,j , see the following formula:

[0069] Z i,j = Vi,j (P x ), where \(i\in\{CLIP, Dinov2\}\) and \(j\in\{16, 23\}\)

[0070] Wherein: P x represents image patches; \(N\) represents the number of image patches; \(V\) i,j represents the \(j\)-th layer of the visual encoder \(i\), and \(D\) i,j represents the feature dimension of the \(j\)-th layer of the visual encoder \(i\);

[0071] Since there is a high similarity in the detailed information of the low-level and middle-level features of CLIP-ViT and DinoV2, directly fusing the low-level, middle-level, and high-level features may lead to an increase in information redundancy and computational overhead. Considering that while retaining detailed information, the middle-level features also possess certain semantic abstraction capabilities and can effectively capture the local structural features and local context information of images, the present invention selects the middle-level features (the 16th layer) and high-level features (the 23rd layer) of the visual encoder as visual representations. This can avoid information redundancy while enriching the amount of information in the visual representation.

[0072] S2: Design a progressive feature fusion module (DualProFusion) to enhance the visual perception ability of the multi-modal large language model by fusing the multi-layer features output by the dual visual encoder architecture;

[0073] To effectively integrate the multi-level visual feature representations extracted by the two visual encoders, CLIP-ViT and DinoV2, the present invention proposes DualProFusion. This module concatenates (concat) multiple visual feature representations in the channel dimension to balance the computational efficiency and the amount of information in the visual representation. Considering that there are differences in the focus of different visual encoders and their respective levels on image features, resulting in significant differences in the numerical scale and gradient distribution of the features, direct fusion may cause problems with training instability. Therefore, the DualProFusion module adopts a three-stage progressive fusion strategy, by fusing the different-level features of the two visual encoders in stages and gradually implementing normalization operations during the fusion process to stabilize the feature distribution.

[0074] The feature fusion of the progressive feature fusion module includes at least the following steps:

[0075] First, concatenate the middle-level features of CLIP-ViT and the middle-level features of DinoV2 along the channel dimension and perform normalization processing through LayerNorm to generate the preliminary fusion feature \(F1\), see the following formula:

[0076] \(F1 = Norm1(Concat[Z DINOv2,16,Z CLIP,16 , dim = channe

[0077] Among them, Concat represents concatenating two feature tensors along the channel dimension, i.e., dim = channe; Norm1 represents normalizing the concatenated features; Z DINOv2,16 is the middle-level feature of DinoV2; Z CLIP,16 is the middle-level feature of CLIP-ViT;

[0078] Subsequently, F1 is concatenated with the high-level feature of DinoV2 and normalized again to obtain the intermediate feature F2, as shown in the following formula:

[0079] F2 = Norm2(Concat[F1, Z DINOv2,23 , dim = channel)

[0080] where Z DINOv2,23 is the high-level feature of DinoV2;

[0081] Finally, F2 is concatenated with the high-level feature of CLIP-ViT and subjected to final normalization to output the fused feature F final , as shown in the following formula:

[0082] F final = Norm3(Concat[F2, Z CLIP,23 , dim = channel)

[0083] where Z CLIP,23 is the high-level feature of CLIP-ViT;

[0084] Through the above phased progressive fusion strategy and layer-by-layer normalization operation, the differences in numerical scale and gradient distribution of different encoder features are effectively aligned, avoiding the problem of gradient explosion caused by direct fusion (experiments show that the training loss fluctuation is reduced by 37%);

[0085] At the same time, on the basis of retaining the dominant features of each encoder, the efficient fusion of global semantics and local details is achieved, significantly improving the model's perception ability of medical images.

[0086] S3: Construct a lightweight multi-modal large language model, using the lightweight Phi-3.5-mini as the large language model (LLM) to reduce the number of model parameters and computational resource requirements;

[0087] Select Phi-3.5-mini as the core language model component. Its 4.6B parameter design not only ensures the model performance but also achieves a significant improvement in computational efficiency. This architecture has a 1.8-fold increase in inference speed compared to traditional 7B models, greatly optimizing the hardware resource utilization efficiency. The connector module uses a two-layer MLP (with GELU activation function) to map to the LLM word vector space, generate visual embeddings, and after concatenating with text tokens, input them into the LLM for multimodal reasoning.

[0088] For the input text data, the large language model module first converts it into a text Token sequence through a tokenizer.

[0089] Subsequently, the text Token sequence T x is concatenated with the visual embedding E x and input into the large language model for joint processing;

[0090] Based on the input visual embedding and text Token sequence, the large language model generates corresponding multimodal response outputs; The method of the present invention selects Phi-3.5-mini as the LLM component, which is a lightweight and efficient open-source model.

[0091] S4: Design a connector module. The connector module follows the design of LLaVA and uses a two-layer multi-layer perceptron (MLP). The connector module includes two linear layers and a GELU activation function;

[0092] The connector module is used to reduce the dimension of the visual feature representation and map it to the word vector space. Specifically:

[0093] The first layer adjusts the dimension of the visual feature representation F final to be consistent with the hidden dimension D t of the large language model (LLM);

[0094] The second layer keeps the dimension D t unchanged and further maps the visual feature representation to the word vector space so that it can be effectively understood by the LLM;

[0095] After being processed by the connector, the final visual embedding is obtained

[0096] S5: Build the Agamotto model based on the dual visual encoder architecture, progressive feature fusion module, connector module, and large language model module;

[0097] The dual visual encoder module uses two pre-trained visual encoders, CLIP-ViT (contrastive learning-based) and DinoV2 (self-supervised learning-based), to extract the global semantic features and local structural features of medical images respectively; the DualProFusion module integrates the multi-level features of the dual encoders through a progressive fusion strategy to enhance the fine-grained information capture ability and reduce visual bias; the connector module uses a multi-layer perceptron (MLP) to map the fused visual features into the word vector space of the LLM to achieve visual-linguistic modality alignment; the large language model module selects the lightweight Phi-3.5-mini as the language model, receives the joint input of visual embeddings and text tokens, and generates multi-modal responses. The overall architecture of the model is shown in Figure 2 。

[0098] S6: Adopt a phased training strategy, and the strategy diagram is as Figure 1 to further optimize and improve the performance of the Agamotto model;

[0099] S6 includes at least the following steps:

[0100] First, in the pre-training stage, freeze the parameters of the dual visual encoder and the large language model, and only train the connector module to achieve visual-linguistic modality alignment;

[0101] Secondly, in the instruction fine-tuning stage: unfreeze all parameters and optimize the medical instruction response ability;

[0102] Furthermore, in the supervised fine-tuning stage, fine-tune according to requirements on the VQA-RAD, SLAKE, and Path-VQA datasets to improve the performance for specific tasks.

[0103] S7: Evaluate the Agamotto model on multiple medical visual question answering benchmarks, verify the model performance in multiple dimensions to ensure the effectiveness and reliability of the Agamotto model in actual medical applications. After passing the verification, use the Agamotto model for medical image visual question answering.

[0104] In multiple medical visual question answering benchmark tests, verify the zero-shot performance and supervised fine-tuning performance of the model in VQA-RAD, SLAKE, Path-VQA, etc. For closed-set questions, use accuracy as the evaluation metric; for open-set questions, use recall to evaluate the proportion of true labels appearing in the generated sequences. Ensure the effectiveness of the model in actual medical applications.

[0105]

[0106] In the formula, TP is the true positive, that is, the number of samples correctly predicted as the positive class; TN is the true negative, that is, the number of samples correctly predicted as the negative class; FP is the false positive, that is, the number of samples incorrectly predicted as the positive class; FN is the false negative, that is, the number of samples incorrectly predicted as the negative class.

[0107] For further verification, it is proposed to refer to Figure 4 Taking medical visual question answering as the core application scenario, the present invention verifies the effectiveness of the proposed Agamotto model through three authoritative medical data sets. Finally, the present invention selects Phi-3.5-mini as the lightweight large language model component, and innovatively adopts the CLIP-ViT and DinoV2 dual visual encoder architectures, combined with the progressive feature fusion strategy (DualProFusion), and trains and evaluates on the processed medical data set.

[0108] The PubMedVision data set is used as the basic training data for the pre-training and instruction fine-tuning phases. After strict medical image screening and question integrity verification, only the single-image-text pair part of this data set is used. A total of 500,000 pairs of high-quality medical image-text pairs are retained in the pre-training phase, and a total of 500,000 pairs of high-quality medical image-text pairs are retained in the instruction fine-tuning phase. After pre-training and instruction fine-tuning, tests are carried out on three medical question-answering data sets, including radiology image question answering and pathology image question answering.

[0109] Table 1 details the division of each data set and key statistical indicators. To ensure the fairness of the evaluation, the present invention adopts a unified standard: the accuracy rate is used as the evaluation index for closed-set questions, and the recall rate is used to evaluate open-set questions. This strict evaluation method avoids the problem of overestimating performance that may be brought about by simplifying open questions into classification tasks.

[0110] Table 2 shows the Zero-Shot performance comparison of the Agamotto model under different configurations. The comparison results show that the Agamotto model demonstrates significant advantages in the medical visual question answering task by fusing multi-level features of dual encoders: in the Zero-Shot setting, the average performance is improved by 7.30% compared with LLaVA-Med-v1.5 and by 6.45% compared with Med-Moe (2.7B×4). It reflects the strong generalization ability of Agamotto in the medical visual question answering task. Notably, in terms of computing resource consumption, the 4.6B-parameter Agamotto only requires 65% of the video memory occupancy of the same-level 7B model, but achieves better performance.

[0111] To further verify the adaptability of Agamotto in downstream tasks, we performed supervised fine-tuning of Agamotto on the training sets of three medical visual question answering benchmarks, namely VQA-RAD, SLAKE, and Path-VQA. The training set includes 27,000 pairs of image-text pairs, which contain radiology images and pathology images.

[0112] Table 3 compares the performance differences between Agamotto and the current state-of-the-art medical MLLMs under the supervised fine-tuning setting. The average performance of Agamotto in the three medical visual question answering benchmarks reaches 77.5%, an improvement of 3.45% compared to the sub-optimal method. These results fully demonstrate that, while maintaining the lightweight of the model, the present invention has successfully achieved a breakthrough in medical multimodal understanding through an innovative dual-encoder architecture and feature fusion strategy, and also proves its feasibility and efficiency in actual deployment.

[0113] Table 1 Details of the datasets used in the invention embodiments

[0114]

[0115] Table 2 Comparison results between the Agamotto model and existing models under the Zero-Shot setting

[0116]

[0117] Table 3 Comparison results between the Agamotto model and existing models in the supervised fine-tuning setting

[0118]

[0119] In summary:

[0120] The method proposed in the present invention effectively improves the ability of medical multimodal large language models in fine-grained perception, alleviates the "visual bias" problem, significantly improves the training stability, and its performance is significantly better than existing methods.

[0121] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, in any regard, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed by the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.

Claims

1. A medical image visual question answering method based on a multimodal large language model, characterized in that: At least include the following steps: S1: Build a dual-vision encoder architecture, integrate CLIP-ViT and DinoV2, and mitigate the "visual bias" of the multimodal large language model through feature complementarity to enhance the medical image representation ability; S2: Design a progressive feature fusion module to fuse the multi-layer features output by the dual-vision encoder architecture to enhance the visual perception ability of the multimodal large language model; S3: Construct a lightweight multimodal large language model, adopt the lightweight Phi-3.5-mini as the large language model to reduce the number of model parameters and the computational resource requirements; S4: Design a connector module. The connector module follows the design of LLaVA and adopts a two-layer multi-layer perceptron. The connector module contains two linear layers and a GELU activation function; S5: Build the Agamotto model based on the dual-vision encoder architecture, progressive feature fusion module, connector module and large language model module; S6: Adopt a phased training strategy to further optimize and improve the performance of the Agamotto model; S7: Evaluate the Agamotto model on multiple medical visual question answering benchmarks, verify the model performance in multiple dimensions to ensure the effectiveness and reliability of the Agamotto model in actual medical applications. After passing the verification, use the Agamotto model for medical image visual question answering.

2. The medical image visual question answering method based on a multimodal large language model according to claim 1, wherein: The dual-vision encoder architecture includes two pre-trained vision encoders, which are designed to process the input medical images and generate corresponding visual representations; The application of the dual-vision encoder architecture at least includes the following steps: Given input image The input image is first divided into several non-overlapping image patches by the CLIP-ViT image processor; Subsequently, the segmented image patches are respectively input into two vision encoders, CLIP-ViT and DinoV2, and each vision encoder outputs corresponding visual feature representations Z at each level. i,j , as shown in the following formula: Z i,j = V i,j (P x ), i ∈ {CLIP, Dinov2}, j ∈ {16, 23} Wherein: P x represents image patches; N represents the number of image patches; V i,j represents the j-th layer of the visual encoder i, D i,j represents the feature dimension of the j-th layer of the visual encoder i.

3. A medical image visual question answering method based on a multimodal large language model according to claim 2, wherein: The feature fusion of the progressive feature fusion module at least includes the following steps: First, splice the middle-level features of CLIP-ViT and the middle-level features of DinoV2 along the channel dimension and perform normalization processing through LayerNorm to generate the preliminary fusion feature F1. See the following formula: F1 = Norm1(Concat[Z DINOv2,16 , Z CLIP,16 , dim = channe Among them, Convat means concatenating two feature tensors along the channel dimension, i.e., dim = channel; Norm1 means normalizing the concatenated features; Z DINOv2,16 is the middle-level feature of DinoV2; Z CLIP,16 is the middle-level feature of CLIP-ViT; Subsequently, splice F1 with the high-level features of DinoV2 and normalize again to obtain the intermediate feature F2. See the following formula: F2 = Norm2(Concat[F1, Z DINOv2,23 , dim = channel) Among them, Z DINOv2,23 is the high-level feature of DinoV2; Finally, F2 is concatenated with the high-level features of CLIP-ViT and finally normalized to output the fused feature F final , as shown in the following formula: F final = Norm3(Concat[F2, Z CLIP,23 , dim = channel) Among them, Z CLIP,23 is the high-level feature of CLIP-ViT; Through the above phased progressive fusion strategy and layer-by-layer normalization operation, the differences in numerical scale and gradient distribution of different encoder features are effectively aligned, avoiding the problem of gradient explosion caused by direct fusion; At the same time, on the basis of retaining the advantageous features of each encoder, the efficient fusion of global semantics and local details is realized, significantly enhancing the model's perception ability of medical images.

4. The medical image visual question answering method based on a multimodal large language model according to claim 3, wherein: For the input text data, the large language model module first converts it into a text Token sequence through a tokenizer Subsequently, the text Token sequence T x is concatenated with the visual embedding E x and input into the large language model for joint processing; The large language model generates corresponding multimodal response outputs based on the input visual embeddings and text Token sequences.

5. The medical image visual question answering method based on a multimodal large language model according to claim 4, wherein: The connector module is used to reduce the dimension of the visual feature representation and map it to the word vector space. Specifically: The first layer adjusts the dimension of the visual feature representation F final to be consistent with the hidden dimension D of the large language model t ; The second layer maintains the dimension D t unchanged and further maps the visual feature representation to the word vector space so that it can be effectively understood by the LLM; After being processed by the connector, a visual embedding is finally obtained 6. The medical image visual question answering method based on a multimodal large language model according to claim 1, characterized in that: S6 at least includes the following steps: First, freeze the parameters of the dual-vision encoder and the large language model during the pre-training stage, and only train the connector module to achieve visual-language modality alignment; Secondly, in the instruction fine-tuning stage: unfreeze all parameters and optimize the medical instruction response ability; Furthermore, during the supervised fine-tuning stage, the performance for specific tasks is improved by fine-tuning according to requirements on the VQA-RAD, SLAKE, and Path-VQA datasets.

Citation Information

Cited By

  • Method for supervising visual language model training by using diffusion model

    CN120580446A