Medical Visual Question Answering Method and Device Based on Causal Reasoning and Large Language Model

Through the method of combining causal reasoning with large language models, the integration of image and text information and causal reasoning in medical visual question-and-answer questions are solved, and more accurate and in-depth question-and-answer results are achieved, improving the system's understanding and answering ability.

CN119537518BActive Publication Date: 2025-07-11TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411176071.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-26
Publication Date
2025-07-11
Estimated Expiration
2044-08-26

AI Technical Summary

Technical Problem

The existing medical visual question-and-answer methods lack a comprehensive understanding of the context of medical images, resulting in a lack of correlation or accuracy in answers, the inability to effectively use image and text information for causal reasoning and comprehensive analysis, the lack of ability to integrate and utilize multimodal data, and the inability to fully mine the correlation information between medical images and text.

Method used

Using a method based on causal reasoning and large language model, we extract initial features from medical images and problem texts, perform feature mapping and fusion, use multi-head attention and front-door adjustment to perform causal intervention, generate medical Q&A results, and provide contextual references in combination with pre-trained large language models.

Benefits of technology

It improves the performance and accuracy of the medical visual question-and-answer system, can understand the questions more comprehensively, provide deep and credible answers, effectively remove confounding factors, and improves the ability to analyze and answer complex medical questions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119537518B_ABST
    Figure CN119537518B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a medical visual question answering method and device based on causal reasoning and large language models, which are applied to the field related to computer vision. The method includes: S1, respectively extracting initial visual features and initial text features from medical images and question texts; S2, performing a feature mapping operation on the initial visual features; S3, fusing the local visual features obtained in step S2 with the initial text features; S4, fusing the fusion features with the global visual features and local visual features to obtain a visual mediating variable; and fusing the fusion features with the initial text features to obtain a text mediating variable; S5, performing a causal intervention operation; S6, generating corresponding question and answer pairs according to training samples; S7, generating medical question answering results. In this way, the present invention solves the technical problems that the current medical visual question answering methods have some limitations, such as the answers to questions lacking relevance or accuracy, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision, natural language processing and intelligent recognition, and in particular to a medical visual question answering method and device based on causal reasoning and large language models. Background Art

[0002] In the fields of medical image processing and visual question answering, with the development and popularization of medical imaging technology, more and more medical image data have been widely collected and applied. However, traditional medical image processing methods are often limited to solving specific tasks and lack an understanding of the overall context and causal relationships. At the same time, although large language models have achieved remarkable results in natural language processing tasks, they face many challenges in medical visual question answering tasks, such as insufficient ability to understand and interpret medical images, and high requirements for the accuracy and richness of medical knowledge. Therefore, current medical visual question answering methods have some limitations, including but not limited to: lack of a comprehensive understanding of the medical image context, resulting in answers to questions that may lack relevance or accuracy; limited ability to express and reason about medical knowledge, unable to make good use of image and text information for causal reasoning and comprehensive analysis; lack of the ability to integrate and utilize multi-modal data, unable to fully explore the correlation information between medical images and texts. Summary of the Invention

[0003] The purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art, and provide a medical visual question answering method and device based on causal reasoning and large language models for the current medical visual question answering task, so as to solve some limitations existing in the current medical visual question answering methods, including but not limited to: lack of a comprehensive understanding of the medical image context, resulting in answers to questions that may lack relevance or accuracy; limited ability to express and reason about medical knowledge, unable to make good use of image and text information for causal reasoning and comprehensive analysis; lack of the ability to integrate and utilize multi-modal data, unable to fully explore the correlation information between medical images and texts, and other technical problems.

[0004] According to the first aspect of the present invention, there is provided a medical visual question answering method based on causal reasoning and large language models, including the following steps:

[0005] S1. Extract initial visual feature f i and initial text feature f q respectively from the medical image and the question text;

[0006] S2. Perform a feature mapping operation on the initial visual feature f i to obtain a global visual feature f ig and a local visual feature f il ;

[0007] S3. Combine the local visual feature f il with the initial text feature f q to perform feature fusion and obtain a fused feature f iq ;

[0008] S4. Combine the fused feature f iq with the global visual feature f ig and the local visual feature f il to perform feature fusion and obtain a visual intermediate variable m i ; and combine the fused feature f iq with the initial text feature f q to perform feature fusion and obtain a text intermediate variable m q ;

[0009] S5. Perform causal intervention operations on the initial visual feature f i and the visual intermediate variable m i , and on the initial text feature f q and the text intermediate variable m q respectively to obtain a visual causal feature f i ' and a text causal feature f' q ;

[0010] S6. Use a pre-trained large language model to generate corresponding question-answer pairs based on training samples as the prompt content for the question-answer process to provide a rich context and reference for the question-answer process;

[0011] S7. Input the visual causal feature f i ', the text causal feature f' q and the text feature f p of the prompt content into the large language model to generate medical question-answer results.

[0012] Furthermore, in step S1, it includes:

[0013] Extract the features of the medical image into a feature vector of w×h×c, where w represents the width of the feature vector, h represents the height of the feature vector, and c represents the number of channels of the feature vector;

[0014] Extract the features of the question text into a feature vector of v×d, where v represents the size of the vocabulary and d represents the dimension of the word embedding.

[0015] Furthermore, in step S2, the global visual feature f ig provides information about the shape, position, and overall state of the organ; the local visual feature f il provides detailed information, and the detailed information includes lesions and abnormalities.

[0016] Further, in step S3, the feature fusion formula is as follows:

[0017] f iq = MLP(MHA(f il , f q , f q ))

[0018] Further, in step S4, the following formula for extracting the mediating variable is used for feature fusion:

[0019] m i = MLP(MHA(f il , f ig , f ig ), MHA(f il , f iq , f iq ))

[0020] m q = MLP(MHA(f q , f q , f q ), MHA(f iq , f q , f q ))

[0021] Further, in step S5, the causal intervention formula is as follows:

[0022]

[0023] Among them, P(A|do(I,Q)) represents the total causal effect of the treatment variables, i.e., the image I and the question Q, on the outcome variable, i.e., the answer A; the do algorithm represents the intervention operation; m is the mediating variable; i and q respectively represent the feature vectors of the image and the question text; P(m|I,Q) represents the conditional probability distribution of the mediating variable m given the image I and the question Q; P(A|m,i,q) represents the conditional probability of generating the answer A given the mediating variable m and the image I and the question Q; P(i,q) represents the joint probability distribution of the image I and the question Q.

[0024] Further, in step S5, the causal intervention operation adopts front-door adjustment, and the formula for front-door adjustment is as follows:

[0025] f i ' = FDA(f i , m i ), f' q = FDA(f q , m q )

[0026] Among them, FDA represents the front-door adjustment process.

[0027] Further, in step S6, the generating of corresponding question-and-answer pairs by using the pre-trained large language model according to the training samples includes: generating question-and-answer pairs according to the training samples input to the model and the template types;

[0028] Among them, the template types are determined by the types of questions and are divided into open-ended questions and closed-ended questions;

[0029] Among them, in step S7, different cross-entropy losses are selected for closed-ended questions and open-ended questions, and are defined as follows:

[0030]

[0031] Among them, Lc is the cross-entropy loss of closed-ended questions, Lo is the cross-entropy loss of closed-ended questions, a t represents the true answer, T represents the training data, F() represents the overall question-and-answer framework, N represents the length of the answer, a 1:n-1 represents the partial output, p t represents the prompt text, and θ represents the model parameters; meanwhile, in order to ensure the consistency of the model's prediction of the original features and the causal features, the following loss function is defined:

[0032] L cau = KL(F(f i ', f' q ), F(f i , f q ))

[0033] The combined loss function is defined as:

[0034] L closed = L c + L cau , L open = L o + L cau

[0035] According to the second aspect of the present invention, there is also provided a medical visual question answering device based on causal reasoning and a large language model, including:

[0036] A basic network for respectively extracting initial visual features f i and initial text features f q from medical images and question texts; among them, the basic network includes a pre-trained image encoder and a text encoder;

[0037] A feature selector for performing a feature mapping operation on the initial visual feature f i to obtain a global visual feature f ig and a local visual feature fil ;

[0038] An information integration module for fusing the local visual feature f il with the initial text feature f q through multi-head attention to obtain a fused feature f iq ;

[0039] An intermediate variable extraction module for fusing the fused feature f iq with the global visual feature f ig and the local visual feature f il through multi-head attention to obtain a visual intermediate variable m i ; and fusing the fused feature f iq with the initial text feature f q through multi-head attention to obtain a text intermediate variable m q ;

[0040] A causal intervention operation module for performing causal intervention operations on the initial visual feature f i and the visual intermediate variable m i , the initial text feature f q and the text intermediate variable m q respectively to obtain a visual causal feature f i ' and a text causal feature f' q ;

[0041] A prompt module composed of pre-trained visual and language backbones, which uses a pre-trained large language model to generate corresponding question-answer pairs according to training samples and template structures as the prompt content for the question-answer process, providing a rich context and reference for the question-answer process;

[0042] A visual question-answer generation module for inputting the visual causal feature f i ', the text causal feature f' q and the text feature f p of the prompt content into a large language model to generate medical question-answer results.

[0043] Furthermore, the causal intervention operation module includes two front-door adjustment units, and each front-door adjustment unit is composed of two attention fusion modules;

[0044] The causal intervention operation module respectively inputs the initial visual feature f i and the visual intermediate variable m i , the initial text feature f q and the text intermediate variable m q into the two front-door adjustment units for causal intervention operations to obtain a visual causal feature fi 'with the text causal feature f' q 。

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] 1. The present invention proposes a medical visual question answering method that combines causal reasoning with a large language model, comprehensively and accurately analyzes and explains the causal relationship between medical images and related text information, and improves the performance and accuracy of the medical visual question answering system.

[0047] 2. The present invention can make full use of the multi-modal features of medical images and text information, enabling the question answering system to more comprehensively understand the questions, and thus providing more in-depth and credible answers.

[0048] 3. The present invention can effectively remove confounding factors in the process of medical visual question answering by emphasizing causal reasoning, and improve the system's ability to accurately understand and analyze medical questions.

[0049] 4. The present invention can provide effective question guidance and information prompts for the medical visual question answering system, thereby improving the system's ability to comprehensively understand and accurately answer medical questions.

[0050] 5. The present invention can combine causal reasoning with a large language model, improve the parsing and answering ability of the medical visual question answering system for complex medical questions, and make the question answering results more accurate and in-depth.

[0051] It should be understood that the content described in the summary of the invention section is not intended to limit the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Combined with the accompanying drawings and referring to the following detailed description, the above and other features, advantages and aspects of the embodiments of the present invention will become more obvious. The drawings are used to better understand the solution and do not constitute a limitation to the present invention. In the drawings, the same or similar reference numerals represent the same or similar elements, where:

[0053] Figure 1 shows a schematic flow chart of a medical visual question answering method based on causal reasoning and a large language model according to an embodiment of the present invention;

[0054] Figure 2 shows a schematic diagram of the visual question answering network structure of the present invention;

[0055] Figure 3 shows the causal reasoning graph and the front-door adjustment causal graph of the present invention;

[0056] Figure 4The deconfounding module in the causal reasoning module of the present invention is presented;

[0057] Figure 5 The accuracy comparison between the method of the present invention and other methods on the VQA-RAD, SLAKE, and PathVQA datasets is presented;

[0058] Figure 6 The question-and-answer result comparison between the method of the present invention and other methods on the SLAKE dataset is presented;

[0059] Figure 7 The module schematic diagram of a medical visual question answering device based on causal reasoning and large language model according to an embodiment of the present invention is shown. Detailed implementation manners

[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0061] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0062] Figure 1 The schematic flowchart of a medical visual question answering method based on causal reasoning and large language model according to an embodiment of the present invention is shown; Figure 2 The schematic diagram of the visual question answering network structure of the embodiment of the present invention is shown. Please refer to the attached Figure 1 and Figure 2 As shown, a medical visual question answering method based on causal reasoning and large language model includes the following steps:

[0063] S1. Extract initial visual feature f i and initial text feature f q .

[0064] In this step S1, the initial visual feature f i and the initial text feature f q are respectively extracted from the medical image and the question text through the basic network module. These features may include local and global features of the image and semantic information of the question. Among them, the basic network module is composed of a pre-trained image encoder and a text encoder, as Figure 1 shown.

[0065] In some embodiments, the image encoder extracts the features of a medical image as a feature vector of w×h×c, where w represents the width of the feature vector, h represents the height of the feature vector, and c represents the number of channels of the feature vector. The text encoder extracts the features of the input text as a feature vector of v×d, where v represents the size of the vocabulary and d represents the dimension of the word embedding.

[0066] S2. Perform a feature mapping operation on the initial visual feature f i to obtain the global visual feature f ig and the local visual feature f il .

[0067] In this step S2, specifically, a feature selector is used to perform a feature mapping operation on the initial visual feature f i to obtain the global visual feature f ig and the local visual feature f il , as Figure 3 shown. Among them, f ig provides information about the shape, position, and overall state of the organ. f il provides more detailed information, such as lesions, abnormalities, or other important details.

[0068] S3. Perform feature fusion on the local visual feature f il and the initial text feature f q to obtain the fused feature f iq .

[0069] Preferably, in this step S3, the local visual feature f il and the initial text feature f q are subjected to feature fusion through multi-head attention to obtain the fused feature f iq , and through this feature fusion, the association between the text and the local visual information can be made more explicit and tight.

[0070] Specifically, in some embodiments, the feature fusion formula is as follows:

[0071] f iq = MLP(MHA(f il , f q , f q ))

[0072] S4. Perform feature fusion on the fused feature f iq with the global visual feature f ig and the local visual feature f il to obtain the visual intermediate variable m i ; and perform feature fusion on the fused feature f iq with the initial text feature f qPerform feature fusion to obtain the text mediating variable m q .

[0073] In this step S4, according to the needs of causal inference theory, a mediating variable is first constructed. The mediating variable fully integrates feature information at different levels and from different sources. The above two mediating variables m i , m q provide a richer and more comprehensive input for subsequent causal inference operations. Then, causal intervention is performed to eliminate the influence of confounding factors and achieve the causal effect. After the causal intervention operation, the true causal features can be obtained. The causal intervention formula is as follows:

[0074]

[0075] Among them, P(A|do(I,Q)) represents the total causal effect of the treatment variables, i.e., the image I and the question Q, on the outcome variable, i.e., the answer A; the do algorithm represents the intervention operation; m is the mediating variable; i and q respectively represent the feature vectors of the image and the question text; P(m|I,Q) represents the conditional probability distribution of the mediating variable m given the image I and the question Q; the mediating variable represents the hidden layer features, and these features affect the answer A by extracting and combining the information of I and Q; P(A|m,i,q) represents the conditional probability of generating the answer A given the mediating variable m and the image I and the question Q, and this part expresses how the image and the question jointly affect the final answer under the action of the mediating variable; P(i,q) represents the joint probability distribution of the image I and the question Q, reflecting the natural distribution of the respective features of the image and the question without intervention.

[0076] Preferably, in some embodiments, in this step S4, the fused feature f iq is fused with the global visual feature f ig and the local visual feature f il through multi-head attention to obtain the visual mediating variable m i ; and the fused feature f iq is fused with the initial text feature f q through multi-head attention to obtain the text mediating variable m q . Specifically, the following formula for extracting the mediating variable is used for feature fusion:

[0077] m i = MLP(MHA(f il , f ig , f ig ), MHA(f il , f iq , f iq ))

[0078] m q = MLP(MHA(f q , f q , f q ), MHA(f iq , f q , f q ))

[0079] S5. Respectively perform causal intervention operations on the initial visual feature f i and the visual intermediate variable m i , and the initial text feature f q and the text intermediate variable m q to obtain the visual causal feature f i ' and the text causal feature f' q .

[0080] Preferably, in some embodiments, the causal intervention operation in step S5 adopts front-door adjustment. By intervening in the original feature and the intermediate variable through front-door adjustment, the deconfounding process of causal inference can be achieved. Among them, the formula for front-door adjustment is as follows:

[0081] f i ' = FDA(f i , m i ), f' q = FDA(f q , m q )

[0082] Among them, FDA represents the front-door adjustment process.

[0083] S6. Use the pre-trained large language model to generate corresponding question-and-answer pairs according to the training samples, as the prompt content for the question-and-answer process, providing a rich context and reference for the question-and-answer process.

[0084] Specifically, in step S6, generate question-and-answer pairs according to the training samples input to the model and the template type; among them, the template type is determined by the type of the question and is divided into open-ended questions and closed-ended questions. Provide prompt content according to the input samples to assist the causal inference process, thereby improving the effectiveness in generating answers.

[0085] S7. Input the visual causal feature f i ', the text causal feature f' q and the text feature f p of the prompt content into the large language model to generate medical question-and-answer results.

[0086] In this step S7, the visual causal feature f i ' and the text causal feature f' qText feature f of the prompt content p Input into the large language model for the medical visual question - answering process. By combining causal features with rich context information, the causal effect is enhanced, the accuracy of question understanding and the quality of answers are improved, thereby enhancing the question - answering performance.

[0087] Preferably, in some embodiments, in step S7, different cross - entropy losses are selected for closed - ended questions and open - ended questions, and are defined as follows:

[0088]

[0089] Where Lc is the cross - entropy loss for closed - ended questions, Lo is the cross - entropy loss for open - ended questions, a t represents the true answer, T represents the training data, F() represents the overall question - answering framework, N represents the length of the answer, a 1:n-1 represents the partial output, p t represents the prompt text, θ represents the model parameters; meanwhile, in order to ensure the consistency between the prediction of the original features and the prediction of the causal features by the model, the following loss function is defined:

[0090] L cau =KL(F(f i ',f' q ),F(f i ,f q ))

[0091] The combined loss function is defined as:

[0092] L closed =L c +L cau ,L open =L o +L cau

[0093] According to the above embodiments, the present invention proposes a medical visual question - answering method that combines causal reasoning with a large language model, comprehensively and accurately analyzes and interprets the causal relationship between medical images and relevant text information, improves the performance and accuracy of the medical visual question - answering system. By making full use of the multi - modal features of medical images and text information, the question - answering system can understand questions more comprehensively, thereby providing more in - depth and reliable answers. By emphasizing causal reasoning, confounding factors in the medical visual question - answering process are effectively removed, and the system's ability to accurately understand and analyze medical questions is improved. By combining causal reasoning with a large language model, the parsing and answering ability of the medical visual question - answering system for complex medical questions is improved, making the question - answering results more accurate and in - depth.

[0094] The above is the introduction of the method embodiments. The following further illustrates the solution of the present disclosure through device embodiments.

[0095] Figure 7 Fig. shows a schematic block diagram of a medical visual question answering device based on causal reasoning and large language models according to an embodiment of the present invention. As Figure 7 shown, a medical visual question answering device based on causal reasoning and large language models includes:

[0096] A basic network module 11 for respectively extracting an initial visual feature f i and an initial text feature f q from a medical image and a question text; wherein, as Figure 2 shown, the basic network module 11 includes a pre-trained image encoder and a text encoder.

[0097] A feature selector 12 for performing a feature mapping operation on the initial visual feature f i to obtain a global visual feature f ig and a local visual feature f il , as Figure 4 shown.

[0098] An information integration module 13 for fusing the local visual feature f il with the initial text feature f q to obtain a fused feature f iq .

[0099] An intermediate variable extraction module 14 for fusing the fused feature f iq with the global visual feature f ig and the local visual feature f il to obtain a visual intermediate variable m i ; and fusing the fused feature f iq with the initial text feature f q to obtain a text intermediate variable m q , as Figure 3 and Figure 4 shown.

[0100] A causal intervention operation module 15 for respectively performing causal intervention operations on the initial visual feature f i and the visual intermediate variable m i , the initial text feature f q and the text intermediate variable m q to obtain a visual causal feature f i ' and a text causal feature f' q .

[0101] Specifically, in some embodiments, such as Figures 2 to 4 shown, the causal intervention operation module 15 includes two front-door adjustment units, and each front-door adjustment unit is composed of two attention fusion modules; the causal intervention operation module 15 respectively inputs the initial visual feature f i and the visual mediator variable m i , the initial text feature f q and the text mediator variable m q into the two front-door adjustment units for causal intervention operations to obtain the visual causal feature f i ' and the text causal feature f' q .

[0102] The hint module 16, which is composed of pre-trained visual and language backbones, uses a pre-trained large language model to generate corresponding question-answer pairs according to training samples and template structures as the hint content for the question-answer process, providing a rich context and reference for the question-answer process; among them, the hint module 16 provides hint content according to the input sample to assist the causal reasoning process, thereby improving the effectiveness in generating answers.

[0103] The visual question-answer generation module 17 is used to input the visual causal feature f i ', the text causal feature f' q and the text feature f p of the hint content into the large language model to generate medical question-answer results.

[0104] The causal reasoning module (the mediator variable extraction module 14 and the causal intervention operation module 15) and the hint module 16 provided in this example can be added to the existing medical visual network framework, thereby effectively improving the accuracy of model question-answering without adding too much computational cost.

[0105] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0106] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0107] To verify the effectiveness of the present invention, the following test experiments were specifically conducted. Test configuration: The hardware environment for the experiments in this article is AMD Ryzen 9 5900k + NVIDIA GeForce RTX 3090 24GB, and the software environment is Windows 10 x64 + CUDA 12.2 + CuDNN 7.1 + Pytorch 2.1.1 + Python 3.8.

[0108] Dataset: The datasets used in the experiments of the present invention are the VQA-RAD, SLAKE, and PathVQA datasets. These datasets contain rich medical images and corresponding question-answer pairs. Table 1 presents the basic information of the datasets selected by the present invention.

[0109] Table 1 Basic information of the datasets selected by the present invention

[0110]

[0111] The effects produced by the present invention are as Figure 6 shown, demonstrating that by adding a causal reasoning module and a prompting module, the interference of confounding factors is effectively eliminated, and the question-answering effect of the model is better. Additionally, as Figure 5 shown, we can see that compared with other methods, the medical visual question-answering method proposed by the present invention, which combines causal reasoning with a large language model, effectively removes the interference of confounding factors in the medical visual question-answering process and achieves advanced results in three large public datasets. It can also be applied to other medical image processing and natural language processing tasks.

[0112] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed by the present invention can be achieved. No limitations are imposed herein.

[0113] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A medical visual question answering method based on causal reasoning and large language models, characterized in that Including the following steps: S1. Extract initial visual feature f and initial text feature f from medical images and problem texts respectively i q ;​ S2. Perform a feature mapping operation on the initial visual feature f i to obtain the global visual feature f ig and the local visual feature f il ; S3. Fuse the local visual feature f il with the initial text feature f q to obtain a fused feature f iq ; S4. Fuse the fused feature f iq with the global visual feature f ig and the local visual feature f il to obtain a visual intermediate variable m i ; and fuse the fused feature f iq with the initial text feature f q to obtain a text intermediate variable m q ; S5. Respectively perform causal intervention operations on the initial visual feature f i and the visual mediating variable m i , the initial text feature f q and the text mediating variable m q to obtain the visual causal feature f i ' and the text causal feature f q '; S6. Using a pre-trained large language model to generate corresponding question-and-answer pairs based on training samples as the prompt content for the question-and-answer process, providing a rich context and reference for the question-and-answer process; S7. Input the visual causal feature f i ', the text causal feature f q ', and the text feature f of the prompt content p into the large language model to generate medical Q&A results.

2. The method according to claim 1, wherein: In step S1, it includes: Extracting features of the medical image into a feature vector of w×h×c, where w represents the width of the feature vector, h represents the height of the feature vector, and c represents the number of channels of the feature vector; Extracting features of the question text into a feature vector of v×d, where v represents the size of the vocabulary and d represents the dimension of the word embedding.

3. The method according to claim 1, wherein: In step S2, the global visual feature f ig provides information about the shape, position, and overall state of the organ; the local visual feature f il provides detailed information, and the detailed information includes lesions and abnormalities.

4. The method according to claim 1, wherein: In step S3, the feature fusion formula is as follows: f iq = MLP(MHA(f il , f q , f q ))。 5. The method according to claim 1, wherein: In step S4, the following formula for extracting the mediating variable is used for feature fusion: m i = MLP(MHA(f il , f ig , f ig ), MHA(f il , f iq , f iq )) m q = MLP(MHA(f q , f q , f q ), MHA(f iq , f q , f q ))。 6. The method according to claim 1 or 5, characterized in that: In step S5, the causal intervention formula is as follows: Where P(A|do(I,Q)) represents the total causal effect of the treatment variables, i.e., the image I and the question Q, on the outcome variable, i.e., the answer A; the do algorithm represents the intervention operation; m is the mediating variable; i and q respectively represent the feature vectors of the image and the question text; P(m|I,Q) represents the conditional probability distribution of the mediating variable m given the image I and the question Q; P(A|m,i,q) represents the conditional probability of generating the answer A given the mediating variable m and the image I and the question Q; P(i,q) represents the joint probability distribution of the image I and the question Q.

7. The method according to claim 6, wherein: In step S5, the causal intervention operation adopts front-door adjustment, and the formula for the front-door adjustment is as follows: f i ' = FDA(f i , m i ), f q ' = FDA(f q , m q ) Where FDA represents the front-door adjustment process.

8. The method according to claim 1, characterized in that: Where In step S6, the generating of corresponding question-and-answer pairs using a pre-trained large language model based on training samples includes: generating question-and-answer pairs according to the training samples input to the model and the template type; Where the template type is determined by the type of the question and is divided into open-ended questions and closed-ended questions; Where in step S7, different cross-entropy losses are selected for closed-ended questions and open-ended questions, and are defined as follows: Among them, Lc is the cross-entropy loss of the closed-ended question, Lo is the cross-entropy loss of the closed-ended question, a t represents the true answer, T represents the training data, F() represents the overall question-answering framework, N represents the length of the answer, a 1:n-1 represents the partial output, p t represents the prompt text, θ represents the model parameters; at the same time, in order to ensure the consistency of the model's prediction of the original features and the causal features, the following loss function is defined: L cau = KL(F(f i ', f q '), F(f i , f q )) The joint loss function is defined as: L closed = L c + L cau , L open = L o + L cau。 9. A medical visual question answering device based on causal reasoning and large language models, characterized in that, Including: The basic network is used to extract initial visual feature f from medical images and problem texts respectively i and initial text feature f q ; wherein, the basic network includes a pre-trained image encoder and a text encoder; A feature selector for performing a feature mapping operation on the initial visual feature f i to obtain a global visual feature f ig and a local visual feature f il ; An information integration module for integrating the local visual feature f il with the initial text feature f q to perform feature fusion and obtain a fused feature f iq ; The mediating variable extraction module is used to combine the fused feature f iq with the global visual feature f ig and the local visual feature f il to perform feature fusion and obtain the visual mediating variable m i ; and combine the fused feature f iq with the initial text feature f q to perform feature fusion and obtain the text mediating variable m q ; Causal intervention operation module, which is used to perform causal intervention operations on the initial visual feature f i and the visual mediation variable m i , the initial text feature f q and the text mediation variable m q respectively, so as to obtain the visual causal feature f i ' and the text causal feature f q '; A prompt module, which is composed of pre-trained visual and language backbones, uses a pre-trained large language model to generate corresponding question-and-answer pairs according to training samples and template structures as the prompt content for the question-and-answer process, providing a rich context and reference for the question-and-answer process; A visual question-answering generation module for inputting the visual causal feature f i ', the text causal feature f q ', and the text feature f p of the prompt content into a large language model to generate medical Q&A results.

10. The device according to claim 9, wherein: Where The causal intervention operation module includes two front-door adjustment units, and the front-door adjustment unit is composed of two attention fusion modules; The causal intervention operation module respectively inputs the initial visual feature f i and the visual mediation variable m i , the initial text feature f q and the text mediation variable m q into two front-door adjustment units for causal intervention operations, and obtains the visual causal feature f i ' and the text causal feature f q '.