Method and System for Constructing an Interpretable Visual Question Answering Model Based on Factual Scenarios
Through the interpretability visual question-and-answer model construction method based on fact scenes, the visual question-and-answer model is iteratively updated using weight backpropagation and open source machine learning library, which solves the interpretability problem of deep learning models in the field of medical image question-and-answer field, and improves the model's reasoning ability and interpretability.
Patent Information
- Application Number
- CN202211623149.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-16
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-12-16
AI Technical Summary
Existing deep learning models cannot explain the output results of the post-learning model in the field of medical image question and answers, and lack interpretability.
Through the interpretable visual question-and-answer model construction method based on fact scenes, the visual question-and-answer model is iteratively updated using weight backpropagation and open source machine learning library, image and text features are extracted, and interpretable visual question-and-answer model is generated.
It enhances the model's inference ability, can capture image and text features more granularly, improves the interpretability of the model, and reduces the possibility of wrong answers.
Smart Images

Figure CN116306681B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of visual question answering, and more particularly, to a method and system for constructing an interpretable visual question answering model based on a factual scenario. Background Art
[0002] In recent years, computer vision and natural language processing have made great progress in the fields of images and texts respectively. The research field combining the two is the field of visual question answering. The purpose of the Visual Question Answering (VQA) task is to predict the answer to a question associated with a given image. Inspired by VQA, the exploration of medical VQA has attracted much attention in recent years. Medical VQA means that both the source of the image and the question are medical images and clinical medical questions related to the images. Recent research has shown that interpretability determines the accuracy of the predicted answer, and medical VQA has a higher requirement for interpretability compared to general-domain VQA because incorrect answer prediction may bring catastrophic consequences.
[0003] However, although there are already relevant studies on the interpretability techniques of neural network CNN and recurrent neural network RNN, there are few studies on the interpretability in the field of visual question answering, especially in the field of medical image question answering. For example, when asking "What abnormalities are there in this medical image?" or "How many abnormalities are there in the image?", there should be a reliable interpretability method to verify the predicted answer. Such verification should be based on the overall medical VQA system, rather than just a visualization display of the attention mechanism for the image and text. Such an interpretability method has not been fully explored. Therefore, it is necessary to study the interpretability techniques in the field of medical image question answering.
[0004] Causal reasoning can be used for the interpretability of the model. Most current deep learning models are learned in a data-driven manner based on statistical models. Although this black-box method can directly learn the implicit correlations from the data, it cannot explain the results output by the model after learning.
[0005] In view of this, the present application is specifically proposed. Summary of the Invention
[0006] The technical problem to be solved by the present invention is that in the prior art, the deep learning model cannot explain the output result of the model after learning. The purpose is to provide a method and system for constructing an interpretable visual question answering model based on a factual scenario, which can achieve interpretability of the output result of deep learning after the model performs deep learning.
[0007] The present invention is achieved by the following technical solutions:
[0008] Method for constructing interpretable visual question answering model based on factual scenarios, the method steps include:
[0009] Obtain a first data set and a second data set, where the first data set is an image-text pair data set, and the second data set is a visual question answering data set;
[0010] Construct a visual question answering model, and pre-train the visual question answering model through the first data set to obtain an image feature extraction network and a text feature extraction network;
[0011] Process the image feature extraction network by using the weight backpropagation method to obtain image counterfactual samples;
[0012] Process the text feature extraction network by using an open-source machine learning library to obtain text counterfactual samples;
[0013] Introduce adversarial semi-factual samples of images and texts, and iteratively update the visual question answering model by combining the image counterfactual samples and the text counterfactual samples to obtain a visual question answering prediction model;
[0014] Extract the feature data in the second data set, and verify the visual question answering prediction model through the feature data to obtain an interpretable visual question answering model.
[0015] In the traditional field of visual question answering technology, most deep learning models learn based on the data-driven method of statistical models. Although this black-box method can directly learn the implicit correlations through data, it is impossible to explain the results output by the model after learning. The present invention provides a method for constructing an interpretable visual question answering model based on factual scenarios. By using weight backpropagation and an open-source machine learning library to extract relevant networks respectively, and continuously updating and iteratively optimizing the visual question answering model with the extracted networks, the problem of weak model interpretability in current visual question answering research is solved, enabling the model to preserve key causal information to enhance the model's reasoning ability and capture image features and text features more granularly.
[0016] Preferably, the sub-steps for obtaining the image feature extraction network and the text feature extraction network include:
[0017] In the visual question answering model, extract the image features in the first data set through the ResNet50 network to obtain image features;
[0018] Embed the question text words through the GloVe model, and then input the embedded model into a 1024D LSTM network to obtain text features;
[0019] Both the image features and the text features are processed through a bilinear attention network to obtain an image feature extraction network and a text feature extraction network.
[0020] Preferably, the sub-steps for obtaining the image counterfactual sample include:
[0021] Process the image feature extraction network using the Weighted Backpropagation (WBP) method to obtain a causal saliency map.
[0022] Combined with the L1 norm, approximate the pixel values in the causal saliency map to 0 to obtain the image counterfactual sample.
[0023] Preferably, the sub-steps for the text counterfactual sample include:
[0024] Process the text feature extraction network using the open-source machine learning library SHAP to obtain the importance score for each word in the question text associated with the image.
[0025] Combined with the L1 norm, uniformly replace the word with the highest score with MASK to obtain the text counterfactual sample.
[0026] Preferably, the pre-training is specifically: in the gradient calculation stage of the visual question answering model, optimize the symmetric loss function by using the cosine similarity.
[0027] Preferably, the sub-steps for obtaining the visual question answering prediction model include:
[0028] Derive the network layer parameters through the loss function of the original sample, the positive sample loss function, the counterfactual sample loss function, and the L1 norm, and backpropagate along the gradient to minimize the loss function value. Continuously iterate and update the relevant parameters to obtain the visual question answering prediction model.
[0029] Preferably, in the image-text pair dataset, the image-text data is the data composed of an image and its corresponding relevant question and answer, and the image-text dataset is a set composed of several image-text data.
[0030] The present invention also provides an interpretable visual question answering model construction system based on a factual scenario, including a data acquisition module, a pre-training module, a first processing module, a second processing module, an iterative update module, and a verification module;
[0031] The data acquisition module is used to acquire a first dataset and a second dataset. The first dataset is an image-text pair dataset, and the second dataset is a visual question answering dataset;
[0032] The pre-training module is used to construct a visual question answering model and pre-train the visual question answering model through the first data set to obtain an image feature extraction network and a text feature extraction network;
[0033] The first processing module is used to process the image feature extraction network by using the weight backpropagation method to obtain image counterfactual samples;
[0034] The second processing module is used to process the text feature extraction network by using an open-source machine learning library to obtain text counterfactual samples;
[0035] The iterative update module is used to introduce adversarial semi-factual samples of images and texts, and iteratively update the visual question answering model by combining the image counterfactual samples and the text counterfactual samples to obtain a visual question answering prediction model;
[0036] The verification module is used to extract feature data from the second data set and verify the visual question answering prediction model through the feature data to obtain an interpretable visual question answering model.
[0037] Preferably, the pre-training module includes an image feature extraction module, a text feature extraction module, and a network processing module.
[0038] The image feature extraction module is used to extract image features in the visual question answering model through the ResNet50 network for the image features in the first data set to obtain image features;
[0039] The text feature extraction module is used to embed question text words through the GloVe model and input the embedded model into a 1024D LSTM network to obtain text features;
[0040] The network processing module is used to process both the image features and the text features through a bilinear attention network to obtain an image feature extraction network and a text feature extraction network.
[0041] The present invention also provides a computer storage medium, on which a computing program is stored. When the computer program is executed by a processor, the method described above is implemented.
[0042] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0043] A method and system for constructing an interpretable visual question answering model based on a factual scenario provided by an embodiment of the present invention extract relevant networks through weight backpropagation and an open-source machine learning library respectively, and continuously update and iterate the visual question answering model with the extracted networks to solve the problem of weak model interpretability in current visual question answering research, enabling the model to preserve key causal information to enhance the model's reasoning ability and capture image features and text features more granularly. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0045] Figure 1 It is a schematic flowchart of the construction method;
[0046] Figure 2 It is a framework diagram of the visual question answering model;
[0047] Figure 3 It is a causal reasoning intervention strategy based on a factual scenario;
[0048] Figure 4 It is an interpretable reasoning effect diagram on a benchmark dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the embodiments and the drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and do not limit the present invention.
[0050] In the following description, a large number of specific details are set forth in order to provide a thorough understanding of the present invention. However, it is obvious to those of ordinary skill in the art that the present invention does not have to be practiced with these specific details. In other embodiments, well-known structures, circuits, materials, or methods are not specifically described to avoid obscuring the present invention.
[0051] Throughout the specification, references to "an embodiment", "embodiments", "an example" or "examples" mean that a particular feature, structure, or characteristic described in connection with the embodiment or example is included in at least one embodiment of the present invention. Thus, the phrases "an embodiment", "embodiments", "an example" or "examples" appearing throughout the specification do not necessarily all refer to the same embodiment or example. In addition, the specific features, structures, or characteristics may be combined in any suitable combination and / or sub-combination in one or more embodiments or examples. Further, those of ordinary skill in the art should understand that the diagrams provided herein are for illustrative purposes only and are not necessarily drawn to scale. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0052] In the description of the present invention, the orientation or positional relationship indicated by terms such as "front", "rear", "left", "right", "upper", "lower", "vertical", "horizontal", "high", "low", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as limiting the protection scope of the present invention.
[0053] Embodiment 1
[0054] In the field of traditional visual question answering technology, most deep learning models learn based on the data-driven method of statistical models. Although this black-box method can directly learn the implicit correlations through data, it is unable to explain the results output by the model after learning.
[0055] This embodiment discloses a method for constructing an interpretable visual question answering model based on a factual scenario. By using weight backpropagation and an open-source machine learning library to extract relevant networks respectively, and then continuously updating and iteratively optimizing the visual question answering model with the extracted networks, it solves the problem of weak model interpretability in current visual question answering research, enables the model to preserve key causal information to enhance the model's reasoning ability, and captures image features and text features more granularly. The flow schematic diagram of the construction method of this embodiment is as Figure 1 shown, and the method steps include:
[0056] S1: Obtain a first data set and a second data set. The first data set is an image-text pair data set, and the second data set is a visual question answering data set; in the image-text pair data set, the image-text data is data composed of an image and its corresponding related questions and answers, and the image-text data set is a set composed of several image-text data. This embodiment takes the obtained medical image visual question answering model as an example.
[0057] S2: Construct a visual question answering model, and pre-train the visual question answering model through the first data set to obtain an image feature extraction network and a text feature extraction network;
[0058] The sub-steps obtained by the image feature extraction network and the text feature extraction network include: in the visual question answering model, extract the image features in the first data set through the ResNet50 network to obtain image features; embed the question text words through the GloVe model, and then input the embedded model into the 1024D LSTM network to obtain text features; process both the image features and the text features through a bilinear attention network to obtain an image feature extraction network and a text feature extraction network. The pre-training is specifically: in the gradient calculation stage of the visual question answering model, optimize the symmetric loss function by using cosine similarity.
[0059] After the medical image is sent into the model, it will first enter the ResNet50 network for image feature extraction. After the question text is word-embedded by the GloVe model, the size of each word vector is 300 dimensions. It is sent into the 1024D LSTM network to generate question text features. The forward propagation equation of the LSTM unit with a forget gate is:
[0060] f t =σ(W fx x t +W fh h t-1 +b f )
[0061] i t =(W ix x t +W ih h t-1 +b i )
[0062] o t =σ(W ox x t +W oh h t-1 b o )
[0063]
[0064]
[0065] h t =o t ⊙σ(c t )
[0066] where f t, i t , o t are respectively the forget gate, input gate, and output gate of the control state. W and b are the weight biases of the three gates, and c t is the cell state of the LSTM.
[0067] For feature fusion, the present invention uses a bilinear attention network to fuse visual information and language information, and the joint representation of the fused features is:
[0068]
[0069]
[0070] where U and V are linear embeddings, p is a learnable mapping vector. For the visual question answering model, the overall framework is as Figure 2 shown.
[0071] S3: Process the image feature extraction network using the weight backpropagation method to obtain an image counterfactual sample;
[0072] The sub-steps for obtaining the image counterfactual sample include: processing the image feature extraction network using the weight backpropagation WBP method to obtain a causal saliency map; combining the L1 norm, approximately replacing the pixel values in the causal saliency map with 0 to obtain the image counterfactual sample.
[0073] The specific process is as follows: Design a causal intervention strategy to generate counterfactual causal examples during model training to strengthen causal relevance. Given an image input x, the causal saliency mapping s m (x) with the answer label y = m, where n = 1,..., M, M is the number of classes. Causal intervention is to remove the causal information in x contained in s m (x) (replacing the significant pixel values with zeros), and then use it as the counterfactual causal example in the image modality; given a question text input t, the causal saliency mapping s m (t) with the answer label y = m. Causal intervention is to remove the causal information in t contained in s m (t) (replacing the significant words with [MASK]), and then use it as the counterfactual causal example label in the text modality now.
[0074] To generate a saliency map in the original picture pixel space and provide information for decision-driven features, the following describes weight backpropagation, which is a new and efficient computational saliency map scheme applicable to any neural architecture and can evaluate the contribution of each pixel to the final class-specific prediction.
[0075] Consider a vector input and a linear mapping. Let x lis the internal representation of the data of the l-th layer, where l = 0 is the input layer, i.e., x 0 = x, and l = L is the second-to-last logit layer before the softmax transformation, i.e., To assign relative importance to each hidden unit in the l-th layer, symbolically decompose all the transformations after l into an operator denoted by which is called the saliency matrix and satisfies: where x L is an M-dimensional vector corresponding to M different classes in y. Although represented in a matrix form, there is a slight abuse of notation. For example, the instantiation of the operator effectively depends on the input x, so all non-linearities are effectively absorbed into it. For an object related to a given label y = m, its causal features are contained in the interaction between the m-th row of and the input x, i.e.: where s m (x) k represents the k-th element of the saliency map s m (x), is a single element of. Calculating a key observation is that it can be done recursively. Specifically, let g l (x l ) be the transformation of the l-th layer, such as affine transformation, convolution, activation function, normalization, etc., then there is:
[0076] This means that can be recursively calculated in the following way
[0077] where G(·) is the update rule. The update rules for common transformations in deep networks are listed in Table I.
[0078]
[0079] In this embodiment: Restrict the image occlusion operation so that the local image of the replaced part (i.e., the causal part of the image that affects the model output) is as small as possible. Any causal saliency map that satisfies the causal relationship, regardless of the size of the occluded part, is a valid occluded causal saliency map. Occluding the entire image and only occluding the lesion part actually cover up the causal relationship, which is not conducive to the interpretability of the model. To avoid this situation, use the L1 norm to encourage the causal part of each image to only account for a very small part of the entire image.
[0080] S4: Process the text feature extraction network using an open-source machine learning library to obtain text counterfactual samples;
[0081] The text counterfactual sample sub-step includes: using the open-source machine learning library SHAP to process the text feature extraction network to obtain the importance score of each word in the question text associated with the image; combining the L1 norm, and replacing the word with the highest score with MASK by agreement to obtain the text counterfactual sample.
[0082] Generate the saliency map for the original question text. SHapley SHAP is a general-purpose model interpretability framework. It was proposed and created inspired by game theory. Classical methods include Shapley regression values and Shapley sampling values. Shapley regression values retrain the model on the feature subsets when calculating the feature contributions. For feature i, first generate all feature sets that include i and exclude i, then retrain and calculate the prediction results, and thus calculate the average of the contributions of feature i:
[0083]
[0084] Shapley sampling values avoid the process of repeatedly training new models and approximate the above formula by sampling. And Quantitative Input Influence is a more general algorithmic interpretation framework, where the part of the feature contributions is still approximated by sampling to obtain Shapley values. Specify the interpretation model as: where g is the interpretation model, z' ∈ {0, 1} M is the joint vector, and M is the maximum length of the vector. is the contribution of feature j (Shapley values). The joint vector characterizes which feature combinations the selected data points have. 0 represents not including the feature, and 1 represents including the feature.
[0085] In this embodiment, Step2 and Step3 also include: in order to remove causal information from the picture input x i and the question text input t i and obtain the counterfactual sample and This method applies the following occlusion method: where T() is the masking function: where ω, ω, σ > 0 are the threshold and scaling parameters, which simply control the range and pixel values of the occlusion. And define the following objective function: where f θ is the prediction model, is the counterfactual sample loss function to be optimized, The text counterfactual sample of the text vector replaces the element with the highest score in the vector with 0, represents the flipping of the class label, i.e., l(x, t, y; f θ ) = -l(x, t, y; f θ ).
[0086] Note that the above objective function may lead to degenerate solutions. That is, for any causal significant graph that satisfies the causal relationship, regardless of the size of the occlusion area, it is a causal significant graph of effective occlusion. Occluding the entire image and only the lesion area actually covers up the causal relationship, which is not conducive to the interpretability of the model. To avoid this situation, the L1 norm is used to encourage only a small part of the causal part of each image to account for the entire image: L reg = ||s(·)||1.
[0087] S5: Introduce adversarial semi-factual samples of images and texts, and iteratively update the visual question answering model by combining the image counterfactual samples and the text counterfactual samples to obtain a visual question answering prediction model;
[0088] The sub-steps for obtaining the visual question answering prediction model include: taking the derivative of the network layer parameters through the loss function of the original sample, the positive sample loss function, the counterfactual sample loss function, and the L1 norm, and backpropagating along the gradient to minimize the loss function value, and continuously iteratively updating the relevant parameters to obtain the visual question answering prediction model.
[0089] To avoid the interference brought by the intervention strategy itself, that is, the model does not learn to capture causal correlations, but learns to predict the intervention operations (occluding pictures) given to it. For example, when the model detects that the input is censored, it can learn to change the prediction regardless of whether the image lacks causal features, which may affect the discrimination result. Therefore, an adversarial control group is introduced to randomly occlude the non-causal associated parts of the image and the question to obtain semi-factual samples x i ' and t i ', x’ i = x i - T(s m (x j )) ⊙ x i , i ≠ j
[0090] such as Figure 3As shown, according to the assumption of causal relationship, the counterfactual samples obtained after causal intervention will predict wrong answers, while the factual samples and semi-factual samples of the original input will predict correct answers. The derivative of the network layer parameters is calculated through different loss functions and propagated backward along the gradient to minimize the loss function value for parameter update. At the same time, the causal saliency map obtained through the weight backpropagation technique will become more accurate gradually with the deepening of training, which is the explanation in the model training process.
[0091] In this embodiment, the loss function specifically includes four loss functions, namely the loss function for original sample classification, the loss function for positive sample classification, the loss function for counterfactual sample classification, and the L1 norm function in S2. It should be noted that the loss function of the counterfactual sample should be negative because the classification of the counterfactual sample is the result of causal intervention and cannot be classified into the correct category. The objective function of the semi-factual sample is: The objective function of the original sample is: The total objective function is: L = L Cls -L Neg +L reg +L Pos , thus optimizing the model parameters through the total objective function can help the model capture the causal relationship in the samples and make the model have stronger interpretability.
[0092] S6: Extract the feature data in the second dataset, and verify the visual question answering prediction model through the feature data to obtain an interpretable visual question answering model.
[0093] In this embodiment, specifically as Figure 4 shown, this method also improves the interpretability of the model. In Figure 4 , visualization techniques are used to describe the test process to reveal the interpretability of the proposed method. First, the answer distributions of two specific question patterns are compared, and then the most important regions are shown on the test input using feature maps.
[0094] In Figure 4 , in the first row, this method shows the ability to capture the causal relationship for the question pattern "is there an abnormal". This is a closed-ended question with "yes" or "no" candidate answers, and most of the closed-ended questions in the train set have "no" answers. For the test input from VQA-RAD, there is an abnormality in the shoulder bone density (red rectangle). Due to the distribution imbalance, the baseline method almost always answers "no", while the method proposed in this patent outputs about 80% "yes". The method of the present invention seems to infer the abnormality of the shoulder bone density by accurately locating the correct region, while the baseline model gets the wrong answer because it does not see the abnormal region in the image. This unsatisfactory performance may be due to language bias.
[0095] Furthermore, in Figure 4 the second row of, a similar situation also occurred for the question type "What is abnormal in the CT scan?". More than 50% of the answers in the training set were "cystic teratoma", and only 10% of the answers were "colon cancer". For the test input from SLAKE, there was a tumor abnormality in the colon region, and the method of the present invention accurately identified the lesion by capturing the causal part of the model. However, the lesion obtained by the baseline model was incorrect. In terms of answer prediction, the baseline model seemed to only derive "cystic teratoma" from the answer distribution of the training set, while the method proposed in this patent inferred the correct answer "colon cancer" based on the correct lesion, although the distribution of "colon cancer" in the training set was low. These two examples demonstrate that the method of the present invention is effective for various Med-VQA datasets, especially for Med-VQA datasets with language bias.
[0096] The factual scenario is that when people conduct causal reasoning on events occurring around them, they often have a thought process of "if a certain condition is changed, then the result will not occur (if…then…)" or "even if a certain condition is not changed, the result will still occur (but…for…)". This kind of thinking activity of negating an already occurred event and constructing another possible hypothesis is called counterfactual thinking. The term "counterfactual" can be abstractly explained as that an event may have different results under different conditions. Corresponding to this, there are also semi-factual and factual. Table 4.1 vividly explains these three situations with an example of bank lending.
[0097] A method for constructing an interpretable visual question answering model based on a factual scenario disclosed in this embodiment solves the problem of weak model interpretability existing in current visual question answering research, enables the model to preserve key causal information to enhance the model's reasoning ability, and captures image features and text features more granularly. In the embodiment, the model is tested with the benchmark datasets VQA-RAD and SLAKE, and the model of the present invention has achieved competitive results, especially good results in open-ended questions, and the interpretability of the model is also applicable to visual question answering models in other fields.
[0098] Embodiment 2
[0099] This embodiment discloses a system for constructing an interpretable visual question answering model based on a factual scenario. This embodiment is to implement the construction method in Embodiment 1, including a data acquisition module, a pre-training module, a first processing module, a second processing module, an iterative update module, and a verification module;
[0100] The data acquisition module is used to acquire a first data set and a second data set. The first data set is an image-text pair data set, and the second data set is a visual question answering data set;
[0101] The pre-training module is used to construct a visual question answering model and pre-train the visual question answering model through the first data set to obtain an image feature extraction network and a text feature extraction network;
[0102] The first processing module is used to process the image feature extraction network by using the weight backpropagation method to obtain an image counterfactual sample;
[0103] The second processing module is used to process the text feature extraction network by using an open-source machine learning library to obtain a text counterfactual sample;
[0104] The iterative update module is used to introduce adversarial semi-factual samples of images and texts, and iteratively update the visual question answering model by combining the image counterfactual sample and the text counterfactual sample to obtain a visual question answering prediction model;
[0105] The verification module is used to extract feature data from the second data set and verify the visual question answering prediction model through the feature data to obtain an interpretable visual question answering model.
[0106] The pre-training module includes an image feature extraction module, a text feature extraction module, and a network processing module.
[0107] The image feature extraction module is used to extract image features from the first data set through the ResNet50 network in the visual question answering model to obtain image features;
[0108] The text feature extraction module is used to embed question text words through the GloVe model and input the embedded model into a 1024D LSTM network to obtain text features;
[0109] The network processing module is used to process both the image features and the text features through a bilinear attention network to obtain an image feature extraction network and a text feature extraction network.
[0110] Embodiment III
[0111] This embodiment discloses a system for constructing an interpretable visual question answering model, on which a computing program is stored. When the computer program is executed by a processor, the method described in Embodiment I is implemented.
[0112] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0113] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program issued instructions. These computer program issued instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the issued instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0114] These computer program issued instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the issued instructions stored in the computer-readable memory generate a manufactured article including the issued instruction means, and the issued instruction means realizes the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0115] These computer program issued instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the issued instructions executed on the computer or other programmable devices provide steps for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0116] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for constructing an interpretable visual question answering model based on factual scenarios, characterized in that The method steps include: Obtain a first data set and a second data set, where the first data set is an image-text pair data set, and the second data set is a visual question answering data set; Construct a visual question answering model, and pre-train the visual question answering model through the first data set to obtain an image feature extraction network and a text feature extraction network; Process the image feature extraction network using the weight backpropagation method to obtain image counterfactual samples; The sub-steps for obtaining the image counterfactual samples include: Process the image feature extraction network using the weight backpropagation WBP method to obtain a causal saliency map; Combined with the L1 norm, approximate the pixel values in the causal saliency map to 0 to obtain the image counterfactual samples; Process the text feature extraction network using an open-source machine learning library to obtain text counterfactual samples; The sub-steps of the text counterfactual samples include: Process the text feature extraction network using the open-source machine learning library SHAP to obtain the importance scores of each word in the question text associated with the image; Combined with the L1 norm, uniformly replace the word with the highest score with MASK to obtain the text counterfactual samples; Introduce adversarial semi-factual samples of images and texts, and iteratively update the visual question answering model in combination with the image counterfactual samples and the text counterfactual samples to obtain a visual question answering prediction model; Extract the feature data in the second data set, and verify the visual question answering prediction model through the feature data to obtain an interpretable visual question answering model.
2. The method for constructing an interpretable visual question answering model based on a factual scenario according to claim 1, wherein The sub-steps for obtaining the image feature extraction network and the text feature extraction network include: In the visual question answering model, extract the image features in the first data set through the ResNet50 network to obtain image features; Embed the question text words through the GloVe model, and then input the embedded model into a 1024D LSTM network to obtain text features; Process both the image features and the text features through a bilinear attention network to obtain an image feature extraction network and a text feature extraction network.
3. The method for constructing an interpretable visual question answering model based on a factual scenario according to claim 1, wherein The pre-training is specifically: in the gradient calculation stage of the visual question answering model, optimize the symmetric loss function by using cosine similarity.
4. The method for constructing an interpretable visual question answering model based on a factual scenario according to claim 1, wherein The sub-steps for obtaining the visual question answering prediction model include: Derive the network layer parameters through the loss function of the original sample, the positive sample loss function, the counterfactual sample loss function, and the L1 norm, and backpropagate along the gradient to minimize the loss function value, and continuously iterate and update the relevant parameters to obtain a visual question answering prediction model.
5. The method for constructing an interpretable visual question answering model based on a factual scenario according to any one of claims 1 to 4, characterized in that In the image-text pair data set, the image-text data is data composed of an image and its corresponding related questions and answers, and the image-text data set is a set composed of several image-text data.
6. A system for constructing an interpretable visual question answering model based on factual scenarios, characterized in that, It includes a data acquisition module, a pre-training module, a first processing module, a second processing module, an iterative update module, and a verification module; The data acquisition module is used to acquire a first data set and a second data set. The first data set is an image-text pair data set, and the second data set is a visual question answering data set; The pre-training module is used to construct a visual question answering model and pre-train the visual question answering model through the first data set to obtain an image feature extraction network and a text feature extraction network; The first processing module is used to process the image feature extraction network by using the weight backpropagation method to obtain an image counterfactual sample; The sub-steps for obtaining the image counterfactual sample include: Process the image feature extraction network by using the weight backpropagation WBP method to obtain a causal saliency map; Combine the L1 norm and approximately replace the pixel point values in the causal saliency map with 0 to obtain the image counterfactual sample; The second processing module is used to process the text feature extraction network by using an open-source machine learning library to obtain a text counterfactual sample; The sub-steps of the text counterfactual sample include: Process the text feature extraction network by using the open-source machine learning library SHAP to obtain the importance score of each word in the question text associated with the image; Combine the L1 norm and uniformly replace the word with the highest score with MASK to obtain the text counterfactual sample. The iterative update module is used to introduce the adversarial semi-factual samples of the image and the text, and iteratively update the visual question answering model by combining the image counterfactual sample and the text counterfactual sample to obtain a visual question answering prediction model; The verification module is used to extract the feature data in the second data set and verify the visual question answering prediction model through the feature data to obtain an interpretable visual question answering model.
7. The system for constructing an interpretable visual question answering model based on a factual scenario according to claim 6, wherein The pre-training module includes an image feature extraction module, a text feature extraction module, and a network processing module. The image feature extraction module is used to extract the image features in the first data set through the ResNet50 network in the visual question answering model to obtain image features; The text feature extraction module is used to embed the question text words through the GloVe model, and then input the embedded model into a 1024D LSTM network to obtain text features; The network processing module is used to process both the image features and the text features through a bilinear attention network to obtain an image feature extraction network and a text feature extraction network.
8. A computer storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, the method described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Visual question and answer method based on capsule self-guide collaborative attention mechanism
CN113515615A
Image and text joint embedded multi-modal cultural resource processing method
CN113516118A