A Visual Reasoning Method with Self-Evolving Inference Chains Based on a Consistency Self-Evaluation Strategy

The self-evolving reasoning framework with a consistency evaluation strategy addresses the challenges of black-box deep learning models by generating coherent explanations for visual language tasks without human annotations, improving model explanation reliability and efficiency.

CN117076621BActive Publication Date: 2025-07-15NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310832818.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-09
Publication Date
2025-07-15
Estimated Expiration
2043-07-09

AI Technical Summary

Technical Problem

Existing deep learning models for visual language tasks, such as image description and visual question answering, are often black-box systems, making it difficult to gain user trust, and current explanation methods face challenges like inaccurate explanations due to separate decision and explanation modules, lack of logical consistency, and the need for costly and time-consuming human annotations.

Method used

A self-evolving reasoning framework with a consistency evaluation strategy, comprising an answer-explanation prompt module and a self-criticism reinforcement module, generates detailed explanations without human annotations by using a CLIP encoder, multi-layer perceptrons, and sequence sampling to expand the search space, with answer scores acting as rewards.

Benefits of technology

The method enhances model self-explanation capabilities by automatically generating coherent and reliable explanations, achieving state-of-the-art performance on multiple benchmarks without relying on human-labeled data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117076621B_ABST
    Figure CN117076621B_ABST
Patent Text Reader

Abstract

The present invention discloses a self-evolving visual reasoning method for an inference chain based on a consistency self-evaluation strategy, which includes two parts. The first part is the answer-explanation prompt module, and the second part is the self-critical reinforcement module. In the first part, a basic answer template is first established to obtain a basic answer score, and then an explanation generation template is constructed. In the second part, the search space is expanded by introducing a sequential sampling algorithm, and a set of candidate explanations are generated. At the same time, the answer is regarded as a reward to encourage the model to output more detailed explanations. The framework of the present invention can benefit from a large number of question-answer pairs without manual annotation of explanations, thereby further improving the self-explanation ability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of language-image multimodal fusion, and particularly relates to a visual reasoning method with self-evolving inference chains. Background Art

[0002] Deep neural networks have made significant breakthroughs in various vision-language tasks, such as image captioning and visual question answering. Unfortunately, most of them are black-box systems, which makes it challenging to gain user trust. Explaining the decision-making process of deep neural networks is a long-standing fundamental problem. Some methods rely on attention mechanisms or gradient-based localization to obtain visual explanations, which can highlight the image regions that contribute more to the predicted answer. However, simple visualization cannot explain how these regions support the prediction of the answer. Instead, natural language explanation tasks can explain the model's decision-making process by generating natural language sentences. Language-based explanations are easier for users to understand and can also help researchers optimize the model structure.

[0003] Recently, some natural language explanation models have achieved good results. They can guide the model to generate natural language statements and explain how the model obtains the answer. It is mainly divided into two major categories of methods: post-hoc explanation methods and self-rationalization methods. These two paradigms are still limited by the following challenges: 1) For the first paradigm, since the decision-making model and the explanation part are two independent modules, it inevitably leads to inaccurate explanations of the inference process of the decision-making model. 2) Due to the lack of explicit logical relationship modeling, the direct self-rationalization framework has been proven to have problems of logical inconsistency. 3) The above strategies all require a large number of manually annotated explanations, which are both expensive and time-consuming to collect. Summary of the Invention

[0004] To overcome the deficiencies of the prior art, the present invention provides a visual reasoning method with self-evolving inference chains based on a consistency self-evaluation strategy, including two parts. The first part is the answer-explanation prompting module, and the second part is the self-criticism reinforcement module. In the first part, first establish a basic answer template to obtain the basic answer score, and then construct an explanation generation template. In the second part, introduce a sequence sampling algorithm to expand the search space and generate a set of candidate explanations. At the same time, regard the answer as a reward to encourage the model to output more detailed explanations. The framework of the present invention can benefit from a large number of question-answer pairs without manually annotated explanations, thereby further improving the model's self-explanation ability.

[0005] The technical solution adopted by the present invention to solve its technical problems includes the following steps:

[0006] Step 1: Given an image and a natural language question where qt denotes the t-th word, T is the length of the question, and W×H×3 represents the size of the image;

[0007] Use the classification token of the CLIP visual encoder to obtain the image features of image I; then, use a set of multi-layer perceptrons to convert the image features into The calculation formula is as follows:

[0008] v1, v2, …, v S = MLP(CLIP(I)) (1)

[0009] For question Q, each word q t will be mapped to the corresponding word embedding through a pre-trained image captioning model

[0010] Finally, obtain the image and question sequence Z:

[0011]

[0012] Step 2: For the given image and question sequence Z, use the natural language label "The answer is" to prompt the model to generate the correct answer. By concatenating the language prompt, obtain the basic answer template Z a = [Z; <answer>, where [;] and <answer>respectively represent the splicing operation and specific language prompts;

[0013] During the training process, the basic answer template Z a and the true answer label are fed into the VL model, where a n is the nth answer token, and N represents the length of the answer label;

[0014] Calculate the answer loss conditional on Z a in an autoregressive manner:

[0015]

[0016] where θ represents the parameters of the VL model; obtain the average probability as the basic answer score;

[0017] Use the natural language token "The reason is" to encourage the model to generate a free-text explanation that explains the generation template Z e = [Z; <reason>is constructed, wherein <reason>is a natural language token;

[0018] For the tokenized samples, follow Equation (3) to calculate the explanatory loss L in an autoregressive manner with cross-entropy loss e ;

[0019] Step 3: Use beam search to sample the top K words from the VL model probability distribution at each time step and keep these sequences with the highest probability; then, these generated sentences are integrated into the candidate explanations for each QA pair without manually annotated explanations, where and K represent the k-th rationale and the size of beam search, respectively;

[0020] For the tokenized samples, construct the corresponding set of explanations in the same way Finally, since the tokenized and untokenized samples are integrated into a mini-batch during training, simplify R f and R s to to indicate the candidate explanations for each sample;

[0021] Step 4: Adopt a reinforcement learning method to achieve end-to-end training. Given a sample and the corresponding candidate explanations R, design a new input template:

[0022]

[0023] where represents the template with the possible underlying explanation r k added; then, feed this template into the model to obtain the average probability of the answer i.e., the explanatory answer score, where A is the ground truth answer of the sample;

[0024] Meanwhile, use the average probability output by the base answer template as the base score, and calculate the reinforcement loss L by applying reinforcement learning r , and the gradient is calculated as follows:

[0025]

[0026] where p θ (r k ) is the probability of the k-th explanation;

[0027] Based on the above calculations, when the answer score is higher than the score of the base answer template , this gradient will tend to increase the probability of the k-th underlying explanation; finally, for the tokenized samples, append the manually annotated underlying explanations to the candidate explanations R and predict the answer through cross-entropy loss, i.e., L ea ;

[0028] During the training process, the overall loss function is expressed as follows:

[0029] L = L a + L e + L ea + λL r (6)

[0030] where λ is used to balance the cross - entropy loss and the reinforcement loss;

[0031] During inference, the model will first generate the rationale for the QA pair and then use the explanation template to obtain the corresponding answer.

[0032] Preferably, the dimension size c = 768 and the image sequence length S = 10.

[0033] The beneficial effects of the present invention are as follows:

[0034] The method of the present invention first uses a prompting mechanism to encourage the model to generate answers and candidate explanations. At the same time, a novel self - criticism reinforcement module is designed to convert the answer score into a reward to evaluate these possible reasons. In addition, the framework of the present invention can benefit from a large number of question - answer pairs without manual annotation of explanations, thereby further improving the model's self - explanation ability. Through automatic measurement and manual evaluation, the method of the present invention has reached the state - of - the - art level in multiple benchmark tests and provides a new paradigm for research in this direction. Brief Description of the Drawings

[0035] Figure 1 is the overall framework diagram of the present invention.

[0036] Figure 2 is the illustration of the natural language explanation results of the semi - supervised visual question - answering based on self - criticism learning in the embodiment of the present invention. Detailed Embodiment

[0037] The present invention will be further described below in conjunction with the drawings and embodiments.

[0038] The present invention proposes a new semi - supervised visual question - answering - natural language explanation method based on self - criticism learning. This method can model the logical relationship between answer - explanation pairs and evaluate the generated rationale through answer rewards. This strategy effectively improves the logical consistency and reliability of the explanations.

[0039] Such as Figure 1 As shown in the figure, the main modules of the technical solution of the present invention include the following: The first part is the answer-explanation prompt module, and the second part is the self-criticism reinforcement module. In the first part, first establish a basic answer template to obtain the basic answer score, and then construct an explanation generation template. In the second part, the search space is expanded by introducing a sequence sampling algorithm, and a set of candidate explanations are generated. At the same time, the answer is regarded as a reward to encourage the model to output more detailed explanations.

[0040] The natural language explanation method includes the following main steps:

[0041] Step 1: Given an image and a natural language question where q t represents the t-th word, T is the length of the question, and W×H×3 represents the size of the image. The goal is to predict the answer and generate a corresponding text-independent fundamental explanation. The CLIP visual encoder and a pre-trained image captioning model are used as the basic backbone. During pre-training, the VL model uses the image embeddings from CLIP as a prefix and fine-tunes the language model to generate image captions. The image I and the question Q are regarded as the prefix of the answer-explanation sequence. Specifically, first apply the "classification token" in ViT-B and CLIP to obtain the image features. Then, use a set of lightweight and simple multi-layer perceptrons to convert the image features into where the dimension size c = 768 and the image sequence length S = 10. The above calculation formula is as follows:

[0042] v1, v2, …, v S = MLP(CLIP(I)) (1)

[0043] Note that only the mapping network is updated during training, while the original visual encoder parameters from CLIP will remain frozen. For the question Q, each word q t will be mapped to the corresponding word embedding through the pre-trained image captioning model Finally, obtain the image and question sequence Z:

[0044]

[0045] Step 2: The prompting mechanism can maintain the same optimization objective between the pre-training task and the downstream task. Considering convenience and interpretability, handcrafted prompts are used as templates to enable the model to generate answers or explanations. First, a basic answer template is established to obtain the basic answer score, and then the visual question answering task is regarded as a generation task so that the model can generate answers without a predefined answer space. Specifically, for a given image and question sequence Z, the natural language token "the answer is" is used to prompt the model to generate the correct answer. By concatenating the language prompts, the basic answer template Z can be obtained. a = [Z; <answer>, where [;] and <answer>respectively represent the splicing operation and specific language prompts. During training, the basic answer template and the ground-truth answer label are fed into the VL model, where a n is the n-th answer token, and N represents the length of the answer label. Next, the answer loss conditioned on Z a is calculated in an autoregressive manner:

[0046]

[0047] where θ represents the parameters of the VL model. Additionally, based on the metrics of the ground-truth answer, the average probability is obtained as the basic answer score.

[0048] To generate a reasonable rationale, the language prompt "The reason is" is used to encourage the model to generate free-form text explanations. Similar to the basic answer template, the explanation generation template Z e =[Z; <reason>is constructed, wherein <reason>It is a natural language token. For tokenized samples, follow Equation (3) to calculate the explanation loss L in an autoregressive manner with cross-entropy loss. e Since there are no available human-annotated explanations, no loss is calculated for unlabeled QA samples.

[0049] Step 3: To obtain logically consistent rationales, expand the search space by introducing a sequential sampling algorithm and generate a set of candidate explanations. Additionally, the answer score is regarded as a reward to encourage the model to output more detailed explanations. It should be noted that the above operations will be implemented on both labeled and unlabeled samples.

[0050] Thanks to the language-based prompting strategy and pre-trained VL model, the explanation generation template can easily guide the model to generate human-readable sentences. Therefore, for unlabeled QA samples, directly apply the explanation generation template Z e to generate candidate rationales. More specifically, use beam search to sample the top K words from the probability distribution of the VL model at each time step and keep these sequences with the highest probability. Then, these generated sentences are integrated into the candidate explanations for each QA pair without human-annotated explanations, where and K represent the k-th rationale and the size of beam search respectively. Additionally, for labeled samples, a similar mechanism is used to construct the corresponding set of explanations Applying the above operations to all samples can bring a larger search space. The expanded search space provides more possibilities for generating reliable rationales. At the same time, it can avoid overfitting. For labeled samples, although the model can rely on human explanations for training, these labels are still one-sided and subjective. Using the sequential sampling strategy can prevent the model from overfitting to these specific annotations. Finally, since labeled and unlabeled samples are integrated into a mini-batch during training, simplify R f and R s to to indicate the candidate explanations for each sample.

[0051] Ideal fundamental explanations can help the model better infer the answer. It is considered that the answer score can be converted into a self-criticism reward to evaluate these candidate explanations. Considering that this is a non-differentiable operation, a reinforcement learning method is adopted to achieve end-to-end training. Specifically, given a sample and the corresponding candidate explanation R, a new input template is designed:

[0052]

[0053] where represents adding the possible fundamental explanation r k template. Then, this template is fed back into the model to obtain the average probability of the answer which is the explanatory answer score, where A is the ground truth answer of the sample. Meanwhile, the average probability output using the base answer template is used as the base score. By applying reinforcement learning, the gradient is calculated as follows:

[0054]

[0055] where p θ (r k ) is the probability of the k-th explanation. Based on the above calculation, when the answer score is higher than the score of the base answer template , this gradient will tend to increase the probability of the k-th fundamental explanation. Finally, for the labeled samples, the manually annotated fundamental explanations are appended to the candidate explanations R, and the answer is predicted through the cross-entropy loss, i.e., L ea .

[0056] During the training process, the overall loss function can be expressed as follows:

[0057] L = L a + L e + L ea + λL r (6)

[0058] where λ is used to balance these two different types of losses (i.e., cross-entropy loss and reinforcement loss). During inference, the model will first generate the rationale for the QA pair and then use the explanation template to obtain the corresponding answer. Specific embodiments:

[0060] The present invention provides a new semi-supervised visual question answering natural language understanding method, which evaluates candidate explanations through answer rewards to improve the logical consistency between the answer and the reason. The specific process is as follows:

[0061] 1. Extraction of image features

[0062] Given an image in a natural scene, the whole image is adjusted to 224×224×3 and input into the feature extraction network for forward propagation. The "classification token" in ViT-B and CLIP is used to obtain the image features, and then a set of lightweight and simple multi-layer perceptrons are used to convert the image features into where the dimension size c = 768 and the image sequence length S = 10. Finally, visual features of 10×16×16×768 are obtained.

[0063] 2. Extraction of natural language question features

[0064] The natural language question information is decomposed into words, and the corresponding feature vectors of each word are obtained after word embedding. Natural language question where q t represents the t-th word, and T is the length of the question. For question Q, each word q t will be mapped to the corresponding word embedding through a pre-trained image captioning model

[0065] 3. Feature fusion using cross-modal attention

[0066] Fuse the image features and natural language question information from the above two steps to obtain an image and question sequence

[0067]

[0068] 4. Basic answer template

[0069] Use handcrafted prompts as templates to enable the model to generate answers or explanations. First, establish a basic answer template to obtain a basic answer score, and then regard the visual question answering task as a generation task so that the model can generate answers without a predefined answer space. Specifically, for the given image and question sequence Z, use the natural language token "The answer is" to prompt the model to generate the correct answer. By concatenating the language prompts, the basic answer template Z a = [Z; <answer>, where [;] and <answer>respectively represent the splicing operation and specific language prompts.

[0070] 5. Explanation generation template

[0071] To generate reasonable justifications, the language prompt "The reason is" is used to motivate the model to generate free-text explanations. Similar to the basic answer template, the explanation generation template Z e = [Z; <reason>is constructed, wherein <reason>Is a natural language token.

[0072] 6. Candidate Explanation Generation

[0073] Therefore, for unlabeled QA samples, directly apply the explanation generation template to generate candidate rationales. Use beam search to sample the top K words from the VL model probability distribution at each time step and keep these sequences with the highest probability. Then, these generated sentences are integrated into the candidate explanations for each QA pair without manually annotated explanations, where and k represent the k-th rationale and the size of beam search, respectively. For labeled samples, a similar mechanism is used to construct the corresponding set of explanations.

[0074] 7. Explanation Generation Template

[0075] Convert the answer score to a self-criticism reward to evaluate these candidate explanations. Considering that this is a non-differentiable operation, adopt a reinforcement learning method to achieve end-to-end training. Given a sample and the corresponding candidate explanation R, a new input template is designed:

[0076] 8. Model Training

[0077] The entire training process is end-to-end training. Experiments are mainly conducted on two different VQANLE datasets: VQA-X and A-OKVQA. At the same time, since explanation annotation is both expensive and time-consuming, also use the large-scale VQA v2.0 and OK-VQA datasets to construct a semi-supervised learning paradigm. Each image is first preprocessed by CLIP (e.g., including image resizing, center cropping, and normalization). At the same time, fix the ViT-B weights from the CLIP visual encoder to speed up training. For the mapping network, the image sequence length S = 10 and the embedding size is 768. Use the AdamW optimizer with a weight decay of 1e-5, and set the batch size and beam size K to 4 and 2, respectively. The weight coefficient λ is set to 10. Train all models for 30 epochs on 4 1080Ti GPUs with a learning rate of 1e-5.

[0078] 8. Model Application

[0079] After the above training process, multiple models can be obtained. Select the optimal model (the one with the best test effect on the test set) for application. For the input image and question, just adjust the image to a size of 224×224 and normalize it, and perform word segmentation on the sentence, then it can be used as the input of the model. The parameters of the entire network model remain fixed. As long as the input image data and language data are input and propagated forward, the image and language feature vectors V and E, as well as the image question sequence Z, can be obtained in sequence. Then, it is automatically passed into the answer explanation hint module and the self-criticism learning reinforcement module, and the prediction result can be directly obtained. The actual effect diagram is as Figure 2 shown. Based on this method, an explanation about visual question answering can be given efficiently.< / reason> < / reason> < / answer> < / answer> < / reason> < / reason> < / answer> < / answer> < / reason> < / reason> < / answer> < / answer>

Claims

1. A visual reasoning method for self-evolution of an inference chain based on a consistency self-evaluation strategy, characterized in that, It includes the following steps: Step 1: Given an image and a natural language question , where represents the -th word, is the length of the question, represents the size of the image; Obtain the image features of the picture using the classification tokens of the CLIP visual encoder ; Then, use a set of multi-layer perceptrons to convert the image features into , where the dimension size , the length of the image sequence , and the calculation formula is as follows: For the problem , each word will be mapped to the corresponding word embedding through a pre-trained image captioning model ; Finally, an image and a question sequence are obtained : Step 2: For the given image and question sequence , use the natural language label "The answer is" to prompt the model to generate the correct answer by concatenating language prompts, obtaining the basic answer template , where and represent the concatenation operation and the specific language prompt respectively; During the training process, the basic answer template and the true answer label are fed into the VL model, where is the th answer token, indicating the length of the answer label; Calculate the answer loss conditioned on in an autoregressive manner: Among them represent the parameters of the VL model; obtain the average probability as the basic answer score; Using the natural language token "because" to prompt the model to generate free-text explanations, an explanation generation template is constructed, where is a natural language token; For the labeled samples, follow the formula , and calculate the interpretability loss in an autoregressive manner with cross-entropy loss ; Step 3: Use beam search to sample the top words from the VL model probability distribution at each time step and keep these sequences with the highest probability; then, these generated sentences are integrated into the candidate explanations of each QA pair without manually annotated explanations, where and represent the th basic principle and the size of the beam search, respectively; For the labeled samples, the corresponding explanation sets are also constructed ; Finally, since the labeled and unlabeled samples are integrated into a mini-batch during training, combine and into to indicate the candidate explanations for each sample; Step 4: Implement end-to-end training using reinforcement learning, given samples and corresponding candidate explanations , design a new input template: Among them indicates that a possible fundamental explanation is added of the template; then, this template is fed back into the model to obtain the average probability of the answer , that is, the explanatory answer score, where is the true answer of the sample; Meanwhile, the average probability output using the basic answer template is used as the basic score, and the reinforcement loss is calculated by applying reinforcement learning , and the gradient is calculated as follows: Among them is the probability of the nth explanation; Based on the above calculations, when the answer score is higher than the score of the basic answer template , this gradient will tend to increase the probability of the th fundamental explanation; finally, for the labeled samples, the manually annotated fundamental explanations are attached to the candidate explanations , and the answer is predicted through the cross-entropy loss, that is ; During the training process, the overall loss function is expressed as follows: wherein is used to balance the cross-entropy loss and the reinforcement loss; During inference, the model will first generate the rationale for the QA pair and then use the explanation template to obtain the corresponding answer.

Citation Information

Patent Citations

  • Visual question and answer method based on cross-modal pre-training feature enhancement

    CN114663677A

  • Unsupervised hashing method for cross-modal video-text retrieval with clip

    WO2023004206A1