Visual reasoning method based on retrieval enhancement generation and thinking chain technology

By introducing retrieval enhancement generation and thinking chain technology into the visual inference model, the semantic inconsistency problem when external knowledge is fused with image and text reasoning processes is solved, and higher accuracy and interpretability are achieved, and the model's performance in complex scenarios is improved.

CN119990334AActive Publication Date: 2025-05-13XIANGJIANG LAB

Patent Information

Application Number
CN202510462070.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

Existing visual inference models have semantic inconsistency when fusing external knowledge with image and text inference processes, resulting in low quality of inference results and lack of transparency and interpretability, making it difficult to gain trust in application scenarios that require high interpretability.

Method used

The visual reasoning method based on search-enhanced generation and thinking chain technology is adopted, and the input image is preprocessed and segmented, combined with thinking chain technology is used to perform step-by-step reasoning, and the search-enhanced generation technology is used to retrieve knowledge fragments from the external knowledge base to ensure the semantic consistency between the image and the generated inference text.

Benefits of technology

By enriching the semantic information of the image and ensuring semantic consistency, the accuracy and stability of the visual inference model are significantly improved, while enhancing the transparency and interpretability of the inference process, improving the performance of the model in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990334A_ABST
    Figure CN119990334A_ABST
Patent Text Reader

Abstract

The invention relates to a visual reasoning method based on retrieval enhancement generation and thinking chain technology, and the method comprises the steps: carrying out the preprocessing of an input original image, and dividing the preprocessed original image into a plurality of regions of interest; performing word segmentation and word embedding on the input question to obtain question feature representation; gradually reasoning each region of interest by using a thinking chain technology, and combining the obtained reasoning texts in sequence to generate a multi-step reasoning text; performing retrieval in an external knowledge base based on the problem and the multi-step reasoning text by adopting a retrieval enhancement generation technology to obtain knowledge fragments; inputting the multi-step reasoning text and the knowledge fragment into the optimized generative model to obtain a preliminary reasoning result; and a BERT pre-training model is adopted to check the logic consistency of the preliminary reasoning result, and after the check is reasonable, the preliminary reasoning result is simplified through the BERT pre-training model to obtain a final reasoning result. The method can effectively improve the accuracy and stability of the visual reasoning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of visual reasoning technology, and in particular to a visual reasoning method based on retrieval enhancement generation and thought chain technology. Background Art

[0002] Although many studies in recent years have tried to introduce external knowledge bases (such as encyclopedias) into visual reasoning to enhance the reasoning ability of the model, the existing knowledge introduction methods still have some limitations. Current methods often call external knowledge as an independent module without effectively integrating it with the reasoning process of images and texts. This approach not only increases the complexity of the model, but may also lead to semantic inconsistencies between external knowledge and image content or reasoning text, thereby affecting the quality of reasoning results. In addition, many models only rely on static, predefined knowledge bases, which has great limitations for dealing with unknown fields or tasks that require real-time reasoning.

[0003] In addition to the challenge of efficiently combining external knowledge, existing models are often viewed as a "black box" during reasoning, lacking a transparent reasoning process and interpretability. This black box nature makes it difficult for users to understand the reasoning logic of the model, and in application scenarios that require high interpretability, this black box nature becomes a significant obstacle. The model's reasoning process lacks visualization and explanation methods, which limits its trust and acceptance in practical applications. Summary of the invention

[0004] Based on this, it is necessary to provide a visual reasoning method based on retrieval enhancement generation and thought chain technology, which includes: S1: preprocessing the input original image, and dividing the preprocessed original image into multiple regions of interest; Perform word segmentation and word embedding on the input question to obtain the feature representation of the question; S2: using the thinking chain technology to perform step-by-step reasoning on each of the regions of interest, and combining the obtained reasoning texts in sequence to generate a multi-step reasoning text; S3: extracting features from the preprocessed original image and the multi-step inference text respectively to obtain an image feature vector and a text feature vector; constructing a loss function based on the similarity between the image feature vector and the text feature vector, and optimizing the generation model based on the loss function; S4: using the retrieval enhancement generation technology to search in an external knowledge base based on the question and the multi-step reasoning text to obtain knowledge fragments; inputting the multi-step reasoning text and the knowledge fragments into the optimized generation model together to obtain preliminary reasoning results; S5: Use the BERT pre-training model to verify the logical consistency of the preliminary reasoning result. After the verification is reasonable, simplify the preliminary reasoning result through the BERT pre-training model to obtain the final reasoning result.

[0005] Beneficial effects: This method enriches the semantic information of the image through thought chain technology and retrieval enhancement generation technology, and ensures the semantic consistency between the image and the generated reasoning text, thereby effectively improving the accuracy and stability of the visual reasoning model. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0007] Figure 1 This is a flowchart of the visual reasoning method based on retrieval enhancement generation and thought chain technology in an embodiment of the present application. DETAILED DESCRIPTION

[0008] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present application, so the present application is not limited by the specific embodiments disclosed below.

[0009] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0010] like Figure 1 As shown, this embodiment provides a visual reasoning method based on retrieval enhancement generation and thought chain technology, the method comprising: S1: preprocessing the input original image, and dividing the preprocessed original image into multiple regions of interest.

[0011] Specifically, in order to improve the stability and convergence speed of model training, the input original image is preprocessed, and the preprocessing includes normalization processing; the formula for normalization processing is: ; ; ; in, represents the normalized image; represents the original image; Represents the mean pixel value of the original image; represents the standard deviation of the pixel value of the original image; M represents the total number of pixels in the original image; Indicates the original image i The value of pixels; Mask-R-CNN is used to divide the normalized image into multiple regions of interest, each of which contains a semantic part of the image, which helps the model better focus on important information in the image.

[0012] Perform word segmentation and word embedding on the input question to obtain the question feature representation.

[0013] The obtaining of the problem feature representation comprises: Segment the question into word sequences, and use the BERT pre-trained model to convert the word sequences into embedding vectors; Add position encoding to the embedding vector to get the initial representation of the problem, calculated as: ; in, represents the initial representation of question Q; Embedding vector representing question Q; Represents the positional encoding of question Q; The initial representation is input into the BERT pre-trained model to obtain the problem feature representation.

[0014] Example: Assume that the input original image content is: a man holding a ring in his hand, kneeling on one knee in a heart-shaped wreath facing a smiling woman. After image segmentation, the following regions of interest can be obtained: R1: man; R2: woman; R3: ring; R4: heart-shaped wreath; R5: man's kneeling posture; R6: woman's expression.

[0015] Suppose the input question is: "What is this man doing?", and its word sequence is ["this", "man", "is", "doing", "what", "behavior"].

[0016] S2: Use the thinking chain technology to perform step-by-step reasoning on each of the regions of interest, and combine the obtained reasoning texts in sequence to generate a multi-step reasoning text.

[0017] Specifically, generating a multi-step reasoning text includes: A convolutional neural network is used to extract features from each region of interest to obtain feature representations corresponding to each region of interest. The calculation formula is: ; in, Indicates the area of ​​interest The corresponding feature representation is, , d Represents feature dimension; represents a convolutional neural network; The feature representation corresponding to each region of interest is input into the generation model to generate the inference text corresponding to each region of interest. The calculation formula is: ; in, Indicates the area of ​​interest The corresponding reasoning text; represents a generative model; The reasoning texts corresponding to all the regions of interest are combined in a logical order to generate the multi-step reasoning text, and the calculation formula is: ; in, Represents a multi-step reasoning text, which contains a detailed description of each important area of ​​the image and the reasoning process, enriching the semantic information of the image; Indicates a text link operation; Represents the inference text corresponding to the nth region of interest.

[0018] In this embodiment, the generation model includes LSTM or Transformer.

[0019] Continuing with the above example, feature representation is extracted for each region of interest and inference text is generated: s1: a man is kneeling on one knee; s2: a woman is standing in front of the man; s3: the man is holding a ring in his hand; s4: they are in a heart-shaped wreath, which usually symbolizes romantic love; s5: the woman is smiling, which may indicate joy or acceptance. The enriched semantic information is combined to obtain multi-step inference text: a man is kneeling on one knee, the man is holding a ring in his hand, a woman is standing in front of the man, the woman is smiling, which usually indicates acceptance or happiness, and they are in a heart-shaped wreath, which usually symbolizes romantic love.

[0020] S3: extract features from the preprocessed original image and the multi-step inference text respectively to obtain an image feature vector and a text feature vector; construct a loss function based on the similarity between the image feature vector and the text feature vector, and optimize the generation model based on the loss function.

[0021] Specifically, obtaining the image feature vector and the text feature vector includes: The preprocessed original image is divided into m image blocks of fixed size, each image block is expanded into a vector, and linear embedding is performed to obtain the feature representation of each image block. The calculation formula is: ; in, Represents the feature representation of the i-th image block; represents a linear transformation; Represents a flattening operation; represents the i-th image block; The second position code is added to the feature representation of each image block to obtain the initial features of each image block. The calculation formula is: ; in, Indicates i The initial features of the image blocks; Indicates i a second position code of an image block; The initial features of all image blocks are input into Transformer to obtain the image feature representation, which is calculated as follows: ; in, represents image feature representation, , m represents the number of image blocks; Represents the Transformer model; The multi-step reasoning text is segmented to obtain Words , convert each word into a word vector , and add the third position code to each word vector to obtain the initial text features of each word vector. The calculation formula is: ; in, Indicates j The initial text features of word vectors; Indicates j word vectors; Indicates j The third position encoding of the word vector; The initial text features are input into Transformer to obtain text feature representation, and the calculation formula is: ; in, represents text feature representation, , Indicates the number of words.

[0022] In order to deeply integrate the image feature representation with the text feature representation, this embodiment introduces a cross attention mechanism. First, based on the image feature representation and the text feature representation, the first attention weight matrix from image to text is calculated, and the calculation formula is: ; in, Represents the first attention weight matrix from image to text; represents the softmax activation function; Represents the query matrix, which is used to convert the input feature representation F (text feature representation or image feature representation) into a query vector, calculate the similarity with the key, and determine the feature parts that need to be focused on; The linear transformation matrix representing the key converts the input feature representation F into a key vector, and calculates the dot product with the query vector to measure the correlation between different modalities; represents the scaling factor to prevent the inner product value from being too large; T represents transpose; The text feature representation is weighted based on the first attention weight matrix to obtain an image feature vector, which is calculated as follows: ; in, represents the image feature vector; The linear transformation matrix representing the value converts the input feature representation F into a value vector for the final weighted summation to generate a new feature representation; Based on the image feature representation and the text feature representation, a second attention weight matrix from text to image is calculated, and the calculation formula is: ; in, Represents the second attention weight matrix from text to image; represents the softmax activation function; Represents the query matrix, which is used to convert the input feature representation F (text feature representation or image feature representation) into a query vector, calculate the similarity with the key, and determine the feature parts that need to be focused on; The linear transformation matrix representing the key converts the input feature representation F into a key vector, and calculates the dot product with the query vector to measure the correlation between different modalities; represents the scaling factor to prevent the inner product value from being too large; T represents transpose; The image feature representation is weighted based on the second attention weight matrix to obtain a text feature vector, which is calculated as follows: ; in, Represents a text feature vector; The linear transformation matrix representing the value converts the input feature representation F into a value vector for the final weighted summation to generate a new feature representation.

[0023] The text feature representation generated by thought chain reasoning represents the intermediate reasoning steps generated by the model during the reasoning process. These features contain rich semantic information and logical relationships. However, relying solely on text features for reasoning may not fully utilize the visual information in the image. By introducing the cross-attention mechanism, the text feature representation can be effectively combined with the image feature representation to enhance the semantic expression and contextual association of the text feature representation.

[0024] The cross-attention mechanism enables the fusion process of text feature representation and image feature representation to be summarized, and can dynamically focus on the area in the image that is related to the current text reasoning step. This attention mechanism ensures that each step of text reasoning can use the visual clues in the image to improve the accuracy of the reasoning task. For example, when identifying the "marriage proposal scene", the "ring" and "kneeling on one knee" in the text feature representation can be associated with specific visual areas in the image (such as the position of the ring and the kneeling posture) through the cross-attention mechanism, making the reasoning more consistent with the actual visual content.

[0025] Furthermore, the loss function is constructed based on the similarity between the image feature vector and the text feature vector, and the generation model is optimized based on the loss function, including: Calculate the first cosine similarity between the image feature vector and the text feature vector; Calculating the second cosine similarity between the image feature vector and the text feature vector of the negative sample; Constructing a contrast loss function based on the first cosine similarity and the second cosine similarity; Adding a regularization term to the contrast loss function to obtain the loss function; The loss function is minimized, the model parameters of the generative model are updated, and an optimized generative model is obtained.

[0026] In this embodiment, the loss function is expressed as: ; ; ; ; in, represents the loss function; represents the contrast loss function; represents the regularization term; represents cosine similarity; represents the image feature vector; Represents a text feature vector; represents the temperature parameter; N represents the number of negative samples; Indicates i The text feature vector of negative samples; represents the regularization coefficient; The model parameter update formula is: ; in, represents the updated model parameters; Represents the model parameters before updating; represents the learning rate; Represents the loss function The gradient of the model parameters.

[0027] S4: Using the retrieval enhancement generation technology, a search is performed in an external knowledge base based on the question and the multi-step reasoning text to obtain knowledge fragments; the multi-step reasoning text and the knowledge fragments are input into the optimized generation model together to obtain preliminary reasoning results.

[0028] Specifically, the question and the multi-step reasoning text are input into the retrieval model, and the text related to the question is retrieved from the external knowledge base based on the retrieval enhancement generation technology to obtain the knowledge fragment, which is expressed as: ; in, Represents a piece of knowledge; represents the retrieval model; Q represents the question; represents multi-step reasoning text; D represents an external knowledge base (such as Wikipedia, Baidu Encyclopedia, etc.); The retrieval enhancement generation technology includes a retrieval method based on the FAISS or BM25 model.

[0029] The multi-step reasoning text and the knowledge fragment are input into the optimized generation model to obtain a preliminary reasoning result, which is expressed as: ; in, Indicates preliminary reasoning results; Represents the optimized generative model.

[0030] Continuing with the above example, based on the words "man", "kneeling on one knee", "holding a ring in hand" and the question "What is this man doing?" in the multi-step reasoning text, we can retrieve fragments from the external knowledge base: "holding a ring in hand is usually for proposing", "kneeling on one knee is a common gesture for proposing", "there are usually heart-shaped wreaths in proposal scenes", and combine the retrieved knowledge fragments with the multi-step reasoning text to generate a more detailed and accurate preliminary reasoning result: a man is kneeling on one knee, holding a ring in his hand, a woman is standing in front of the man, the woman is smiling, a smile usually means acceptance or happiness, they are in a heart-shaped wreath, and the heart-shaped wreath usually symbolizes romantic love. This posture and background indicate that the man is proposing to the woman.

[0031] S5: Use the BERT pre-training model to verify the logical consistency of the preliminary reasoning result. After the verification is reasonable, simplify the preliminary reasoning result through the BERT pre-training model to obtain the final reasoning result.

[0032] Specifically, the step includes: Input any sentence pair in the preliminary reasoning result into the BERT pre-trained model, and output the causal relationship determination result of the two sentences in the sentence pair, expressed as: ; in, Expressing sentences With sentence The result of the causal relationship determination between the two sentences is a probability value ranging from 0 to 1; when the probability value is higher than the preset probability, it is determined that there is a reasonable causal relationship between the two sentences; when the probability value is greater than 0 and less than or equal to the preset probability, it is determined that there is an unreasonable causal relationship between the two sentences; when the probability value is equal to 0, it is determined that there is no causal relationship between the two sentences; represents the sigmoid activation function; W represents the weight of the BERT pre-training model; T represents transposition; Represents the sentence vector after the transformation of the i-th sentence in the preliminary reasoning result; Represents the sentence vector after the j-th sentence in the preliminary reasoning result is transformed; Represents the concatenation of sentence vectors; b represents the bias term; If it is determined that there is a reasonable causal relationship between the two sentences, the two sentences are marked with a reasonable causal relationship, and the causal relationship determination results of the two sentences in the remaining sentence pairs are continued to be determined until all the sentence pairs in the preliminary reasoning results are determined; If it is determined that there is an unreasonable causal relationship between the two sentences, optimizing the generation model; If it is determined that there is no causal relationship between the two sentences, the two sentences are marked as having no causal relationship, and the causal relationship determination results of the two sentences in the remaining sentence pairs are continued to be determined until all the sentence pairs in the preliminary reasoning results are determined; After the preliminary reasoning result is verified, the preliminary reasoning result is simplified through the BERT pre-training model to obtain the final reasoning result.

[0033] Continuing with the above example, the preliminary inference result is simplified to: the man is proposing to the woman through the BERT pre-trained model.

[0034] The visual reasoning method based on retrieval enhancement generation and thought chain technology provided in this embodiment has the following beneficial effects: 1. Enhance the multi-step reasoning ability of the model: By introducing the thinking chain technology, the visual reasoning model can reason step by step, avoiding the errors or omissions that may be caused by the reasoning process in traditional methods. Each step of reasoning has clear logic and intermediate steps, which improves the accuracy of reasoning and the interpretability of the reasoning process. When dealing with complex, multi-step reasoning tasks, it can effectively reduce the occurrence of erroneous reasoning.

[0035] 2. Enhance the common sense reasoning ability of the model: The introduction of retrieval enhancement generation technology enables the model to obtain relevant common sense or background information from the external knowledge base, making up for the lack of contextual support in images and texts, thereby generating more accurate reasoning results. This improves the performance of the model in complex scenes containing symbolic information.

[0036] 3. Improve reasoning performance in complex scenarios: Combining multi-step reasoning with common sense reasoning, this method can handle more complex reasoning tasks. By combining thought chain technology and retrieval enhancement generation technology, this method can not only perform multi-level and multi-step reasoning, but also retrieve rich background information from external knowledge bases, greatly improving the diversity and robustness of reasoning tasks.

[0037] 4. Enhanced cross-modal understanding and reasoning capabilities: By combining contrastive learning with the cross-attention mechanism, this method further enhances the cross-modal understanding and reasoning capabilities between images and text. After deep fusion of image feature representation and text feature representation, the reasoning model can more accurately capture the relationship between images and text, improving the model's reasoning capabilities in complex scenarios.

[0038] 5. Improve the practicality of the model in practical applications. The model can adapt to a variety of visual reasoning tasks, such as visual question answering, image description generation, and visual common sense reasoning, and has broad application prospects.

[0039] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0040] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be construed as limiting the scope of the patent application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent application shall be subject to the attached claims.

Claims

1. A visual reasoning method based on retrieval enhancement generation and thought chain technology, characterized in that: include: S1: preprocessing the input original image, and dividing the preprocessed original image into multiple regions of interest; Perform word segmentation and word embedding on the input question to obtain the feature representation of the question; S2: using the thinking chain technology to perform step-by-step reasoning on each of the regions of interest, and combining the obtained reasoning texts in sequence to generate a multi-step reasoning text; S3: extracting features from the preprocessed original image and the multi-step inference text respectively to obtain an image feature vector and a text feature vector; constructing a loss function based on the similarity between the image feature vector and the text feature vector, and optimizing the generation model based on the loss function; S4: using the retrieval enhancement generation technology to search in an external knowledge base based on the question and the multi-step reasoning text to obtain knowledge fragments; inputting the multi-step reasoning text and the knowledge fragments into the optimized generation model together to obtain preliminary reasoning results; S5: Use the BERT pre-training model to verify the logical consistency of the preliminary reasoning result. After the verification is reasonable, simplify the preliminary reasoning result through the BERT pre-training model to obtain the final reasoning result.

2. The visual reasoning method based on retrieval enhancement generation and thought chain technology according to claim 1 is characterized in that: The preprocessing includes normalization processing; Mask-R-CNN is used to divide the normalized image into multiple regions of interest, each of which contains a semantically significant part of the image.

3. The visual reasoning method based on retrieval enhancement generation and thought chain technology according to claim 1 is characterized in that: The obtaining of the problem feature representation comprises: Segment the question into word sequences, and use the BERT pre-trained model to convert the word sequences into embedding vectors; Adding position encoding to the embedding vector to obtain an initial representation of the problem; The initial representation is input into the BERT pre-trained model to obtain the problem feature representation.

4. The visual reasoning method based on retrieval enhancement generation and thought chain technology according to claim 1 is characterized in that: Generating a multi-step reasoning text comprises: Using a convolutional neural network to extract features from each region of interest, and obtaining feature representations corresponding to each region of interest; Input the feature representation corresponding to each region of interest into the generation model to generate the reasoning text corresponding to each region of interest; The reasoning texts corresponding to all the regions of interest are combined in a logical order to generate the multi-step reasoning text.

5. The visual reasoning method based on retrieval enhancement generation and thought chain technology according to claim 1 is characterized in that: The obtaining of the image feature vector and the text feature vector comprises: The preprocessed original image is divided into m image blocks of fixed size, each image block is expanded into a vector, and linear embedding is performed to obtain the feature representation of each image block; a second position code is added to the feature representation of each image block to obtain the initial feature of each image block; Inputting the initial features of all image blocks into the Transformer to obtain the image feature representation; Segmenting the multi-step reasoning text to obtain multiple words, converting each word into a word vector, and adding a third position code to each word vector to obtain initial text features of each word vector; Inputting the initial text features into Transformer to obtain text feature representation; Based on the image feature representation and the text feature representation, calculating a first attention weight matrix from image to text; Weighting the text feature representation based on the first attention weight matrix to obtain an image feature vector; Based on the image feature representation and the text feature representation, calculating a second attention weight matrix from text to image; The image feature representation is weighted based on the second attention weight matrix to obtain a text feature vector.

6. The visual reasoning method based on retrieval enhancement generation and thought chain technology according to claim 1 is characterized in that: The loss function is constructed based on the similarity between the image feature vector and the text feature vector, and the generation model is optimized based on the loss function, including: Calculate the first cosine similarity between the image feature vector and the text feature vector; Calculating the second cosine similarity between the image feature vector and the text feature vector of the negative sample; Constructing a contrast loss function based on the first cosine similarity and the second cosine similarity; Adding a regularization term to the contrast loss function to obtain the loss function; The loss function is minimized, the model parameters of the generative model are updated, and an optimized generative model is obtained.

7. The visual reasoning method based on retrieval enhancement generation and thought chain technology according to claim 6 is characterized in that: The expression of the loss function is: ; ; ; ; in, represents the loss function; represents the contrast loss function; represents the regularization term; represents cosine similarity; represents the image feature vector; Represents a text feature vector; represents the temperature parameter; N represents the number of negative samples; Indicates i The text feature vector of negative samples; represents the regularization coefficient.

8. The visual reasoning method based on retrieval enhancement generation and thought chain technology according to claim 1 is characterized in that: In S4, Inputting the question and the multi-step reasoning text into a retrieval model, retrieving text related to the question in an external knowledge base based on retrieval enhancement generation technology, and obtaining knowledge fragments; The retrieval enhancement generation technology includes a retrieval method based on the FAISS or BM25 model.

9. The visual reasoning method based on retrieval enhancement generation and thought chain technology according to claim 1 is characterized in that: In S5, Input any sentence pair in the preliminary reasoning result into the BERT pre-trained model, and output the causal relationship determination result of the two sentences in the sentence pair; If it is determined that there is a reasonable causal relationship between the two sentences, the two sentences are marked with a reasonable causal relationship, and the causal relationship determination results of the two sentences in the remaining sentence pairs are continued to be determined until all the sentence pairs in the preliminary reasoning results are determined; If it is determined that there is an unreasonable causal relationship between the two sentences, optimizing the generation model; If it is determined that there is no causal relationship between the two sentences, the two sentences are marked as having no causal relationship, and the causal relationship determination results of the two sentences in the remaining sentence pairs are continued to be determined until all the sentence pairs in the preliminary reasoning results are determined; After the preliminary reasoning result is verified, the preliminary reasoning result is simplified through the BERT pre-training model to obtain the final reasoning result.

10. The visual reasoning method based on retrieval enhancement generation and thought chain technology according to claim 1 is characterized in that: The generation model includes LSTM or Transformer.

Citation Information

Patent Citations

  • Cross-modal retrieval model construction and retrieval method based on image-text collaborative attention

    CN114201621A

  • Cross-modal retrieval method for image-text

    CN114461836A

  • Multi-task learning model combining image-text matching and visual reasoning, visual common sense reasoning method and computer equipment

    CN114996502A

  • Visual question and answer method and device and storage medium

    CN115618045A

  • Visual question-answering system and method based on scene graph neural network inference mechanism

    CN117010501A

Cited By

  • Image annotation control method

    CN121330687A