Visual knowledge question-answering method based on question-enhanced knowledge retrieval network

By using a question-based knowledge retrieval network in visual question-answer tasks, identifying the image areas related to the problem and building problem-enhanced queries, the problems of inaccurate image-to-text conversion and insufficient correlation in the prior art are solved, and more accurate and high-quality answer generation is achieved.

CN119938864AActive Publication Date: 2025-05-06GUANGDONG UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510114188.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-06
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

In the prior art In visual question-and-answer task based on external knowledge, the query of image-to-text conversion is not accurate and redundant enough, and the correlation calculation between the query and knowledge is not sufficient to answer the question.

Method used

A method based on problem-enhanced knowledge retrieval network is proposed, which identifies image areas related to the problem through the cross attention mechanism, generates picture titles related to the problem and retains picture objectives related to the problem, constructs a problem-enhanced query, and reorders the retrieved knowledge through the reverse inference reordering search module to enhance the fine-grained interaction between the problem and knowledge.

Benefits of technology

It improves the accuracy and relevance of image-to-text conversion, reduces the loss of key information during query construction, enhances the interaction between questions and knowledge, and improves the answer generation quality of visual question-and-answer answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938864A_ABST
    Figure CN119938864A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge visual question-answering method based on a question-enhanced knowledge retrieval network. The method comprises the following steps: acquiring a to-be-detected image and a corresponding question; inputting a to-be-detected image and a corresponding question into the question-based enhanced knowledge retrieval network for processing, and outputting an answer; wherein the question-based enhanced knowledge retrieval network is used for identifying an image region related to a question by using a cross attention mechanism, generating a picture title related to the question and retaining a picture target related to the question. According to the knowledge retrieval network based on question enhancement, the importance of questions in query construction is enhanced, and rich cross attention between articles and queries is provided in article retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal visual question answering, and in particular to a knowledge visual question answering method based on a question-enhanced knowledge retrieval network. Background Art

[0002] Visual question answering based on external knowledge is a challenging visual question answering task that requires retrieving external knowledge to answer questions about images. Recently, scholars engaged in research in this field have proposed two main paradigms: the multimodal space paradigm and the text space paradigm. The multimodal space paradigm usually uses knowledge embedding and fuses the output of the visual language model with the knowledge graph through a graph convolutional neural network to obtain the final answer, while the text space paradigm uses multimodal models such as image caption generation and target retrieval to convert the visual information in the image into text information, thereby converting the visual question answering task into an approach similar to the open domain question answering task. Specifically, the approach of the text space paradigm is to add the text information obtained from the image conversion to the question to form a query, which is then given to the knowledge retriever to find relevant knowledge, and finally the above text is given to the language model to generate the final answer. The limitation of the multimodal space paradigm is that it is necessary to compress a large amount of rich text knowledge into a much smaller multimodal space, resulting in insufficient data available for training multimodal models, while the text space paradigm converts the visual information of the image into text information and can use the large amount of text knowledge on the Internet for question answering, which makes the latter a mainstream method recently.

[0003] There are two main challenges in the text space paradigm. On the one hand, the queries obtained by converting images to text may be inaccurate and redundant. On the other hand, the relevance between queries and knowledge is calculated by their semantic similarity, but this approach is not fine enough and knowledge with high similarity may not necessarily help answer the corresponding questions. TRiG proposed three levels of image-to-text conversion, including image caption information, image recognition object information, and image character information to construct queries. However, the visual information obtained from their image caption generation tool and object retrieval tool is sometimes useless for answering questions because their method lacks attention to the questions, and some questions only ask about the content related to specific areas of the image. The most widely used DPR (Dense Passage Retrieval) relies on the semantic similarity between knowledge and questions to retrieve knowledge by calculating one-dimensional embedding. However, high-scoring passages may not necessarily help answer questions and are insensitive to fine-grained relevance. Summary of the invention

[0004] In order to solve the technical problems existing in the above-mentioned prior art, the present invention proposes a knowledge visual question answering method based on a question-enhanced knowledge retrieval network.

[0005] To achieve the above object, the present invention provides a knowledge visual question answering method based on a question-enhanced knowledge retrieval network, comprising:

[0006] Get the image to be detected and the corresponding question;

[0007] Input the image to be detected and the corresponding question into the question-enhanced knowledge retrieval network for processing, and output the answer;

[0008] The question-enhanced knowledge retrieval network is used to identify image regions related to the question using a cross-attention mechanism, generate picture captions related to the question, and retain picture targets related to the question.

[0009] Preferably, the question-based enhanced knowledge retrieval network includes:

[0010] Question enhancement query building module: used to obtain the final text query using several different levels of transformations on the image to be detected and the corresponding question;

[0011] Backward Reasoning Re-ranking Retrieval Module: It is used to retrieve top-k relevant knowledge from the text corpus using the DPR retriever given a text query and question;

[0012] Answer output module: used to connect the text output of the question enhancement query construction module and the knowledge retrieved by the reverse reasoning re-ranking retrieval module, and provide them to the T5 model for answer generation.

[0013] Preferably, the question enhancement query building module includes:

[0014] Question-enhanced title generation unit: used to obtain image title information related to the question;

[0015] Question-enhanced label filtering unit: used to filter image label information related to questions based on images and questions;

[0016] Text recognition unit: used to extract character information from images.

[0017] Preferably, the processing process of the question-enhanced title generation unit is:

[0018]

[0019] In the formula, is the oth token of the generated image title information, I i For the image, Q i For the problem, C i is the image title information related to the question, f CG Enhanced title generation unit for questions;

[0020] The processing process of the question-enhanced label screening unit is as follows:

[0021]

[0022] Where, L i is the image label information related to the question, To identify the target’s attribute information, To identify the label information of the target, f LF Enhanced tag filtering unit for questions;

[0023] The processing process of the character recognition unit is as follows:

[0024]

[0025] In the formula, is the oth token of the text information obtained in the image, O i is the character information extracted from the image, f OCR It is the text recognition unit.

[0026] Preferably, obtaining the final text query includes:

[0027] The information output by the question-enhanced title generation unit, the question-enhanced tag screening unit and the text recognition unit are connected to obtain the final overall text query after conversion, which is specifically:

[0028] T i ={C i ,L i ,O i ,Q i};

[0029] Where, T i For the final text query.

[0030] Preferably, using the DPR retriever to retrieve top-k relevant knowledge from a text corpus comprises:

[0031]

[0032] In the formula, r is the correlation score between query and knowledge, Q i For the problem, P i is the knowledge retrieved from DPR, is the question encoder, A knowledge encoder.

[0033] Preferably, after retrieving the top-k related knowledge, the following steps are also included:

[0034] The knowledge retrieved by DPR is reordered using the reverse reasoning reordering mechanism, and the average log-likelihood of the question tokens under the given knowledge is calculated using the T5 model, and logp(T i |P i ) to estimate:

[0035]

[0036] In the formula, β is the parameter of the language model, |Q i | is the number of question tokens, and q is any token in the question.

[0037] Preferably, the processing process of the answer output module is:

[0038] answer = f T5 (T i ,P i );

[0039] Among them, T i For the final text query, P i is the knowledge retrieved from DPR, f T5 is the T5 language model, and answer is the output answer.

[0040] Compared with the prior art, the present invention has the following advantages and technical effects:

[0041] The present invention proposes a question-enhanced knowledge retrieval network, which enhances the importance of questions in query construction and provides rich cross-attention between articles and queries in article retrieval. The question-enhanced query construction module (QQC) and reverse reasoning re-ranking retrieval module (RIR) proposed in the present invention do not require additional annotated data for training and can be extended to the text space paradigm model. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0043] Figure 1 A schematic diagram of the structure of a question-based enhanced knowledge retrieval network according to an embodiment of the present invention;

[0044] Figure 2 A schematic diagram of a label screening unit enhanced for problems in an embodiment of the present invention;

[0045] Figure 3 The figure is a schematic diagram for visualizing an embodiment of the present invention. DETAILED DESCRIPTION

[0046] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0047] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0048] Knowledge-based visual question answering (KB-VQA) requires external knowledge in addition to image content to answer questions. Currently, many works convert all content into text space for knowledge retrieval through text space paradigm retrievers, but there are two main limitations of text space paradigm retrievers in KB-VQA: (1) the query obtained by image-to-text conversion may be inaccurate and redundant due to the lack of questions; (2) the relevance between the query and the supporting knowledge is calculated by semantic similarity, which may not be sufficient to answer the question. To this end, this embodiment proposes a knowledge visual question answering method based on question-enhanced knowledge retrieval network, wherein the question-enhanced knowledge retrieval network (QKRN) consists of a question-enhanced query construction module (QQC) and a reverse reasoning re-ranking retrieval module (RIR). More specifically, the QQC module uses a cross-attention mechanism to locate the visual area related to the question and construct the question-enhanced query. In addition, the RIR module re-ranks the knowledge retrieved from the DPR by calculating the possibility of question generation conditioned on the knowledge. This example demonstrates the excellent performance of the proposed question-enhanced knowledge retrieval network QKRN through a large number of experiments conducted on the OK-VQA and FVQA datasets.

[0049] A knowledge-based visual question answering method based on question-augmented knowledge retrieval network, such as Figure 1 ,include:

[0050] Get the image to be detected and the corresponding question;

[0051] The image to be detected and the corresponding question are input into the question-enhanced knowledge retrieval network for processing, and the answer is output;

[0052] Among them, the question-enhanced knowledge retrieval network is used to identify image regions related to the question using a cross-attention mechanism, generate picture captions related to the question, and retain picture targets related to the question.

[0053] Question-based enhanced knowledge retrieval network includes:

[0054] Question enhancement query building module: used to obtain the final text query using several different levels of transformations on the image to be detected and the corresponding question;

[0055] Backward Reasoning Re-ranking Retrieval Module: It is used to retrieve top-k relevant knowledge from the text corpus using the DPR retriever given a text query and question;

[0056] Answer output module: used to connect the text output of the question enhancement query construction module and the knowledge retrieved by the reverse reasoning re-ranking retrieval module, and provide them to the T5 model for answer generation.

[0057] Specifically, in the question enhancement query construction module, a cross-attention mechanism is used to identify image regions related to the question. Based on these image regions, picture captions related to the question are generated and picture targets related to the question are retained. The method of this embodiment combines comprehensive consideration of images and questions, so it can target image information related to the question, thereby reducing the loss of key information in the query construction process and improving its accuracy. In the reverse reasoning re-ranking retrieval module, a large language model is first used to reversely generate questions based on the knowledge retrieved by DPR. In addition, the generated questions are re-ranked according to their log-likelihood scores to ensure that the retrieved paragraphs have a positive impact on answering the questions and enhance the fine-grained interaction between questions and knowledge during the retrieval process.

[0058] Compared with the previous method (shaded part), in the question enhancement query construction module, the image areas related to the question are identified by using a cross-attention mechanism. Based on these image areas, picture captions related to the question are generated and picture targets related to the question are retained. The method of this embodiment combines comprehensive consideration of images and questions, so it can target image information related to the question, thereby reducing the loss of key information in the query construction process and improving its accuracy. In the reverse reasoning re-ranking retrieval module, a large language model is first used to reversely generate questions based on the knowledge retrieved by DPR. In addition, the generated questions are re-ranked according to their log-likelihood scores to ensure that the retrieved paragraphs have a positive impact on answering the questions and enhance the fine-grained interaction between questions and knowledge during the retrieval process.

[0059] Furthermore, the question enhancement query building module includes:

[0060] Question-enhanced title generation unit: used to obtain image title information related to the question;

[0061] Question-enhanced label filtering unit: used to filter image label information related to questions based on images and questions;

[0062] Character recognition (OCR) unit: used to extract character information from images.

[0063] Furthermore, the processing process of the question-enhanced title generation unit is as follows:

[0064]

[0065] In the formula, is the oth token of the generated image title information, I i For the image, Q i For the problem, C i is the image title information related to the question, f CG Enhanced title generation unit for questions;

[0066] Specifically, given an image I i and question Q i , the question-enhanced caption generation unit (CG) uses the ITE module in BILP and the GradCAM technique from the image I i Sampling and Question Q i The relevant key image regions are used to generate semantically meaningful captions C i .

[0067] In this embodiment, Represents the features extracted from each image region, and represents the features extracted for each question, i is the number of layers of ITE, K is the number of image regions, L is the number of tokens in the question, is the dimension of the image feature of layer i in ITE, is the dimension of the question feature of layer i in ITE. Taking the process of layer i in ITE as an example, the layer i of ITE calculates the cross attention score between each image region and each question token as follows:

[0068]

[0069] In the formula, and are the parameter matrices in each cross-attention head of the i-th layer ITE, W i ∈R L×K is the cross-attention matrix, each row is the cross-attention score between the entire image area and each question token, and the larger the value, the higher the correlation.

[0070] Attention matrix W i It can be seen as ITE calculating the similarity between the image and the question. At the selected layer of the ITE block, GradCAM is used to calculate the derivative of the similarity score and the cross-attention score, and the gradient matrix elements are multiplied by the cross-attention score. The relevance of the k-th image region to the question It can be calculated by taking the average of H attention heads and summing the L text tokens:

[0071]

[0072] Where h is the attention head index and i is the number of layers of ITE.

[0073] Given a region relevance score, an image region is sampled with a probability based on the relevance score, and then a caption is generated from the sampled image region. In order to generate semantically meaningful captions, a short prompt is also fed into the text decoder and repeated N times for each image to generate N different captions. This process can be described by the following formula:

[0074] C i =f caption (s o ,...,s k );

[0075] In the formula, {s o ,...,s k} are the top-k image regions sampled, and their sampling probability is proportional to the correlation score.

[0076] Furthermore, the processing process of the question-enhanced label filtering unit is as follows:

[0077]

[0078] Where, L i is the image label information related to the question, To identify the target’s attribute information, To identify the label information of the target, f LF Enhanced tag filtering unit for questions;

[0079] Specifically, given an image I i and question Q i The enhanced label filtering unit (LF) calculates the overlap between the image area obtained by CG and the label obtained by the target detection model, and filters out irrelevant information to obtain the label information L i .

[0080] Given an image, use the target detection model to obtain the location information and label information of the target area recognized by the image:

[0081] {B o ,...,B n} = f detection (I i );

[0082] In the formula, {B o ,...,B n} is the information obtained by the target detection model.

[0083] Existing object detection models are powerful enough to identify a large number of objects, but too many irrelevant labels will also interfere with the final language model to generate the correct answer. This embodiment proposes a new label screening strategy to screen image labels that are irrelevant to the question based on their relevance to the question.

[0084] like Figure 2 As shown, the regional position information obtained by the target detection model is first converted into information in the same form as the image region in the CG module, and then the degree of overlap between each label and these image regions is calculated, and the labels related to the problem are retained:

[0085] {S o ,...,S m} = f region2patch (B o ,...,B n );

[0086] L i =v thres <{S o ,...,S m}∩{s o ,...,s k};

[0087] In the formula, {S o ,...,S m} is the converted regional location information, v thres is the lowest threshold, ∩ is the operation for calculating the degree of overlap, tags below this value are considered irrelevant and are discarded, while tags above this value are retained.

[0088] To reduce the error while increasing the diversity of information, the image region is resampled for calculation and repeated M times to obtain more accurate labels relevant to the problem.

[0089] Furthermore, the processing process of the OCR unit is:

[0090]

[0091] In the formula, is the oth token of the text information obtained in the image, O i is the character information extracted from the image, f OCR It is the text recognition unit.

[0092] Furthermore, the information output by the question-enhanced title generation unit, the question-enhanced tag screening unit, and the OCR unit is connected to obtain the final overall text query after conversion, which is specifically:

[0093] T i ={C i ,L i ,O i ,Q i};

[0094] Where, T i For the final text query.

[0095] Furthermore, using the DPR retriever to retrieve top-k relevant knowledge from the text corpus includes:

[0096] Given a text query T i and question Q i First, use DPR to retrieve the top-k relevant knowledge from the text corpus. Then, based on this knowledge, prompt the large language model to generate questions. After re-ranking the probability of the large language model generating corresponding questions based on the knowledge (the greater the probability, the more relevant this knowledge is to answer this question), the re-ranked knowledge P is obtained. i ,

[0097] The most widely used retriever is DPR, which uses a BERT model with two different parameters to encode knowledge and queries respectively, calculate the similarity score between them, and retrieve knowledge related to the question. This process is illustrated by the following formula:

[0098]

[0099] Where r is the relevance score between the query and the knowledge. The top k pieces of knowledge are selected based on this relevance score. The relevance between the query and the knowledge is calculated based on their semantic similarity, which may not be sufficient to answer the question. To address this problem, this embodiment proposes a reverse reasoning re-ranking mechanism to re-rank the knowledge retrieved by DPR, thereby providing rich cross-attention between knowledge and questions. p(P i |Q i ). This example uses mathematical principles to prove the feasibility of this method. i |Q i ) Apply Bayes' rule:

[0100] logp(P i |Q i )=logp(Q i |P i )+logp(P i )+c;

[0101] In the formula, p(P i) is the prior probability of the retrieved knowledge, and c is a common constant.

[0102] Assume that the prior probability logp(P i ) is uniform and can be ignored during the reordering process, then the above expression can be simplified to:

[0103] logp(P i |Q i )∝logp(Q i |P i );

[0104] Then, the T5 model is used to calculate the average log-likelihood of the question tokens given the knowledge to estimate logp(T i |P i ):

[0105]

[0106] Where β is the parameter of the language model, |Q i | is the number of question tokens, q is any token in the question, P i It is the knowledge retrieved by DPR.

[0107] The above reranking method allows dataset-independent reranking and also incorporates cross-attention between question and paragraph tokens while forcing the model to interpret every token in the input question. Then, according to logp(Q i |P i ) to re-rank the knowledge retrieved by DPR. It is able to re-rank paragraphs by simply reasoning with an off-the-shelf language model without the need to annotate question-knowledge pairs for fine-tuning.

[0108] Furthermore, the processing process of the answer output module is:

[0109] answer = f T5 (T i ,P i );

[0110] Where, T i is the final text query, f T5 is the T5 language model, and answer is the output answer.

[0111] In terms of loss, QKRN uses autoregressive cross entropy loss to train the T5 model:

[0112]

[0113] In the formula, y i,j,w is the true answer, a i,j,wTo generate answers, N is the batch size, l is the answer length, and |V| is the vocabulary size.

[0114] Table 1 summarizes the experimental results of the QKRN model and existing methods on the OK-VQA dataset, and the following conclusions can be drawn: First, in most cases, the text space paradigm is better than the multimodal space paradigm, because the multimodal data used to train or fine-tune the multimodal model is much less than the pure text data, resulting in poor multimodal feature learning performance and poor performance in answering questions. Second, the method using a very large-scale language model (such as GPT-3) performs best on the OK-VQA dataset, because GPT-3 has more parameters than general language models, and the training and fine-tuning of this model require huge computing resources. Finally, the question-enhanced knowledge retrieval network QKRN performs best in the text space paradigm, because QQC uses the cross-attention mechanism to integrate the problems ignored by the existing methods into the process of information transformation, and RIR promotes a deeper interaction between articles and questions, which is conducive to retrieving more accurate knowledge to answer questions.

[0115] The external knowledge bases involved in Table 1, where C represents ConceptNet, W represents Wikipedia, GI represents Google Images, GS represents Google Search, and GPT-3 represents the knowledge accumulated when training using the model.

[0116] Table 1

[0117]

[0118] As shown in Table 2, QKRN has achieved very good results on the FVQA dataset compared to earlier works, which shows that large language models are able to learn well to answer commonsense VQA questions without access to the provided knowledge graph. Compared with other methods in FVQA, the proposed QKRN model achieves top 3 performance without explicitly designing to exploit the knowledge graph. This demonstrates the power of open-ended answer generation based on large language models. Better performance can be achieved by designing a more specialized retrieval component for the structured knowledge base used in this task. It can be seen that the QKRN model shows strong generalization in constructing different KB-VQA tasks.

[0119] Table 2

[0120]

[0121] like Figure 3As shown in the figure, two visualization examples are shown to illustrate the effectiveness of the method of this embodiment. The first row shows the heat map and output results in the question enhancement title generation module, proving that the method of this embodiment can better find relevant image areas for the problem and the obtained title information (green) is more targeted to the problem than the previous method (red). The second row shows the recognition box information and output results of the question enhancement label screening module, which also illustrates the targeting of this method to the problem.

[0122] The method of this embodiment does not require additional annotated data for training and can be extended to other text space paradigm models. This embodiment also conducted a large number of experiments, and the VQA score on OK-VQA was 54.14 points, and the accuracy on FVQA was 67.98 points, indicating that QKRN is effective for the latest works on KB-VQA.

[0123] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A knowledge visual question answering method based on question-enhanced knowledge retrieval network, characterized in that: include: Get the image to be detected and the corresponding question; Input the image to be detected and the corresponding question into the question-enhanced knowledge retrieval network for processing, and output the answer; The question-enhanced knowledge retrieval network is used to identify image regions related to the question using a cross-attention mechanism, generate picture captions related to the question, and retain picture targets related to the question.

2. The knowledge visual question answering method based on question-enhanced knowledge retrieval network according to claim 1 is characterized in that: The question-based enhanced knowledge retrieval network includes: Question enhancement query building module: used to obtain the final text query using several different levels of transformations on the image to be detected and the corresponding question; Backward Reasoning Re-ranking Retrieval Module: It is used to retrieve top-k relevant knowledge from the text corpus using the DPR retriever given a text query and question; Answer output module: used to connect the text output of the question enhancement query construction module and the knowledge retrieved by the reverse reasoning re-ranking retrieval module, and provide them to the T5 model for answer generation.

3. The knowledge visual question answering method based on question-enhanced knowledge retrieval network according to claim 2 is characterized in that: The question enhancement query building module includes: Question-enhanced title generation unit: used to obtain image title information related to the question; Question-enhanced label filtering unit: used to filter image label information related to questions based on images and questions; Text recognition unit: used to extract character information from images.

4. The knowledge visual question answering method based on question-enhanced knowledge retrieval network according to claim 3 is characterized in that: The processing process of the question enhanced title generation unit is as follows: In the formula, is the oth token of the generated image title information, I i For the image, Q i For the problem, C i is the image title information related to the question, f CG Enhanced title generation unit for questions; The processing process of the question-enhanced label screening unit is as follows: Where, L i is the image label information related to the question, To identify the target’s attribute information, To identify the label information of the target, f LF Enhanced tag filtering unit for questions; The processing process of the character recognition unit is as follows: In the formula, is the oth token of the text information obtained in the image, O i is the character information extracted from the image, f OCR It is the text recognition unit.

5. The knowledge visual question answering method based on question-enhanced knowledge retrieval network according to claim 4 is characterized in that: Obtaining the final text query includes: The information output by the question-enhanced title generation unit, the question-enhanced tag screening unit and the text recognition unit are connected to obtain the final overall text query after conversion, which is specifically: T i ={C i ,L i ,O i ,Q i }; Where, T i For the final text query.

6. The knowledge visual question answering method based on question-enhanced knowledge retrieval network according to claim 2, characterized in that: Retrieving top-k relevant knowledge from a text corpus using the DPR retriever includes: In the formula, r is the correlation score between query and knowledge, Q i For the problem, P i is the knowledge retrieved from DPR, is the question encoder, A knowledge encoder.

7. The knowledge visual question answering method based on question-enhanced knowledge retrieval network according to claim 6, characterized in that: After retrieving the top-k related knowledge, it also includes: The knowledge retrieved by DPR is reordered using the reverse reasoning reordering mechanism, and the average log-likelihood of the question tokens under the given knowledge is calculated using the T5 model, and logp(T i |P i ) to estimate: In the formula, β is the parameter of the language model, |Q i | is the number of question tokens, and q is any token in the question.

8. The knowledge visual question answering method based on question-enhanced knowledge retrieval network according to claim 2, characterized in that: The processing process of the answer output module is: answer=f T5 (T i ,P i ); Among them, T i For the final text query, P i is the knowledge retrieved from DPR, f T5 is the T5 language model, and answer is the output answer.

Citation Information

Patent Citations

  • Intelligent dialogue method, system and equipment based on large model and local knowledge base

    CN116992005A

  • Visual knowledge reasoning question and answer method for multi-source heterogeneous knowledge joint enhancement

    CN117010500A

  • Visual question and answer model and method based on multi-modal retrieval enhancement

    CN119066174A

  • System and method for question answering

    US20240104624A1