Retrieval enhancement generation method based on multi-modal document
By building a multimodal knowledge base and designing a multimodal knowledge searcher and answer generator, the problems of low accuracy and poor interpretability of knowledge-intensive visual question-and-answer in the prior art are solved, and higher accuracy and interpretability are achieved.
Patent Information
- Application Number
- CN202411867298.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-18
AI Technical Summary
The existing end-to-end model and search-enhanced generation systems have low accuracy and lack of interpretability when dealing with knowledge-intensive visual Q&A, especially because of the modal differences in the use of text documents as knowledge carriers, resulting in bottlenecks in the accuracy.
Using a search enhancement generation method based on multimodal documents, a multimodal knowledge base is constructed, text documents and entity pictures are combined into multimodal documents, and a multimodal knowledge searcher and answer generator are designed, and correlation calculation and answer generation are used to use multimodal features.
The accuracy and interpretability of knowledge-intensive visual question-and-answer tasks are improved, and the ability of knowledge retrieval and answer generation is enhanced by eliminating the modal differences between user input and knowledge carrier.
Smart Images

Figure CN119988542A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of knowledge-intensive visual question answering, and in particular, relates to a retrieval enhancement generation method based on multimodal documents. Background Art
[0002] Knowledge-intensive visual question answering is a subtask in visual question answering. Visual question answering is a task of answering relevant natural language questions based on given visual content, such as answering questions about elements in a picture based on a picture. It is a typical multimodal task involving fields such as deep learning, natural language processing, and multimodal learning. In addition to understanding the visual content, knowledge-intensive visual question answering also requires knowledge related to the visual content to correctly answer questions, which is more difficult than traditional visual question answering. This task has potential application scenarios in production and life, such as allowing users to ask knowledge-related questions about objects in photos based on the photos they took. The development of relevant technologies can facilitate users' information queries.
[0003] Existing technical solutions for solving knowledge-intensive visual question answering can be divided into two categories: end-to-end models and retrieval-augmented generation (RAG) systems. The end-to-end model solution uses a single model, such as a multimodal large language model, to handle visual question answering. The composition of a multimodal large language model usually includes a visual encoder, a large language model, and a visual-to-language mapping module. The visual encoder is used to encode visual content into visual features, which are then converted to the input space of the large language model through a mapping module. After training with a large amount of visual language corpus and instructions, the multimodal large language model can generate appropriate text replies based on visual content and text instructions, has visual question answering capabilities, and can correctly answer some knowledge-intensive image-text question answering questions.
[0004] In order to explicitly utilize knowledge when answering knowledge-intensive graphic question-answering questions, improve the accuracy of results and the interpretability of the system, a retrieval-enhanced generation system can be constructed. Retrieval-enhanced generation is a method of improving the performance of knowledge-intensive tasks by adding externally retrieved knowledge to the input of a large language model. In specific implementation, the architecture of a retrieval-enhanced generation system usually consists of a knowledge retriever and an answer generator. Existing retrieval-enhanced generation systems usually use text documents as knowledge carriers. The knowledge retriever matches user input and knowledge documents through a certain algorithm, sorts them by relevance, and returns a set of highly relevant documents. The answer generator usually uses a multimodal large language model, which can use question information and retrieved documents when generating answers, thereby achieving the purpose of utilizing knowledge.
[0005] The end-to-end model solution mainly relies on the capabilities of the multimodal large language model to answer questions, including the ability to understand vision, the ability to use knowledge reserves, and the ability to generate language. Knowledge-intensive visual question answering involves a wealth of visual entities and entity knowledge, which is difficult for a single model to accurately remember and use. Therefore, the end-to-end model solution usually has a low answer accuracy rate. Moreover, since it is usually a black box model, it is difficult to track and improve the specific logic of the model to generate answers, and the interpretability is poor.
[0006] Existing retrieval-enhanced generation systems usually use text documents as carriers of knowledge. When applied to knowledge-intensive text-and-picture question answering, the knowledge retriever needs to use visual information and text questions to retrieve the corresponding text documents. There are modal differences on both sides, which requires the retriever to have strong cross-modal retrieval capabilities. Therefore, the retriever trained using text documents as the knowledge base may be suboptimal. When generating answers, since the retriever may return multiple highly relevant documents, the answer generator needs to use the correct document to generate answers, which also requires the answer generator to have strong cross-modal information screening capabilities. Therefore, the retrieval-enhanced generation system using text documents has a bottleneck in accuracy when processing knowledge-intensive text-and-picture question answering.
[0007] The information disclosed in this background technology section is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as acknowledging or suggesting in any form that the information constitutes the prior art already known to those skilled in the art. Summary of the invention
[0008] In view of the problems existing in the prior art, the object of the present invention is to provide a retrieval enhancement generation method based on multimodal documents.
[0009] In order to achieve the above object, the present invention adopts the following technical solutions: A retrieval enhancement generation method based on multimodal documents, the retrieval enhancement generation method comprising the following steps: S1. Data construction: construct a multimodal knowledge base consisting of multimodal documents; S2. Feature extraction of multimodal knowledge retriever: The question image and question text in the user input are encoded into features by the image encoder and the text encoder respectively, where the image extracts global features and local features, and the text extracts text features; the document image in the multimodal document extracts global features, and the document text extracts text features; S3, feature mapping of multimodal knowledge retriever: map the features extracted in step S2 to the same dimension; the global features and local features of the image are mapped using multi-layer perceptron and Transformer network respectively, and the text is mapped using linear layer; S4, relevance calculation of the multimodal knowledge retriever: each feature on the user input side is dot-multiplied with all the features of the document, and the maximum value is taken as the relevance corresponding to the input feature, and then all the relevances are summed up as the final relevance; S5. Multimodal answer generation: Large language models generate text responses based on multimodal inputs.
[0010] Further, step S1 is specifically as follows: First, all text documents are matched with corresponding entity images to form multimodal documents. Then, encyclopedia page images are extracted in descending order of source accuracy, using encyclopedia image search and general image search until entity images are found. Finally, each text document in the original knowledge base is paired with a corresponding entity image to form a multimodal knowledge base.
[0011] Furthermore, in step S4, for the retrieval of multimodal documents, a mask interaction strategy for multimodal features is designed to optimize the performance, that is, the dot product value of the local features of the user input side image and the global features of the document side image is masked to negative infinity, in order to reduce the interference of irrelevant features on the relevance calculation; specifically: for the question text in the user input q and the problem image I , if a multimodal document contains text d and pictures I d , the correlation between the two r It can be calculated as: ; Among them, the matrix Q is q and I The feature matrix obtained by feature extraction and feature mapping on the user input side is D: d and I d The feature matrix obtained by feature extraction and feature mapping on the document side, l Q and l D are the number of features contained in Q and D respectively; mask represents the mask interaction strategy, which sets the dot product value of the local features of the user input side image and the global features of the document side image to negative infinity, and other values remain unchanged.
[0012] Furthermore, the process of answer generation in step S5 can be formally defined as follows: {image} represents the feature sequence after image mapping, {text} represents the text sequence, and the answer prediction process of the model is expressed as: ; The parameters of the model are θ,{text} o Represents the answer text sequence that needs to be predicted, which is obtained by the model autoregressively predicting the word sequence, including L words; {image} q and {text} q Respectively represent the question image and question text; {image} d 1 and {text} d 1 Respectively represent the image and text of the most relevant multimodal document, and so on into several multimodal documents; {text} inst Represents a reading instruction, and its prompt model answers based on the previous input; {text} o,<i Indicates the number of words that have been generated, {text} o,i Indicates the word currently being generated.
[0013] By adopting the above technical solution, the present invention has the following beneficial effects: This paper uses multimodal documents that combine images and text as knowledge carriers and designs a multimodal retrieval enhancement generation scheme. Compared with existing end-to-end model schemes, this scheme is based on a retrieval enhancement generation framework to ensure the accuracy and interpretability of answers; compared with retrieval enhancement generation schemes that use text documents as knowledge carriers, this scheme adds visual information to documents to construct multimodal documents, and improves the knowledge retriever and answer generator to utilize multimodal documents, thereby improving the accuracy of knowledge-intensive visual question-answering tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0015] Figure 1 A schematic diagram of a multimodal retrieval enhancement generation framework provided by the present invention; Figure 2 A schematic diagram of a multimodal answer generator provided by the present invention. DETAILED DESCRIPTION
[0016] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0017] The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the present invention, and is not used to limit the present invention.
[0018] For knowledge-intensive image-text question-answering tasks, the accuracy is low and lacks explainability when using end-to-end models; when using a retrieval-enhanced generation system with text documents as knowledge carriers, there are modal differences between the visual content and text questions input by the user and the text documents to be retrieved, so there is a bottleneck in accuracy.
[0019] Inspired by the process of humans identifying entities in images by comparing entity images, this paper proposes a retrieval enhancement generation system based on multimodal documents. This paper aims to use multimodal documents composed of entity images and text documents as knowledge carriers to eliminate the modal differences between user input and knowledge carriers, and specifically design corresponding knowledge retrievers and answer generators to utilize multimodal documents, in order to improve the accuracy of knowledge-intensive visual question answering.
[0020] Data construction part: The solution of this application first needs to build a multimodal knowledge base composed of multimodal documents. Since the current knowledge base of knowledge-intensive visual question and answer is usually composed of text documents, this application first needs to match all text documents with corresponding entity images to form multimodal documents. This application tries to extract encyclopedia page images in order from high to low accuracy of the source, using encyclopedia image search and general image search methods until the entity image is found. Finally, this application pairs each text document in the original knowledge base with a corresponding entity image to form a multimodal knowledge base.
[0021] The framework of multimodal retrieval enhancement generation in this application is as follows Figure 1 As shown. The framework of this application consists of a multimodal knowledge retriever and a multimodal answer generator, wherein the multimodal knowledge retriever consists of three parts: feature extraction, feature mapping, and relevance calculation; the multimodal answer generator generates answers by reading multiple multimodal documents. The following application will explain the above four parts in more detail.
[0022] Feature extraction of multimodal knowledge retriever: The question image and question text in the user input are encoded as features by the image encoder and text encoder respectively, where the image extracts global features and local features, and the text extracts text features. The encoders used are all pre-trained and can extract meaningful features from the input of the corresponding modality. For the document side, since it is a multimodal document, in addition to extracting the features of the text part, this application also extracts the global features of the document image.
[0023] Feature mapping of multimodal knowledge retriever: In order to perform subsequent relevance calculations, the features just extracted need to be mapped to the same dimension. The global features and local features of the image are mapped using multi-layer perceptrons and Transformer networks respectively, while the text is mapped using linear layers. After mapping, all features have the same dimension.
[0024] Relevance calculation of multimodal knowledge retriever: The calculation of the relevance between user input and multimodal documents generally follows the late interactive relevance calculation method. Each feature on the user input side is dot-multiplied with all the features of the document, and the maximum value is taken as the relevance corresponding to the input feature. Then all relevances are summed up as the final relevance. For the retrieval of multimodal documents, this application designs a masked interaction strategy for multimodal features to optimize performance, that is, the dot product value of the local features of the image on the user input side and the global features of the image on the document side is masked to negative infinity, in order to reduce the interference of irrelevant features on the relevance calculation. Specifically: For the retrieval of multimodal documents, a mask interaction strategy for multimodal features is designed to optimize performance, that is, the dot product value of the local features of the user input image and the global features of the document image is masked to negative infinity, in order to reduce the interference of irrelevant features on the relevance calculation; specifically: for the question text in the user input q and the problem image I , if a multimodal document contains text d and pictures I d , the correlation between the two r It can be calculated as: ; Among them, the matrix Q is q and I The feature matrix obtained by feature extraction and feature mapping on the user input side is D: d and I d The feature matrix obtained by feature extraction and feature mapping on the document side, l Q and l Dare the number of features contained in Q and D respectively; mask represents the mask interaction strategy, which sets the dot product value of the local features of the user input side image and the global features of the document side image to negative infinity, and other values remain unchanged.
[0025] Multimodal Answer Generator: Figure 2 The model structure of the multimodal answer generator is presented in more detail. When constructing the input of the multimodal large language model, the image is mapped to the input space of the large language model through the visual encoder and the visual-to-language mapping module, and is interlaced with the text to form a multimodal input. The large language model generates a text response based on the multimodal input. When training the model as an answer generator, the present application also adds several multimodal documents with the highest relevance after inputting the question image and question text to form an interlaced image and text input, and constructs reading samples to train the model to generate answers.
[0026] The process of answer generation can be formally defined as follows: {image} represents the feature sequence after image mapping, {text} represents the text sequence, and the answer prediction process of the model is expressed as: ; The parameters of the model are θ ,{text} o Represents the answer text sequence that needs to be predicted, which is obtained by the model autoregressively predicting the word sequence, including L words; {image} q and {text} q Respectively represent the question image and question text; {image} d 1 and {text} d 1 Respectively represent the image and text of the most relevant multimodal document, and so on into several multimodal documents; {text} inst Represents a reading instruction, and its prompt model answers based on the previous input; {text} o,<i Indicates the number of words that have been generated, {text} o,i Indicates the word currently being generated.
[0027] This application will use an example to further illustrate that when the user uses Figure 1When a multimodal query (MultimodalQuery) in the figure asks when the building in the figure will be opened, the multimodal retriever of the present application will find the top K multimodal documents (Top K Multimodal Documents) with the highest relevance from the knowledge base. The top multimodal document corresponds to the correct building "BMW Welt" and contains its opening date October 17, 2007, and the second multimodal document corresponds to the incorrect building "Europa building" and contains its planned completion year 2012. The multimodal answer generator will generate an answer based on the multimodal query and the top K multimodal documents, such as Figure 2 As shown, the model gives the correct answer, which is 2007. This completes the process.
[0028] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A retrieval enhancement generation method based on multimodal documents, characterized in that: The retrieval enhancement generation method comprises the following steps: S1. Data construction: construct a multimodal knowledge base consisting of multimodal documents; S2. Feature extraction of multimodal knowledge retriever: The question image and question text in the user input are encoded into features by the image encoder and the text encoder respectively, where the image extracts global features and local features, and the text extracts text features; the document image in the multimodal document extracts global features, and the document text extracts text features; S3, feature mapping of multimodal knowledge retriever: map the features extracted in step S2 to the same dimension; the global features and local features of the image are mapped using multi-layer perceptron and Transformer network respectively, and the text is mapped using linear layer; S4, relevance calculation of the multimodal knowledge retriever: each feature on the user input side is dot-multiplied with all the features of the document, and the maximum value is taken as the relevance corresponding to the input feature, and then all the relevances are summed up as the final relevance; S5. Multimodal answer generation: Large language models generate text responses based on multimodal inputs.
2. The retrieval enhancement generation method based on multimodal documents according to claim 1 is characterized in that: Step S1 is specifically as follows: First, all text documents are matched with corresponding entity images to form multimodal documents. Then, encyclopedia page images are extracted in descending order of source accuracy, using encyclopedia image search and general image search until entity images are found. Finally, each text document in the original knowledge base is paired with a corresponding entity image to form a multimodal knowledge base.
3. The retrieval enhancement generation method based on multimodal documents according to claim 1 is characterized in that: In step S4, for the retrieval of multimodal documents, a mask interaction strategy for multimodal features is designed to optimize the performance, that is, the dot product value of the local features of the user input side image and the global features of the document side image is masked to negative infinity, in order to reduce the interference of irrelevant features on the relevance calculation; specifically: for the question text in the user input q and the problem image I , if a multimodal document contains text d and pictures I d , the correlation between the two r Calculated as: ; Among them, the matrix Q is q and I The feature matrix obtained by feature extraction and feature mapping on the user input side is D: d and I d The feature matrix obtained by feature extraction and feature mapping on the document side, l Q and l D are the number of features contained in Q and D respectively; mask represents the mask interaction strategy, which sets the dot product value of the local features of the user input side image and the global features of the document side image to negative infinity, and other values remain unchanged.
4. The retrieval enhancement generation method based on multimodal documents according to claim 1 is characterized in that: The process of answer generation in step S5 can be formally defined as follows: {image} represents the feature sequence after image mapping, {text} represents the text sequence, and the answer prediction process of the model is expressed as: ; The parameters of the model are θ ,{text} o Represents the answer text sequence that needs to be predicted, which is obtained by the model autoregressively predicting the word sequence, including L words; {image} q and {text} q Respectively represent the question image and question text; {image} d 1 and {text} d 1 Respectively represent the image and text of the most relevant multimodal document, and so on into several multimodal documents; {text} inst Represents a reading instruction, and its prompt model answers based on the previous input; {text} o,<i Indicates the number of words that have been generated, {text} o,i Indicates the word currently being generated.
Citation Information
Patent Citations
Image-text multi-modal feature representation method and system based on alignment and fusion
CN115661594A
Visual question and answer model and method based on multi-modal retrieval enhancement
CN119066174A
Systems and methods for a vision-language pretraining framework
US20240160853A1
Generative pretraining of multimodal retrieval-augmented visual-language models
WO2024118578A1
Cited By
Retrieval enhancement generation method and device, equipment, medium and product
CN120780818A
Energy field intelligent knowledge base question-answering system based on multi-model cooperation
CN120950657A
Speech recognition method and device based on retrieval enhancement generation and medium
CN121331133A