Multi-modal retrieval enhancement generation method and device and medium

By constructing a feature vector knowledge base and using a multimodal processing model, the problems of modal understanding bias, limited generalization ability and difficulty in modal fusion in VQA technology are solved, and more efficient document retrieval and answer generation are achieved.

CN119961382APending Publication Date: 2025-05-09INSPUR GENERSOFT CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510064620.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Existing VQA technologies are in understanding the problems of deviations between visual and linguistic modalities, limited generalization ability and the difficulty of deep fusion caused by the modal divide.

Method used

By connecting multiple document databases to build a feature vector knowledge base, obtain user questions and process input images, questions and document content based on image feature vector coding models, semantic models and multimodal big models to generate document answers.

Benefits of technology

It improves the ability of visual question-and-answer modal understanding, generalization and cross-modal fusion, and enhances the accuracy and efficiency of document retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961382A_ABST
    Figure CN119961382A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal retrieval enhancement generation method and device and a medium, and belongs to the technical field of data processing. The method comprises the following steps: connecting a plurality of document databases, and constructing a feature vector knowledge base based on the plurality of document databases; obtaining a user question; wherein the user questions comprise input questions and input images; processing the input image and a feature vector knowledge base based on a preset image feature vector coding model to determine a related document set; processing the user input question and the related document set based on a preset semantic model to determine a target document; and processing the user question, the related document set and the target document based on a preset multi-modal large model to generate a document answer. According to the method, the visual question and answer modal understanding, generalization and cross-modal fusion capabilities are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and in particular to a multimodal retrieval enhancement generation method, device and medium. Background Art

[0002] With the rapid development of artificial intelligence technology, visual question answering (VQA), as an important bridge connecting computer vision and natural language processing, has received widespread attention in recent years. VQA systems are designed to understand image content and accurately answer image-related questions, and are widely used in smart education, human-computer interaction, autonomous driving and other fields.

[0003] However, existing VQA technology still faces many challenges in practical applications. First, there is a significant understanding bias between visual and language modalities. Especially when dealing with complex, abstract problems or problems that require in-depth contextual understanding, the system often finds it difficult to accurately integrate visual and language information. Secondly, the generalization ability of the VQA system is limited. When faced with novel or unseen visual scenes and questioning methods, the system performance drops sharply. In addition, the existence of the modality gap makes the deep integration of visual and language information particularly difficult, further restricting the development of VQA technology.

[0004] Therefore, how to improve the visual question answering modal understanding, generalization, and cross-modal fusion capabilities has become a technical problem that needs to be solved urgently. Summary of the invention

[0005] The embodiments of the present application provide a multimodal retrieval enhancement generation method, device and medium to solve the following technical problem: how to improve the visual question answering modal understanding, generalization and cross-modal fusion capabilities.

[0006] In a first aspect, an embodiment of the present application provides a multimodal retrieval enhancement generation method, the method comprising: connecting multiple document databases, and constructing a feature vector knowledge base based on the multiple document databases; obtaining user questions; wherein the user questions include input questions and input images; processing the input images and the feature vector knowledge base based on a preset image feature vector encoding model to determine a related document set; processing the user input questions and the related document set based on a preset semantic model to determine a target document; processing the user questions, the related document set and the target document based on a preset multimodal large model to generate document answers.

[0007] In one implementation of the present application, multiple document databases are connected, and a feature vector knowledge base is constructed based on the multiple document databases, specifically including: connecting multiple document databases, and identifying document titles and paragraph contents in the multiple document databases to construct a knowledge base; using a semantic model to encode the document title and paragraph content respectively to generate a document title feature vector and a paragraph content feature vector; associating the document title feature vector, the paragraph content feature vector, the document title and the paragraph content to construct a feature vector knowledge base.

[0008] In one implementation of the present application, an input image and a feature vector knowledge base are processed based on a preset image feature vector encoding model to determine a set of related documents, specifically including: processing the input image based on the image feature vector encoding model to generate a visual feature vector; matching the similarity of document title feature vectors in a feature vector knowledge base based on the visual feature vector to determine multiple first similarities; and selecting, in order of the first similarities from largest to smallest, document contents corresponding to a preset first number of document title feature vectors to form a first related document.

[0009] In one implementation of the present application, an input image is processed based on an image feature vector encoding model to generate a visual feature vector, specifically including: segmenting the input image to generate multiple image blocks; wherein the multiple image blocks do not overlap with each other; processing the multiple image blocks based on a preset linear transformation algorithm to determine multiple image vectors; and adding sequences to the multiple image vectors to generate a visual feature vector.

[0010] In one implementation of the present application, a user input question and a related document set are processed based on a preset semantic model to determine a target document, specifically including: processing the input question based on the semantic model to generate a question text feature vector; matching the similarity of paragraph content feature vectors in the related document set based on the question text feature vector to determine multiple second similarities; and selecting paragraph content target documents corresponding to a preset second number of majority paragraph content feature vectors in order of the second similarities from large to small.

[0011] In one implementation of the present application, an input question is processed based on a semantic model to generate a question text feature vector, specifically including: decomposing the input question based on the semantic model to generate word groups; comparing the word groups with a preset word vector database to generate a word vector group; adding a sequence to the word vector group to generate a question text feature vector.

[0012] In one implementation of the present application, user questions, related document sets and target documents are processed based on a preset multimodal large model to generate document answers: the input image, user questions and related documents are combined into a comprehensive input prompt word; the comprehensive input prompt word is input into the multimodal large model, and the multimodal large model integrates the comprehensive input prompt words based on the attention mechanism to generate a document answer.

[0013] In one implementation of the present application, after the comprehensive input prompt words are input into the multimodal large model, and the multimodal large model integrates the comprehensive input prompt words based on the attention mechanism to generate a document answer, the method also includes: recording the document answer to a user database corresponding to the user's question; retrieving the user's usage habits corresponding to the user's question based on the user database; processing the document answer and the comprehensive input prompt words based on the user's usage habits to generate an answer.

[0014] In the second aspect, an embodiment of the present application also provides a multimodal retrieval enhancement generation device, the device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: connect to multiple document databases, and build a feature vector knowledge base based on the multiple document databases; obtain user questions; wherein the user questions include input questions and input images; process the input images and the feature vector knowledge base based on a preset image feature vector encoding model to determine a set of related documents; process the user input questions and the set of related documents based on a preset semantic model to determine a target document; process the user questions, the set of related documents and the target document based on a preset multimodal large model to generate document answers.

[0015] In a third aspect, an embodiment of the present application also provides a non-volatile computer storage medium for multimodal retrieval enhanced generation, storing computer executable instructions, wherein the computer executable instructions are configured to: connect multiple document databases, and construct a feature vector knowledge base based on the multiple document databases; obtain user questions; wherein the user questions include input questions and input images; process the input images and the feature vector knowledge base based on a preset image feature vector encoding model to determine a set of related documents; process the user input questions and the set of related documents based on a preset semantic model to determine a target document; process the user questions, the set of related documents, and the target document based on a preset multimodal large model to generate document answers.

[0016] The embodiments of the present application provide a multimodal search enhancement generation method, device, and medium, which at least include the following technical effects:

[0017] By connecting multiple document databases and building a feature vector knowledge base, a large amount of document information from different sources can be integrated to form a comprehensive knowledge base. To a certain extent, the document content can be efficiently indexed and retrieved. By processing the input image and feature vector knowledge base through the preset image feature vector encoding model, the document set related to the input image content can be quickly located, thereby improving the accuracy and efficiency of document retrieval. After obtaining the relevant document set, the user input question is processed through the preset semantic model, and the target document is further determined in combination with the document content, thereby selecting targeted document content to form the target document to a certain extent. The user questions, related document sets and target documents are processed through the preset multimodal large model, so that the generated document answers not only conform to reading habits, but also improve the visual question answering modal understanding, generalization and cross-modal fusion capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0019] Figure 1 A flow chart of a modality retrieval enhancement generation method provided in an embodiment of the present application;

[0020] Figure 2 A schematic diagram of the internal structure of a modal retrieval enhancement generation device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0022] The embodiments of the present application provide a multimodal retrieval enhancement generation method, device and medium to solve the following technical problem: how to improve the visual question answering modal understanding, generalization and cross-modal fusion capabilities.

[0023] The technical solution proposed in the embodiments of the present application is described in detail below with reference to the accompanying drawings.

[0024] Figure 1 A flow chart of modal retrieval enhancement generation provided in an embodiment of the present application. Figure 1 As shown, a modal retrieval enhancement generation method provided in an embodiment of the present application specifically includes the following steps:

[0025] Step 1: Connect multiple document databases and build a feature vector knowledge base based on the multiple document databases.

[0026] First, multiple document databases are connected and the document titles and paragraph contents in multiple document databases are identified to build a knowledge base. Database connection technology (such as JDBC) is used to establish connections with multiple document databases according to the access protocols of each document database. Documents in the document database are traversed to extract the title and paragraph contents of each document.

[0027] Further, the document title and paragraph content are respectively encoded using the semantic model to generate a document title feature vector and a paragraph content feature vector. The document title and paragraph content are encoded according to a preset semantic model to generate corresponding feature vectors, namely, a document title feature vector and a paragraph content feature vector.

[0028] Furthermore, the document title feature vector, the paragraph content feature vector, the document title and the paragraph content are associated to construct a feature vector knowledge base. In the constructed feature vector knowledge base, an entry is created for each document title and paragraph content. The corresponding feature vectors (document title feature vector, paragraph content feature vector) and the original text (document title, paragraph content) are stored in the entry. It can be understood that the association between the document title feature vector, the paragraph content feature vector, the document title and the paragraph content is bidirectional, that is, the corresponding text can be found through the feature vector, and its feature vector can be quickly located through the text. All generated document title feature vectors, paragraph content feature vectors and the corresponding document titles and paragraph contents are imported into this knowledge base.

[0029] Step 2: Obtain user questions; wherein the user questions include input questions and input images.

[0030] Input problem: This is the text problem entered by the user.

[0031] Input image: image data imported by the user.

[0032] For example, a user imports a flowchart and asks, "Convert the above flowchart into text information according to the logic, and the number of words must be controlled within 400 words."

[0033] Step 3: Process the input image and feature vector knowledge base based on a preset image feature vector encoding model to determine a set of related documents.

[0034] First, the input image is processed based on the image feature vector encoding model to generate a visual feature vector, which is specifically described from A1 to A3.

[0035] A1. Segment the input image to generate multiple image blocks; wherein the multiple image blocks do not overlap. Use image segmentation algorithms (such as sliding windows, grid division, etc.) to segment the input image into multiple small blocks of the same size or divided according to specific rules. It is understandable that overlapping image blocks will cause overlapping problems, so it should be ensured that the segmented image blocks do not overlap each other so that they can be processed independently later.

[0036] A2. Process multiple image blocks based on a preset linear transformation algorithm to determine multiple image vectors. Apply the preset linear transformation algorithm to each image block. The linear transformation algorithm converts the pixel values ​​of the image block into a higher-level feature representation, namely an image vector, through a combination of weights and biases. Each image block will generate a corresponding image vector that captures the unique features of the image block.

[0037] A3. Add sequences to multiple image vectors to generate visual feature vectors. Combine all image vectors in a certain order (such as spatial order during segmentation) and add sequence information (such as position encoding) to each image vector to help the model understand the relative position relationship between vectors. The generated visual feature vector is a high-dimensional vector that contains the global and local features of the input image.

[0038] Further, based on the similarity of the visual feature vector to the feature vector of the document title in the feature vector knowledge base, a plurality of first similarities are determined. The similarity between the visual feature vector and each document title feature vector in the knowledge base is calculated (e.g., a cosine similarity algorithm). A plurality of first similarities are generated, each of which corresponds to the degree of similarity between a document title feature vector and a visual feature vector.

[0039] Further, according to the order of the first similarity from large to small, the document contents corresponding to the preset first number of document title feature vectors are selected to form a related document set. The documents are sorted according to the order of the first similarity from large to small. The first preset number of documents with the highest similarity are selected, and the document title feature vectors of these documents are most similar to the visual feature vector of the input image. The contents of these documents are extracted to form a final related document set.

[0040] Step 4: Process the user input question and the feature vector knowledge base based on a preset semantic model to determine a second relevant document.

[0041] First, the input question is processed based on the semantic model to generate a question text feature vector. The specific process is described from B1 to B3.

[0042] B1. Decompose the input question based on the semantic model to generate word groups. Decompose the original question string input by the user into a series of characters, words or phrases (i.e. word groups) through natural language processing technology.

[0043] B2. Compare the preset word vector database based on the word group to generate a word vector group. After obtaining the word group, compare these words with the preset word vector database. The word vector database usually contains a large number of pre-trained word vectors, each of which is a numerical representation of the word in a specific semantic space, which can capture the semantic relationship between words. Through search and mapping, each word will be converted into a vector of fixed dimension to form a word vector group.

[0044] B3. Add the sequence to the word vector group to generate the question text feature vector. In order to capture the order and relationship between the words in the question, the system will arrange the word vector group according to the order in the original question to generate a question text feature vector that can represent the entire question semantics.

[0045] Further, based on the similarity of the question text feature vector matching the paragraph content feature vector in the relevant document set, multiple second similarities are determined. After obtaining the question text feature vector, the question text feature vector is similar to the text feature vector of each paragraph in the feature vector knowledge base (i.e., the relevant document set). That is, the paragraph closest to the question semantics can be found.

[0046] Further, in descending order of the second similarity, the paragraph contents corresponding to the preset second number of paragraph content feature vectors are selected as the target documents. All paragraphs are sorted according to the calculated similarity (i.e., the second similarity), and the documents corresponding to the first preset number of paragraphs with the highest similarity are selected as the target documents.

[0047] Step 5: Process the user question, related document set and target document based on the preset multimodal large model to generate a document answer.

[0048] First, the input image, user question and related documents are combined into a comprehensive input prompt word. The visual and text feature vectors of all objects in the prompt word are collected and concatenated together to form a comprehensive feature vector.

[0049] Furthermore, the comprehensive input prompt words are input into the multimodal large model, and the multimodal large model integrates the comprehensive input prompt words based on the attention mechanism to generate document answers.

[0050] Attention mechanism: The attention mechanism in the multimodal large model can automatically learn the relevance and importance between different input elements (such as key areas in an image, keywords in a text). Through the attention weight, the model can focus on the information that is most helpful in generating answers.

[0051] Based on the information integrated by the attention mechanism, a comprehensive document answer is generated. After the document answer is generated, in order to facilitate user reading, the following methods are also included:

[0052] First, record the document answer to the user database corresponding to the user's question. First, each user corresponds to a user database, which includes information such as user questions, corresponding document answers, and timestamps. When a user asks a question, the system retrieves the answer from the document or knowledge base and records the question and its answer (including multiple candidate answers and the final selected answer) in the database.

[0053] Furthermore, the user usage habits corresponding to the user's question are retrieved based on the user database. Based on the user's historical query records, answer selections and other data, a personalized user profile is constructed. When a user asks a new question, the system first retrieves historical query records similar to the current question and their corresponding user behavior data from the user database.

[0054] Furthermore, the document answers and comprehensive input prompt words are processed based on the user's usage habits to generate answers. The natural language algorithm is used to understand and parse the user's new questions and comprehensive input prompt words, as well as historical queries and questions. Based on the user's usage habits (such as whether the user prefers a short and direct answer or a detailed and comprehensive answer), multiple candidate answers extracted from the document are intelligently sorted or filtered. Based on the sorted or filtered candidate answers, the final answer is generated according to the natural language algorithm.

[0055] The above is an embodiment of the method proposed in this application. Based on the same inventive concept, this application embodiment also provides a modal retrieval enhancement generation device, whose structure is as follows: Figure 2 shown.

[0056] Figure 2 A schematic diagram of the internal structure of a modal retrieval enhancement generation device provided in an embodiment of the present application. Figure 2 As shown, the device includes:

[0057] at least one processor 201;

[0058] and, a memory 202 communicatively connected to the at least one processor;

[0059] The memory 202 stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor 201 to enable at least one processor 201 to:

[0060] Connect multiple document databases and build a feature vector knowledge base based on the multiple document databases; obtain user questions; wherein the user questions include input questions and input images; process the input images and the feature vector knowledge base based on a preset image feature vector encoding model to determine a related document set; process the user input questions and the related document set based on a preset semantic model to determine a target document; process the user questions, the related document set and the target document based on a preset multimodal large model to generate document answers.

[0061] Some embodiments of the present application provide corresponding Figure 1 A non-volatile computer storage medium for enhancing the generation of modal retrieval. The non-volatile computer storage medium stores computer executable instructions, and the computer executable instructions are set to:

[0062] Connect multiple document databases and build a feature vector knowledge base based on the multiple document databases; obtain user questions; wherein the user questions include input questions and input images; process the input images and the feature vector knowledge base based on a preset image feature vector encoding model to determine a related document set; process the user input questions and the related document set based on a preset semantic model to determine a target document; process the user questions, the related document set and the target document based on a preset multimodal large model to generate document answers.

[0063] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the IoT device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0064] The system and medium provided in the embodiments of the present application correspond one-to-one to the method. Therefore, the system and medium also have similar beneficial technical effects to the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the system and medium will not be repeated here.

[0065] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0066] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0067] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0068] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0069] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0070] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0071] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0072] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0073] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A multimodal retrieval enhancement generation method, characterized in that: The method comprises: Connecting multiple document databases and building a feature vector knowledge base based on the multiple document databases; Obtaining a user question; wherein the user question includes an input question and an input image; Processing the input image and the feature vector knowledge base based on a preset image feature vector encoding model to determine a set of related documents; Processing the user input question and the related document set based on a preset semantic model to determine a target document; The user question, related document set and target document are processed based on a preset multimodal macro model to generate a document answer.

2. A multimodal retrieval enhancement generation method according to claim 1, characterized in that: Connecting multiple document databases and building a feature vector knowledge base based on the multiple document databases, specifically including: Connecting multiple document databases and identifying document titles and paragraph contents in the multiple document databases to build a knowledge base; Using the semantic model to encode the document title and paragraph content respectively to generate a document title feature vector and a paragraph content feature vector; The document title feature vector, the paragraph content feature vector, the document title and the paragraph content are associated to construct a feature vector knowledge base.

3. A multimodal retrieval enhancement generation method according to claim 2, characterized in that: The input image and the feature vector knowledge base are processed based on a preset image feature vector encoding model to determine a set of related documents, specifically including: Processing the input image based on an image feature vector encoding model to generate a visual feature vector; Based on the visual feature vector, the similarity of matching the feature vector of the document title in the feature vector knowledge base is determined to determine a plurality of first similarities; According to the order of the first similarities from large to small, document contents corresponding to a preset first number of document title feature vectors are selected to form a related document set.

4. A multimodal retrieval enhancement generation method according to claim 3, characterized in that: Processing the input image based on the image feature vector encoding model to generate a visual feature vector specifically includes: Segmenting the input image to generate a plurality of image blocks; wherein the plurality of image blocks do not overlap with each other; Processing the plurality of image blocks based on a preset linear transformation algorithm to determine a plurality of image vectors; A sequence is added to a plurality of the image vectors to generate a visual feature vector.

5. A multimodal retrieval enhancement generation method according to claim 2, characterized in that: Processing the user input question and the related document set based on a preset semantic model to determine a target document specifically includes: Processing the input question based on the semantic model to generate a question text feature vector; Based on the similarity of matching the question text feature vector to the paragraph content feature vector in the relevant document set, a plurality of second similarities are determined; In descending order of the second similarities, paragraph contents corresponding to a preset second number of the paragraph content feature vectors are selected as target documents.

6. A multimodal retrieval enhancement generation method according to claim 5, characterized in that: Processing the input question based on the semantic model to generate a question text feature vector specifically includes: Decomposing the input question based on the semantic model to generate word groups; Comparing the word group with a preset word vector database to generate a word vector group; Add the sequence to the word vector group to generate a question text feature vector.

7. A multimodal retrieval enhancement generation method according to claim 1, characterized in that: The user question, the related document set and the target document are processed based on a preset multimodal macro model to generate a document answer: Combining the input image, user question and related documents into a comprehensive input prompt word; The comprehensive input prompt words are input into the multimodal large model, and the multimodal large model integrates the comprehensive input prompt words based on the attention mechanism to generate a document answer.

8. A multimodal retrieval enhancement generation method according to claim 7, characterized in that: After the comprehensive input prompt words are input into the multimodal large model, and the multimodal large model integrates the comprehensive input prompt words based on the attention mechanism to generate a document answer, the method further includes: Recording the answer to the document in a user database corresponding to the user's question; Retrieving the user usage habits corresponding to the user question based on the user database; The document answer and the comprehensive input prompt words are processed based on the user's usage habits to generate an answer.

9. A multimodal retrieval enhancement generation device, characterized in that: The device comprises: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: Connecting multiple document databases and building a feature vector knowledge base based on the multiple document databases; Obtaining a user question; wherein the user question includes an input question and an input image; Processing the input image and the feature vector knowledge base based on a preset image feature vector encoding model to determine a set of related documents; Processing the user input question and the related document set based on a preset semantic model to determine a target document; The user question, related document set and target document are processed based on a preset multimodal macro model to generate a document answer.

10. A non-volatile computer storage medium for multimodal retrieval enhancement generation, storing computer executable instructions, characterized in that: The computer executable instructions are configured to: Connecting multiple document databases and building a feature vector knowledge base based on the multiple document databases; Obtaining a user question; wherein the user question includes an input question and an input image; Processing the input image and the feature vector knowledge base based on a preset image feature vector encoding model to determine a set of related documents; Processing the user input question and the related document set based on a preset semantic model to determine a target document; The user question, related document set and target document are processed based on a preset multimodal macro model to generate a document answer.

Citation Information

Cited By

  • Visual question and answer method and device based on multiple document images

    CN120508686A

  • Method for processing multi-modal mixed data and enhancing model generation effect

    CN120541243A

  • Multi-modal data automatic processing and information extraction method and system

    CN120763346A

  • Large model retrieval enhancement generation method and device based on multi-modal knowledge enhancement and medium

    CN121390325A

  • Multi-modal fusion output retrieval question and answer method and device, storage medium and computer program product

    CN121478912A