Hierarchical retrieval method and device based on multi-modal questions and answers and computer equipment

By introducing a hierarchical search method in the multimodal question and answer system, using image feature vectors for text retrieval and integration of results, the problem of inefficiency in the utilization of visual information and knowledge integration of traditional systems is solved, and more efficient multimodal question and answer performance and a wider range of applicable scenarios are achieved.

CN120045683AInactive Publication Date: 2025-05-27CHINA ELECTRONICS RELIABILITY AND ENVIRONMENTAL TESTING INSTITUTE ((THE FIFTH INSTITUTE OF ELECTRONICS MINISTRY OF INDUSTRY AND INFORMATION TECHNOLOGY) (CHINA SAIBAO LABORATORY)

Patent Information

Application Number
CN202510480776.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional multimodal question and answer systems are inefficient in utilizing visual information, complex knowledge integration and rely on reward models, resulting in limited applicable scenarios.

Method used

A hierarchical search method based on multimodal question and answer is provided. By receiving image data and question data to be answered, the preset knowledge base and multimodal input data are encoded to generate feature vectors and image feature vectors of document data. Then, text search is performed in the knowledge base based on image feature vectors, search results are integrated and multimodal question and answer models are input to generate answers.

Benefits of technology

It improves the search effect and knowledge integration efficiency, enhances the performance and flexibility of the multimodal question and answer system, can better adapt to different types of multimodal question and answer tasks, and improves the generalization ability and applicable scenario range of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045683A_ABST
    Figure CN120045683A_ABST
Patent Text Reader

Abstract

The invention relates to a hierarchical retrieval method and device based on multi-modal questions and answers, computer equipment, a computer readable storage medium and a computer program product. The method comprises the steps of receiving multi-modal input data, wherein the multi-modal input data comprises image data and corresponding to-be-answered question data; performing encoding processing on the document data and the multi-modal input data in the preset knowledge base to obtain a feature vector of the document data in the preset knowledge base and an image feature vector of the multi-modal input data; according to feature vectors of document data in a preset knowledge base, performing at least one time of text retrieval in the preset knowledge base for the image feature vectors to obtain a text retrieval result; and integrating the text retrieval results into an input sequence, inputting the input sequence into the multi-mode question and answer model, and generating a corresponding answer text of the to-be-answered question data. By adopting the method, the retrieval effect and the knowledge integration efficiency can be improved, and the performance and the flexibility of the multi-modal question-answering system are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of large model technology, and in particular to a hierarchical retrieval method, apparatus, computer device, computer-readable storage medium and computer program product based on multimodal question and answer. Background Art

[0002] With the development of artificial intelligence, large models (LLMs) have made remarkable achievements in the field of natural language processing, but they face challenges such as knowledge limitations, hallucination problems, and data security in professional fields. Retrieval-augmented generation technology (RAG) came into being. By retrieving relevant knowledge and incorporating prompts, large models can combine knowledge to generate more reasonable answers. The RAG architecture includes the data preparation stage (data extraction, text segmentation, vectorization, data storage) and the application stage (data retrieval, prompt injection), which can effectively solve the problem of insufficient knowledge of large models in specific fields.

[0003] The development of multimodal technology has further promoted the emergence of multimodal RAG. Multimodal models can handle multiple data types (such as text, images, etc.) and align data of different modalities through joint embedding strategies. Multimodal RAG allows the system to retrieve information of multiple modalities and generate answers by combining multimodal data such as text and images, which improves the performance and generalization ability of question-answering systems; however, traditional retrieval strategies do not make sufficient use of visual information, knowledge integration is complex and relies on reward models, and the applicable scenarios are limited, resulting in overall low efficiency.

[0004] To address these issues, there is an urgent need for a hierarchical retrieval method, apparatus, computer equipment, computer-readable storage medium, and computer program product based on multimodal question and answering that can improve retrieval results and knowledge integration efficiency and enhance the performance and flexibility of the multimodal question and answering system. Summary of the invention

[0005] Based on this, it is necessary to provide a hierarchical retrieval method, device, computer equipment, computer-readable storage medium and computer program product based on multimodal question and answering, which can improve the retrieval effect and knowledge integration efficiency, and enhance the performance and flexibility of the multimodal question and answering system in response to the above-mentioned technical problems.

[0006] In a first aspect, the present application provides a hierarchical retrieval method based on multimodal question answering, comprising:

[0007] Receiving multimodal input data, the multimodal input data comprising image data and corresponding question data to be answered;

[0008] Encoding the document data in the preset knowledge base and the multimodal input data to obtain a feature vector of the document data in the preset knowledge base and an image feature vector of the multimodal input data;

[0009] According to the feature vector of the document data in the preset knowledge base, at least one text search is performed in the preset knowledge base for the image feature vector to obtain a text search result;

[0010] The text retrieval results are integrated into an input sequence, and the input sequence is input into a multimodal question-answering model to generate a corresponding answer text for the question data to be answered.

[0011] In one embodiment, the feature vector of the document data includes a feature vector of a document title; performing at least one text search in the preset knowledge base for the image feature vector based on the feature vector of the document data in the preset knowledge base to obtain a text search result includes:

[0012] According to the feature vector of the document title in the preset knowledge base, a preliminary text search is performed in the preset knowledge base for the image feature vector to obtain k documents in the preset knowledge base that are related to the image feature vector;

[0013] Segmenting the k documents to obtain multiple text blocks;

[0014] According to the multiple text blocks, a secondary text search is performed in a preset knowledge base for the question data to be answered to obtain a secondary text search result.

[0015] In one embodiment, the secondary text search is performed in a preset knowledge base for the question data to be answered based on the multiple text blocks to obtain the secondary text search results, including:

[0016] Using a text encoder, encoding the multiple text blocks and the question data to be answered, and generating multiple text block vectors and question vectors of preset dimensions respectively;

[0017] According to the multiple text block vectors, a secondary text search is performed in a preset knowledge base for the question vector to be answered to obtain a secondary text search result.

[0018] In one embodiment, the performing of a secondary text search in a preset knowledge base for the question vector to be answered based on the multiple text block vectors to obtain a secondary text search result includes:

[0019] Calculate the similarity between each text block vector and the question vector to be answered to obtain a second similarity result;

[0020] According to the second similarity result, a secondary text search is performed in a preset knowledge base for the question vector to be answered to obtain a secondary text search result, wherein the secondary text search result includes n text blocks related to the question data to be answered.

[0021] In one embodiment, the step of performing a preliminary text search in the preset knowledge base for the image feature vector based on the feature vector of the document title in the preset knowledge base to obtain k documents in the preset knowledge base related to the image feature vector includes:

[0022] Calculating an inner product between the image feature vector and the feature vector of the document title as a first similarity result;

[0023] According to the first similarity result, an approximate k-nearest neighbor search algorithm is used to perform a preliminary text search in a preset knowledge base for the image feature vector to obtain k documents related to the image feature vector in the preset knowledge base.

[0024] In one embodiment, the input sequence includes a text block vector, an image feature vector, a system-level prompt, and question data to be answered; and inputting the input sequence into a multimodal question-answering model to generate a corresponding answer text for the question data to be answered includes:

[0025] The image feature vector and the text block vector are used as context information, and the system-level prompt and the question data to be answered are used as question-answering input tasks, and are input into a multimodal question-answering model to generate a corresponding answer text for the question data to be answered;

[0026] A data mixing strategy is adopted to fine-tune the multimodal question-answering model based on the corresponding answer text and the multimodal input data.

[0027] In a second aspect, the present application also provides a hierarchical retrieval device based on multimodal question and answer, comprising:

[0028] A data receiving module, used for receiving multimodal input data, wherein the multimodal input data includes image data and corresponding question data to be answered;

[0029] An encoding processing module, used to perform encoding processing on the document data in the preset knowledge base and the multimodal input data to obtain a feature vector of the document data in the preset knowledge base and an image feature vector of the multimodal input data;

[0030] A text retrieval module, used to perform at least one text search in a preset knowledge base for the image feature vector according to the feature vector of the document data in the preset knowledge base, to obtain a text retrieval result;

[0031] The answer output module is used to integrate the text retrieval results into an input sequence, input the input sequence into the multimodal question-answering model, and generate the corresponding answer text for the question data to be answered.

[0032] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0033] Receiving multimodal input data, the multimodal input data comprising image data and corresponding question data to be answered;

[0034] Encoding the document data in the preset knowledge base and the multimodal input data to obtain a feature vector of the document data in the preset knowledge base and an image feature vector of the multimodal input data;

[0035] According to the feature vector of the document data in the preset knowledge base, at least one text search is performed in the preset knowledge base for the image feature vector to obtain a text search result;

[0036] The text retrieval results are integrated into an input sequence, and the input sequence is input into a multimodal question-answering model to generate a corresponding answer text for the question data to be answered.

[0037] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0038] Receiving multimodal input data, the multimodal input data comprising image data and corresponding question data to be answered;

[0039] Encoding the document data in the preset knowledge base and the multimodal input data to obtain a feature vector of the document data in the preset knowledge base and an image feature vector of the multimodal input data;

[0040] According to the feature vector of the document data in the preset knowledge base, at least one text search is performed in the preset knowledge base for the image feature vector to obtain a text search result;

[0041] The text retrieval results are integrated into an input sequence, and the input sequence is input into a multimodal question-answering model to generate a corresponding answer text for the question data to be answered.

[0042] In a fifth aspect, the present application further provides a computer program product, including a computer program, which implements the following steps when executed by a processor:

[0043] Receiving multimodal input data, the multimodal input data comprising image data and corresponding question data to be answered;

[0044] Encoding the document data in the preset knowledge base and the multimodal input data to obtain a feature vector of the document data in the preset knowledge base and an image feature vector of the multimodal input data;

[0045] According to the feature vector of the document data in the preset knowledge base, at least one text search is performed in the preset knowledge base for the image feature vector to obtain a text search result;

[0046] The text retrieval results are integrated into an input sequence, and the input sequence is input into a multimodal question-answering model to generate a corresponding answer text for the question data to be answered.

[0047] The hierarchical retrieval method, device, computer equipment, computer-readable storage medium and computer program product based on multimodal question and answer receive image data and question data to be answered, encode the preset knowledge base and multimodal input data, and obtain the feature vector of document data and the image feature vector. Subsequently, text retrieval is performed in the knowledge base based on the image feature vector, and the retrieval results are integrated and input into the multimodal question and answer model to generate answers. Specifically, through encoding processing, the image and text information are converted into feature vectors, and the alignment and fusion of multimodal data are realized, so that the question and answer system can process image and text input at the same time, and improve the ability to understand complex problems. Text retrieval based on image feature vectors can accurately locate the knowledge base content related to the image, provide more accurate background knowledge for the question and answer model, and make up for the lack of knowledge of the large model in specific fields. After the retrieval results are integrated into an input sequence and input into the multimodal question and answer model, the knowledge integration process is simplified, the efficiency and quality of generating answers are improved, and the overall performance of the system is enhanced. Through the collaborative processing and precise retrieval of multimodal information, the technology can better adapt to different types of multimodal question and answer tasks, and improve the generalization ability and applicable scenario range of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0049] Figure 1 is an application environment diagram of a hierarchical retrieval method based on multimodal question answering in one embodiment;

[0050] Figure 2 is a flowchart of a hierarchical retrieval method based on multimodal question and answer in one embodiment;

[0051] Figure 3 is a flowchart of a hierarchical retrieval method based on multimodal question and answer in another embodiment;

[0052] Figure 4 It is a flowchart of a hierarchical retrieval method based on multimodal question and answer in the most detailed embodiment;

[0053] Figure 5 is a structural block diagram of a hierarchical retrieval device based on multimodal question and answer in one embodiment;

[0054] Figure 6 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0056] The hierarchical retrieval method based on multimodal question answering provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the server 102 communicates with the server 104 via a network. The data storage system can store data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed on the cloud or other network servers.

[0057] The server 104 receives multimodal input data, which includes image data and corresponding question data to be answered; encodes the document data in the preset knowledge base and the multimodal input data to obtain the feature vector of the document data in the preset knowledge base and the image feature vector of the multimodal input data; performs at least one text search in the preset knowledge base for the image feature vector based on the feature vector of the document data in the preset knowledge base to obtain a text search result; integrates the text search result into an input sequence, inputs the input sequence into the multimodal question-answering model of the server 102, and generates the corresponding answer text of the question data to be answered.

[0058] The server 102 may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, IoT devices, and portable wearable devices. The IoT devices may be smart speakers, smart TVs, smart air conditioners, smart car devices, projection devices, etc. Portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. Head-mounted devices may be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server 104 may be an independent physical server, or a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides cloud computing services.

[0059] In an exemplary embodiment, Figure 2 As shown in the figure, a hierarchical retrieval method based on multimodal question answering is provided. Figure 1 The server 104 in the example is used as an example to illustrate the method, which includes the following steps S202 to S208. Among them:

[0060] Step S202, receiving multimodal input data, where the multimodal input data includes image data and corresponding question data to be answered.

[0061] Specifically, multimodal input data refers to input containing multiple types of data, including image data and question data to be answered.

[0062] The image data can be photos, charts, illustrations, or other visual content. These images provide visual information to help the model understand the background or specific content of the question. The question data to be answered refers to the text questions associated with the image data, which are usually specific questions that the user wants the model to answer. The content of the question may be directly related to the image, or it may require a combination of the image and other knowledge to answer.

[0063] Specifically, image data provides visual information to help the model better understand the context or specific content of the question. For example, in a visual question answering task, an image may contain objects, scenes, or charts that need to be recognized. The question data clarifies the user's needs and guides the model to extract relevant information from images and other knowledge sources to generate targeted answers.

[0064] Step S204: Encoding the document data in the preset knowledge base and the multimodal input data to obtain feature vectors of the document data in the preset knowledge base and image feature vectors of the multimodal input data.

[0065] Specifically, encoding is the process of converting data of different modalities into feature vectors that can be understood and processed by computers. Specifically, a text encoder (such as BERT, Contriever, etc.) is used to process the document data in the knowledge base and convert it into feature vectors. These feature vectors can represent the semantic information of the document, which is convenient for subsequent similarity calculation and retrieval.

[0066] Use visual encoders (such as CLIP, ALIGN, etc.) to process the input image data and extract the image feature vector. The image feature vector can represent the visual content of the image, such as objects, scenes, colors, and other information.

[0067] The preset knowledge base is collected from multimodal documents from multiple channels such as the Internet, professional literature databases, and encyclopedias. The document content covers professional knowledge, cultural knowledge, and expert experience in various fields. The document form is a (document, image, text title) triple, where the document contains a detailed text description, the image is related to the document content, and the text title is used to quickly locate the document topic. The collected documents are cleaned to remove noise data and irrelevant information.

[0068] Extract key information from documents, such as text paragraphs, images and their descriptions, to form structured knowledge units for subsequent retrieval and fusion. Perform preprocessing operations such as word segmentation and part-of-speech tagging on the text content of the document to facilitate the calculation of similarity between texts.

[0069] The feature vector of document data is used to represent the semantic information of text in the knowledge base. In the retrieval stage, by calculating the similarity between feature vectors, the document fragments most relevant to the user's question can be quickly found. The image feature vector is used to represent the visual content of the input image. In the retrieval stage, the image feature vector can be similarly calculated with the document title or other text content in the knowledge base to help the system find knowledge fragments related to the image content.

[0070] Step S206: Based on the feature vector of the document data in the preset knowledge base, at least one text search is performed in the preset knowledge base for the image feature vector to obtain a text search result.

[0071] Specifically, the retrieval process is a key link in the multimodal question answering system, and its purpose is to find the most relevant text content from the knowledge base that is most relevant to the input image and question. The retrieval process is usually divided into two stages:

[0072] Through rough search, the search scope is quickly narrowed down to find a collection of documents related to the input image. Through fine search, the paragraphs or text fragments most relevant to the user's question are further found in the documents obtained through rough search. The output of the retrieval process is the text fragments most relevant to the input image and question, which are used as retrieval results for subsequent question-answering generation.

[0073] Step S208, integrating the text retrieval results into an input sequence, inputting the input sequence into a multimodal question-answering model, and generating corresponding answer texts for the question data to be answered.

[0074] Specifically, the most relevant text fragments are selected from the text retrieval results, and these text fragments are arranged in a certain logical order to form a coherent context. For example, they can be spliced ​​into a long text, or sorted by relevance. Some auxiliary information, such as the title or source of the document, can also be added to enhance the completeness and credibility of the context.

[0075] The input sequence is the input to the multimodal question answering model, which is a large pre-trained language model capable of processing text and images. It processes the input sequence in the following way:

[0076] The image feature vector and the text feature vector are fused together to form a unified semantic representation. For example, the image features are converted into a form compatible with the text features through the MLP (Multi-layer Perceptron) Adapter. The model understands the background and requirements of the question based on the input context information (retrieved text fragments, image features, and question text). Based on the understood context, the model generates the answer text for the user's question.

[0077] Finally, the answer text output by the multimodal question answering model is a direct answer to the user's question. This answer combines image information and knowledge base information to generate a more accurate and detailed answer.

[0078] In the above-mentioned hierarchical retrieval method based on multimodal question and answer, by receiving image data and question data to be answered, the preset knowledge base and multimodal input data are encoded and processed to obtain the feature vector of the document data and the image feature vector. Subsequently, text retrieval is performed in the knowledge base based on the image feature vector, and the retrieval results are integrated and input into the multimodal question and answer model to generate answers. Specifically, through encoding processing, the image and text information are converted into feature vectors, and the alignment and fusion of multimodal data are realized, so that the question and answer system can process image and text input at the same time, and improve the ability to understand complex problems. Text retrieval based on image feature vectors can accurately locate the knowledge base content related to the image, provide more accurate background knowledge for the question and answer model, and make up for the lack of knowledge of the large model in specific fields. After the retrieval results are integrated into the input sequence and input into the multimodal question and answer model, the knowledge integration process is simplified, the efficiency and quality of generating answers are improved, and the overall performance of the system is enhanced. Through the collaborative processing and precise retrieval of multimodal information, this technology can better adapt to different types of multimodal question and answer tasks, and improve the generalization ability and applicable scenario range of the system.

[0079] In an exemplary embodiment, Figure 3 As shown, the feature vector of the document data includes the feature vector of the document title; according to the feature vector of the document data in the preset knowledge base, at least one text search is performed in the preset knowledge base for the image feature vector to obtain the text search result, including:

[0080] Step S302, based on the feature vector of the document title in the preset knowledge base, a preliminary text search is performed in the preset knowledge base for the image feature vector to obtain k documents in the preset knowledge base that are related to the image feature vector;

[0081] Step S304, segmenting the k documents to obtain multiple text blocks;

[0082] Step S306: performing a secondary text search in a preset knowledge base for the question data to be answered based on the multiple text blocks to obtain a secondary text search result.

[0083] Specifically, first, a preliminary text search is performed to calculate the similarity (such as cosine similarity) between the image feature vector and the feature vector of the Chinese document title in the preset knowledge base. Based on the similarity score, the top k documents most relevant to the image feature vector are returned, and the topics of these documents are highly relevant to the image content.

[0084] Next, a secondary text search is performed to segment each document into multiple text blocks (such as paragraphs or sentences). The segmentation method can be based on semantic integrity (such as paragraph segmentation) or fixed-length segmentation. Each text block contains local semantic information of the document.

[0085] In short, the result of coarse retrieval is k documents related to the image, and the result of fine retrieval is n text blocks in these documents that are most relevant to the user's question.

[0086] In this embodiment, a hierarchical search strategy is used to quickly locate relevant documents first, and then accurately search within the documents, thereby avoiding the inefficiency of complex searches directly in a large-scale knowledge base. Coarse search uses image feature vectors to quickly filter out relevant documents, and fine search combines question text to further filter out the most relevant text blocks to ensure that the search results are highly relevant to the question. Image feature vectors play a guiding role in coarse search, ensuring that the search results are relevant to the image content; question text plays a screening role in fine search, ensuring that the search results are consistent with the question semantics.

[0087] In an exemplary embodiment, a secondary text search is performed in a preset knowledge base for the question data to be answered based on the multiple text blocks to obtain secondary text search results, including:

[0088] Using a text encoder, encoding the multiple text blocks and the question data to be answered, and generating multiple text block vectors and question vectors of preset dimensions respectively;

[0089] According to the multiple text block vectors, a secondary text search is performed in a preset knowledge base for the question vector to be answered, and a secondary text search result is obtained.

[0090] Specifically, a text encoder is a tool that converts text data into fixed-dimensional vectors that can capture the semantic information of the text. Common text encoders include:

[0091] BERT: Encodes text through the Transformer architecture to capture the contextual information of the text. Contriever: A text encoder optimized for retrieval tasks that can generate vectors suitable for similarity calculations.

[0092] (1) Encode the text block: Use the text encoder to encode each text block and generate a vector of fixed dimension (text block vector). Each text block is converted into a vector that captures the semantic information of the text block.

[0093] (2) Encode the question data to be answered: Use the same text encoder to encode the question data to be answered raised by the user to generate a question vector that captures the semantic information of the user's question.

[0094] Next, multiple text block vectors (T1, T2, ..., Tn) and the question vector to be answered (Q) are taken as input. The similarity between each text block vector and the question vector is calculated. Common similarity calculation methods include: cosine similarity, dot product and Euclidean distance, and the output is the similarity score between each text block and the question (second similarity result).

[0095] In this embodiment, the accuracy of retrieval and the efficiency of knowledge utilization are significantly improved through fine-grained text block matching and semantic similarity calculation. It not only provides more accurate context information for the multimodal question-answering model to generate more accurate answers, but also improves the overall performance and flexibility of the system, enabling it to better adapt to the needs of complex multimodal question-answering tasks.

[0096] In an exemplary embodiment, based on multiple text block vectors, a secondary text search is performed in a preset knowledge base for the question vector to be answered, and a secondary text search result is obtained, including:

[0097] Calculate the similarity between each text block vector and the question vector to be answered to obtain a second similarity result;

[0098] According to the second similarity result, a secondary text search is performed in a preset knowledge base for the question vector to be answered to obtain a secondary text search result, which includes n text blocks related to the question data to be answered.

[0099] Specifically, the k documents obtained by the rough search are divided into multiple text blocks, and each text block is converted into a text block vector by an encoder (such as Contriever). The question text raised by the user is processed by the text encoder to obtain the question vector to be answered. The similarity between each text block vector and the question vector is calculated. Commonly used similarity calculation methods include: Cosine similarity: measures the angle between two vectors in the semantic space. The closer the value is to 1, the more similar they are. Euclidean distance: measures the straight-line distance between two vectors. The smaller the value, the more similar they are. Dot product: directly calculates the inner product of two vectors. The larger the value, the more similar they are.

[0100] The result is a second similarity result, that is, the similarity score between each text block and the question. According to the second similarity result, n text blocks that are most relevant to the question are selected from all text blocks. These text blocks are further screened from the k documents obtained from the rough search and are closer to the semantic requirements of the user's question.

[0101] In this embodiment, by calculating the second similarity between the text block and the question vector, the fine search can filter out the text block that is most relevant to the question, ensuring that the answers generated subsequently have higher accuracy and relevance. The text blocks filtered out by the fine search are the most relevant fragments in the knowledge base that are most relevant to the question. These fragments are input into the multimodal question-answering model as context information, which can help the model better understand the background of the question and generate more accurate answers. Compared with the entire document returned by the coarse search, the text blocks returned by the fine search are more focused on the core content of the question, avoiding the interference of irrelevant information.

[0102] In an exemplary embodiment, according to the feature vector of the document title in the preset knowledge base, a preliminary text search is performed in the preset knowledge base for the image feature vector to obtain k documents related to the image feature vector in the preset knowledge base, including:

[0103] Calculate the inner product between the image feature vector and the document title feature vector as the first similarity result;

[0104] According to the first similarity result, an approximate k-nearest neighbor search algorithm is used to perform a preliminary text search in a preset knowledge base for the image feature vector, and k documents related to the image feature vector in the preset knowledge base are obtained.

[0105] Specifically, the first similarity result (the inner integral of each document title and the image feature vector). Use the approximate k-nearest neighbor search algorithm to quickly find the k documents most similar to the image feature vector from the knowledge base. The approximate k-NN algorithm is an efficient retrieval method that can quickly find the k vectors most similar to the query vector in a large-scale data set while reducing the computational complexity. Return the k documents most relevant to the image feature vector.

[0106] In this embodiment, by calculating the similarity between the image feature vector and the document title, the rough search can quickly find the document set related to the input image and narrow the scope of the subsequent search. By calculating the inner product between the image feature vector and the feature vector of the document title, a first similarity result is obtained. According to the first similarity result, the approximate k-NN algorithm is used to find the k documents most relevant to the image feature vector. Using the approximate k-NN algorithm, the retrieval task can be completed efficiently in a large-scale knowledge base, avoiding the inefficiency of performing complex calculations directly in all documents. The k documents obtained by the rough search will be used as the input of the fine search to further screen out the text blocks most relevant to the user's question.

[0107] In an exemplary embodiment, the input sequence includes a text block vector, an image feature vector, a system-level prompt, and question data to be answered; the input sequence is input into a multimodal question-answering model to generate a corresponding answer text for the question data to be answered, including:

[0108] The image feature vector and the text block vector are used as context information, and the system-level prompts and the question data to be answered are used as the question-answering input task. They are input into the multimodal question-answering model to generate the corresponding answer text for the question data to be answered.

[0109] A data mixing strategy is adopted to fine-tune the multimodal question-answering model based on the corresponding answer text and multimodal input data.

[0110] Specifically, the input sequence is the input of the multimodal question-answering model and contains a variety of information to help the model generate accurate answers. Specifically, the input sequence includes the following parts:

[0111] (1) An image feature vector is a vector obtained by encoding an input image using a visual encoder (such as CLIP). It represents the visual content of the image and helps the model understand the visual context of the problem.

[0112] (2) Text block vectors refer to the text blocks that are most relevant to the user's question obtained through detailed retrieval. These text blocks are encoded into vectors through a text encoder (such as Contriever). They provide knowledge background related to the question and help the model generate more accurate answers.

[0113] (3) System-level prompts refer to instructions used to guide the model in generating answers, which may include question type, answer format requirements, task description, etc. They help the model understand the specific requirements of the task and improve the accuracy and relevance of the generated answers.

[0114] (4) Question data refers to the specific question text raised by the user. As the core input of the question-answering task, it guides the model to generate targeted answers.

[0115] Image feature vectors and text block vectors as context information: These vectors provide visual and textual context related to the question, helping the model understand the background knowledge of the question. System-level prompts and question data as Q&A input tasks: System-level prompts and question text directly guide the model to generate answers. All the above information is integrated into an input sequence and input into the multimodal question-answering model. The model generates answers to user questions through internal mechanisms (such as feature fusion, context understanding, etc.).

[0116] In order to further optimize the performance of the multimodal question-answering model, the system adopts a data mixing strategy for fine-tuning. Specifically, question-answer pairs that require external knowledge to answer are mixed with question-answer pairs that do not require external knowledge to form a fine-tuning dataset.

[0117] In this example, different types of data are mixed to ensure that the model performs well when dealing with questions that require external knowledge while maintaining performance on other tasks. The multimodal question answering model is fine-tuned using the above mixed dataset. The parameters of the model are optimized so that it can better utilize the retrieved text blocks and image features to generate accurate answers. Through fine-tuning, the performance of the model in multimodal question answering tasks is improved, especially when dealing with questions that require external knowledge.

[0118] like Figure 4 As shown, the most detailed embodiment of the present application is:

[0119] 1. Implementation environment:

[0120] The hierarchical retrieval architecture in this embodiment is applied in a network environment including a server and a terminal. The server is responsible for processing multimodal input data and generating answers, and the terminal is used for user interaction and data display. The specific environment is as follows:

[0121] Server: Deploys a multimodal question-answering system, including a visual encoder, a hierarchical retrieval module, and a multimodal question-answering model.

[0122] Terminal: Users upload images and questions through the terminal, and receive and display answers generated by the system.

[0123] Knowledge base: stores multimodal documents, including text, images, and title information, for retrieval and knowledge supplementation.

[0124] 2. Multimodal input data processing:

[0125] Receiving multimodal input data:

[0126] The system receives multimodal input data uploaded by users, including image data and corresponding question data to be answered.

[0127] Image encoding:

[0128] Use the pre-trained CLIP visual encoder to encode the image data and extract the image feature vector. The image feature vector can represent the visual content of the image, such as objects, scenes, colors, textures, and other information.

[0129] Question code:

[0130] The CLIP text encoder is used to encode user questions and generate question vectors. The question vector captures the semantic information of the question and is used for subsequent retrieval and answer generation.

[0131] 3. Hierarchical search:

[0132] Rough search:

[0133] Document title encoding: Encode the titles of all documents in the knowledge base using the CLIP text encoder to generate a title feature vector.

[0134] Similarity calculation: Calculate the inner product between the image feature vector and each document title feature vector as a similarity measure.

[0135] Approximate k-nearest neighbor search: Use the approximate k-NN algorithm to quickly find the answer output module 508k answer output module 508 documents that are most similar to the image feature vector from the knowledge base. The returned answer output module 508k answer output module 508 documents are used as coarse retrieval results for subsequent fine retrieval.

[0136] Detailed search:

[0137] Document segmentation: Each document obtained by rough retrieval is segmented into multiple text blocks, each of which contains a certain amount of text content.

[0138] Text block encoding: Use the Contriever architecture to encode each text block and generate a text block vector. Figure 4 As shown, R2: Inspection contents include appearance inspection, electrical performance test, functional test, X-ray inspection, etc.; R3: Appearance inspection to see if the pins are broken and the solder joints are cold. X-ray inspection of the internal structure to see if there are wafer cracks or line short circuits;

[0139] Similarity calculation: Calculate the inner product between each text block vector and the question vector as a similarity measure.

[0140] Text block selection: Based on the similarity results, the answer output module 508n answer output module 508 text blocks that are most relevant to the question are selected as detailed search results. These text blocks contain the key information required to answer the question.

[0141] 4. Input Integration and Answer Generation:

[0142] Input Integration:

[0143] Integrate image feature vectors, text block vectors, system-level prompts (such as instructions for question-answering tasks, format requirements, etc.) and user questions into a complete input sequence. Image feature vectors and text block vectors are used as context information, and system-level prompts and user questions are used as input for question-answering tasks. Alternatively, all retrieved knowledge fragments can be directly concatenated to form a long text, which is input into the large model as additional context information. This can simplify the steps of knowledge fusion and reduce errors that may be introduced by weight calculation and sorting. At the same time, the large model itself has certain text understanding and generation capabilities, and can extract useful information from a longer context and generate accurate answers.

[0144] Answer generation:

[0145] Use a pre-trained multimodal question-answering model (such as a Transformer-based model) to process the input sequence. The model converts the image feature vector into a form compatible with text features through the MLP answer output module 508Adapter, realizing the fusion of visual information and text information.

[0146] The model generates answer text for user questions based on the integrated input sequence. For example, whether the appearance pins are deformed and the solder joints are intact. The wafer has no cracks and the circuit has no short circuits. It is recommended to conduct high-temperature aging tests in the next step to observe the performance changes of the chip at high temperatures, so as to determine the cause of failure and improvement measures.

[0147] 5. Model fine-tuning and optimization:

[0148] Data blending strategy:

[0149] The question-answer pairs that require external knowledge to answer are mixed with the general question-answering data to form a fine-tuning dataset. The data mixing method ensures that the model performs well on questions that require external knowledge while maintaining performance on other tasks.

[0150] Fine-tuning training:

[0151] During fine-tuning, the model is trained on a mixed dataset to optimize its ability to exploit the retrieved knowledge fragments. A combination of zero-shot learning and fine-tuning training is used to quickly deploy and optimize model performance.

[0152] Search technology optimization:

[0153] Use efficient retrieval algorithms (such as approximate nearest neighbor search) to speed up the retrieval process and improve retrieval efficiency. Embed document titles and text blocks to optimize retrieval accuracy and speed.

[0154] 6. Implementation Effect:

[0155] Through the above hierarchical retrieval architecture and implementation method, the system can efficiently retrieve the most relevant knowledge fragments for the input image and question from a large-scale knowledge base and generate accurate and rich answers. The specific effects are as follows:

[0156] Improved search accuracy: Through hierarchical search strategies, the system can gradually narrow the search scope to ensure that the search results are highly relevant to the question.

[0157] Knowledge integration optimization: It simplifies the knowledge integration process, reduces errors introduced by improper knowledge processing, and improves the quality of answers.

[0158] Wide range of applicable scenarios: This architecture is suitable for a variety of multimodal question-answering tasks, especially visual question-answering tasks that require external knowledge assistance, and has high flexibility and scalability.

[0159] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0160] Based on the same inventive concept, the embodiment of the present application also provides a hierarchical retrieval device based on multimodal question and answer for implementing the hierarchical retrieval method based on multimodal question and answer involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more embodiments of the hierarchical retrieval device based on multimodal question and answer provided below can refer to the limitations of the hierarchical retrieval method based on multimodal question and answer above, and will not be repeated here.

[0161] In an exemplary embodiment, Figure 5 As shown, a hierarchical retrieval device based on multimodal question and answer is provided, comprising:

[0162] A data receiving module 502 is used to receive multimodal input data, where the multimodal input data includes image data and corresponding question data to be answered;

[0163] The encoding processing module 504 is used to perform encoding processing on the document data in the preset knowledge base and the multimodal input data to obtain a feature vector of the document data in the preset knowledge base and an image feature vector of the multimodal input data;

[0164] A text search module 506 is used to perform at least one text search in the preset knowledge base for the image feature vector according to the feature vector of the document data in the preset knowledge base, and obtain a text search result;

[0165] The answer output module 508 is used to integrate the text retrieval results into an input sequence, input the input sequence into the multimodal question-answering model, and generate the corresponding answer text for the question data to be answered.

[0166] In an exemplary embodiment, the feature vector of the document data includes a feature vector of the document title;

[0167] The text retrieval module 506 is also used to perform a preliminary text search in the preset knowledge base for the image feature vector based on the feature vector of the document title in the preset knowledge base, and obtain k documents related to the image feature vector in the preset knowledge base; segment the k documents to obtain multiple text blocks; and perform a secondary text search in the preset knowledge base for the question data to be answered based on the multiple text blocks to obtain a secondary text retrieval result.

[0168] In an exemplary embodiment, the text retrieval module 506 is also used to use a text encoder to encode multiple text blocks and question data to be answered, and respectively generate multiple text block vectors and question vectors of preset dimensions; based on the multiple text block vectors, a secondary text search is performed in a preset knowledge base for the question vectors to obtain a secondary text retrieval result.

[0169] In an exemplary embodiment, the text retrieval module 506 is also used to calculate the similarity between each text block vector and the question to be answered vector to obtain a second similarity result; based on the second similarity result, a secondary text search is performed in a preset knowledge base for the question to be answered vector to obtain a secondary text retrieval result, and the secondary text retrieval result includes n text blocks related to the question to be answered data.

[0170] In an exemplary embodiment, the text retrieval module 506 is also used to calculate the inner product between the image feature vector and the feature vector of the document title as a first similarity result; based on the first similarity result, an approximate k-nearest neighbor search algorithm is used to perform preliminary text retrieval in a preset knowledge base for the image feature vector, and obtain k documents related to the image feature vector in the preset knowledge base.

[0171] In an exemplary embodiment, the input sequence includes text block vectors, image feature vectors, system-level prompts, and question data to be answered;

[0172] The text retrieval module 506 is also used to use the image feature vector and the text block vector as context information, and the system-level prompts and the question data to be answered as the question-answering input task, and input them together into the multimodal question-answering model to generate the corresponding answer text for the question data to be answered; and adopt a data mixing strategy to fine-tune the multimodal question-answering model based on the corresponding answer text and the multimodal input data.

[0173] Each module in the hierarchical retrieval device based on multimodal question and answer can be implemented in whole or in part by software, hardware and their combination. Each module can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each module above.

[0174] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 6As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store multimodal input data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a hierarchical retrieval method based on multimodal question and answer is implemented.

[0175] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0176] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:

[0177] Receiving multimodal input data, the multimodal input data including image data and corresponding question data to be answered;

[0178] Encoding the document data in the preset knowledge base and the multimodal input data to obtain a feature vector of the document data in the preset knowledge base and an image feature vector of the multimodal input data;

[0179] According to the feature vector of the document data in the preset knowledge base, at least one text search is performed in the preset knowledge base for the image feature vector to obtain a text search result;

[0180] The text retrieval results are integrated into an input sequence, which is then fed into a multimodal question-answering model to generate the corresponding answer text for the question data to be answered.

[0181] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:

[0182] The feature vector of the document data includes the feature vector of the document title;

[0183] According to the feature vector of the document title in the preset knowledge base, a preliminary text search is performed in the preset knowledge base for the image feature vector to obtain k documents related to the image feature vector in the preset knowledge base;

[0184] Segment k documents to obtain multiple text blocks;

[0185] According to the multiple text blocks, a secondary text search is performed in a preset knowledge base for the question data to be answered, and a secondary text search result is obtained.

[0186] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:

[0187] Using a text encoder, encoding the multiple text blocks and the question data to be answered, and generating multiple text block vectors and question vectors of preset dimensions respectively;

[0188] According to the multiple text block vectors, a secondary text search is performed in a preset knowledge base for the question vector to be answered, and a secondary text search result is obtained.

[0189] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:

[0190] Calculate the similarity between each text block vector and the question vector to be answered to obtain a second similarity result;

[0191] According to the second similarity result, a secondary text search is performed in a preset knowledge base for the question vector to be answered to obtain a secondary text search result, which includes n text blocks related to the question data to be answered.

[0192] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:

[0193] Calculate the inner product between the image feature vector and the document title feature vector as the first similarity result;

[0194] According to the first similarity result, an approximate k-nearest neighbor search algorithm is used to perform a preliminary text search in a preset knowledge base for the image feature vector, and k documents related to the image feature vector in the preset knowledge base are obtained.

[0195] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:

[0196] The input sequence includes text block vectors, image feature vectors, system-level prompts, and question data to be answered;

[0197] The image feature vector and the text block vector are used as context information, and the system-level prompts and the question data to be answered are used as the question-answering input task. They are input into the multimodal question-answering model to generate the corresponding answer text for the question data to be answered.

[0198] A data mixing strategy is adopted to fine-tune the multimodal question-answering model based on the corresponding answer text and multimodal input data.

[0199] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0200] Receiving multimodal input data, the multimodal input data including image data and corresponding question data to be answered;

[0201] Encoding the document data in the preset knowledge base and the multimodal input data to obtain a feature vector of the document data in the preset knowledge base and an image feature vector of the multimodal input data;

[0202] According to the feature vector of the document data in the preset knowledge base, at least one text search is performed in the preset knowledge base for the image feature vector to obtain a text search result;

[0203] The text retrieval results are integrated into an input sequence, which is then fed into a multimodal question-answering model to generate the corresponding answer text for the question data to be answered.

[0204] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0205] The feature vector of the document data includes the feature vector of the document title;

[0206] According to the feature vector of the document title in the preset knowledge base, a preliminary text search is performed in the preset knowledge base for the image feature vector to obtain k documents related to the image feature vector in the preset knowledge base;

[0207] Segment k documents to obtain multiple text blocks;

[0208] According to the multiple text blocks, a secondary text search is performed in a preset knowledge base for the question data to be answered, and a secondary text search result is obtained.

[0209] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0210] Using a text encoder, encoding the multiple text blocks and the question data to be answered, and generating multiple text block vectors and question vectors of preset dimensions respectively;

[0211] According to the multiple text block vectors, a secondary text search is performed in a preset knowledge base for the question vector to be answered, and a secondary text search result is obtained.

[0212] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0213] Calculate the similarity between each text block vector and the question vector to be answered to obtain a second similarity result;

[0214] According to the second similarity result, a secondary text search is performed in a preset knowledge base for the question vector to be answered to obtain a secondary text search result, which includes n text blocks related to the question data to be answered.

[0215] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0216] Calculate the inner product between the image feature vector and the document title feature vector as the first similarity result;

[0217] According to the first similarity result, an approximate k-nearest neighbor search algorithm is used to perform a preliminary text search in a preset knowledge base for the image feature vector, and k documents related to the image feature vector in the preset knowledge base are obtained.

[0218] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0219] The input sequence includes text block vectors, image feature vectors, system-level prompts, and question data to be answered;

[0220] The image feature vector and the text block vector are used as context information, and the system-level prompts and the question data to be answered are used as the question-answering input task. They are input into the multimodal question-answering model to generate the corresponding answer text for the question data to be answered.

[0221] A data mixing strategy is adopted to fine-tune the multimodal question-answering model based on the corresponding answer text and multimodal input data.

[0222] In one embodiment, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the following steps:

[0223] Receiving multimodal input data, the multimodal input data including image data and corresponding question data to be answered;

[0224] Encoding the document data in the preset knowledge base and the multimodal input data to obtain a feature vector of the document data in the preset knowledge base and an image feature vector of the multimodal input data;

[0225] According to the feature vector of the document data in the preset knowledge base, at least one text search is performed in the preset knowledge base for the image feature vector to obtain a text search result;

[0226] The text retrieval results are integrated into an input sequence, which is then fed into a multimodal question-answering model to generate the corresponding answer text for the question data to be answered.

[0227] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0228] The feature vector of the document data includes the feature vector of the document title;

[0229] According to the feature vector of the document title in the preset knowledge base, a preliminary text search is performed in the preset knowledge base for the image feature vector to obtain k documents related to the image feature vector in the preset knowledge base;

[0230] Segment k documents to obtain multiple text blocks;

[0231] According to the multiple text blocks, a secondary text search is performed in a preset knowledge base for the question data to be answered, and a secondary text search result is obtained.

[0232] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0233] Using a text encoder, encoding the multiple text blocks and the question data to be answered, and generating multiple text block vectors and question vectors of preset dimensions respectively;

[0234] According to the multiple text block vectors, a secondary text search is performed in a preset knowledge base for the question vector to be answered, and a secondary text search result is obtained.

[0235] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0236] Calculate the similarity between each text block vector and the question vector to be answered to obtain a second similarity result;

[0237] According to the second similarity result, a secondary text search is performed in a preset knowledge base for the question vector to be answered to obtain a secondary text search result, which includes n text blocks related to the question data to be answered.

[0238] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0239] Calculate the inner product between the image feature vector and the document title feature vector as the first similarity result;

[0240] According to the first similarity result, an approximate k-nearest neighbor search algorithm is used to perform a preliminary text search in a preset knowledge base for the image feature vector, and k documents related to the image feature vector in the preset knowledge base are obtained.

[0241] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0242] The input sequence includes text block vectors, image feature vectors, system-level prompts, and question data to be answered;

[0243] The image feature vector and the text block vector are used as context information, and the system-level prompts and the question data to be answered are used as the question-answering input task. They are input into the multimodal question-answering model to generate the corresponding answer text for the question data to be answered.

[0244] A data mixing strategy is adopted to fine-tune the multimodal question-answering model based on the corresponding answer text and multimodal input data.

[0245] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0246] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0247] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0248] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A hierarchical retrieval method based on multimodal question answering, characterized in that: The method comprises: Receiving multimodal input data, the multimodal input data comprising image data and corresponding question data to be answered; Encoding the document data in the preset knowledge base and the multimodal input data to obtain a feature vector of the document data in the preset knowledge base and an image feature vector of the multimodal input data; According to the feature vector of the document data in the preset knowledge base, at least one text search is performed in the preset knowledge base for the image feature vector to obtain a text search result; The text retrieval results are integrated into an input sequence, and the input sequence is input into a multimodal question-answering model to generate a corresponding answer text for the question data to be answered.

2. The method according to claim 1, characterized in that The feature vector of the document data includes a feature vector of the document title; performing at least one text search in the preset knowledge base for the image feature vector based on the feature vector of the document data in the preset knowledge base to obtain a text search result includes: According to the feature vector of the document title in the preset knowledge base, a preliminary text search is performed in the preset knowledge base for the image feature vector to obtain k documents in the preset knowledge base that are related to the image feature vector; Segmenting the k documents to obtain multiple text blocks; According to the multiple text blocks, a secondary text search is performed in a preset knowledge base for the question data to be answered to obtain a secondary text search result.

3. The method according to claim 2, characterized in that The method of performing a secondary text search in a preset knowledge base for the question data to be answered based on the multiple text blocks to obtain secondary text search results includes: Using a text encoder, encoding the multiple text blocks and the question data to be answered, and generating multiple text block vectors and question vectors of preset dimensions respectively; According to the multiple text block vectors, a secondary text search is performed in a preset knowledge base for the question vector to be answered to obtain a secondary text search result.

4. The method according to claim 3, characterized in that The step of performing a secondary text search in a preset knowledge base for the question vector to be answered based on the multiple text block vectors to obtain a secondary text search result includes: Calculate the similarity between each text block vector and the question vector to be answered to obtain a second similarity result; According to the second similarity result, a secondary text search is performed in a preset knowledge base for the question vector to be answered to obtain a secondary text search result, wherein the secondary text search result includes n text blocks related to the question data to be answered.

5. The method according to claim 2, characterized in that: The method of performing a preliminary text search in the preset knowledge base for the image feature vector according to the feature vector of the document title in the preset knowledge base to obtain k documents in the preset knowledge base related to the image feature vector includes: Calculating an inner product between the image feature vector and the feature vector of the document title as a first similarity result; According to the first similarity result, an approximate k-nearest neighbor search algorithm is used to perform a preliminary text search in a preset knowledge base for the image feature vector to obtain k documents related to the image feature vector in the preset knowledge base.

6. The method according to claim 3, characterized in that The input sequence includes a text block vector, an image feature vector, a system-level prompt, and question data to be answered; the inputting the input sequence into the multimodal question-answering model to generate a corresponding answer text for the question data to be answered includes: The image feature vector and the text block vector are used as context information, and the system-level prompt and the question data to be answered are used as question-answering input tasks, and are input into a multimodal question-answering model to generate a corresponding answer text for the question data to be answered; A data mixing strategy is adopted to fine-tune the multimodal question-answering model based on the corresponding answer text and the multimodal input data.

7. A hierarchical retrieval device based on multimodal question and answer, characterized in that: The device comprises: A data receiving module, used for receiving multimodal input data, wherein the multimodal input data includes image data and corresponding question data to be answered; An encoding processing module, used to perform encoding processing on the document data in the preset knowledge base and the multimodal input data to obtain a feature vector of the document data in the preset knowledge base and an image feature vector of the multimodal input data; A text retrieval module, used to perform at least one text search in a preset knowledge base for the image feature vector according to the feature vector of the document data in the preset knowledge base, to obtain a text retrieval result; The answer output module is used to integrate the text retrieval results into an input sequence, input the input sequence into the multimodal question-answering model, and generate the corresponding answer text for the question data to be answered.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Text retrieval method and device, equipment, medium and product

    CN116932701A

  • Retrieval enhanced knowledge base question answering method, device and equipment based on secondary sorting and medium

    CN118445409A

  • Short text query expansion enhancement retrieval method based on knowledge base hierarchical tree structure

    CN118861088A

  • Multi-round dialogue processing method and system based on RAG and knowledge graph

    CN118885627A

  • Data set generation and model training method and device and storage medium

    CN119294467A

Cited By

  • Knowledge reasoning method and device for integrated circuit wafer process and medium

    CN120744140A

  • A knowledge reasoning method, device and medium for an integrated circuit wafer process

    CN120744140B

  • Large model industry knowledge question-answering method and system supporting multi-modal input

    CN121119171A

  • Question and answer method and device, equipment, storage medium and program product

    CN121542382A