Data retrieval method and device, electronic equipment and storage medium

By introducing a meta-database and a vector database into RAG technology, and using file and page identifiers labeled with text vectors to query page images, the problem of neglecting the correlation between text content and image content is solved, thus improving the response quality of the pre-trained language model.

CN121029948APending Publication Date: 2025-11-28CHENGDU SKSPRUCE TECH

Patent Information

Application Number
CN202511243884.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-11-28

Smart Images

  • Figure CN121029948A_ABST
    Figure CN121029948A_ABST
Patent Text Reader

Abstract

The invention discloses a data retrieval method and device, electronic equipment and a storage medium, relates to the field of artificial intelligence, and can store summary information of each document file by using a metadatabase and convert each page of each document file to obtain a page image. And storing a text vector generated by using each page of each document file by using a vector database, wherein the page image and the text vector are labeled with a file identifier corresponding to the document file and a page identifier corresponding to the page. When document searching is carried out, a text vector matched with document file query information can be retrieved in a vector database, then a corresponding page image is queried in a metadatabase by using a file identifier and a page identifier marked by the text vector, and then the page image is input into a multi-modal processing model for answer processing. The multi-modal processing model is helped to understand the association relationship between the text content and the image content through the page image, so that the answer quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, and in particular to a data retrieval method and device, electronic equipment and storage medium. BACKGROUND

[0002] To improve the accuracy of pre-training language models in answering user questions, retrieval-augmented generation (RAG) technology has emerged. Specifically, when the model needs to generate text or answer questions, it can retrieve relevant information from a large set of documents based on this technology, and use the retrieved information to guide text generation, thereby improving the quality of the answer.

[0003] In related technologies, traditional RAG technology usually separates and stores the text content and image content in the document file, ignoring the relevance between the text content and the image content, which can easily affect the quality of the model's answer. SUMMARY

[0004] The purpose of the present application is to provide a data retrieval method and device, electronic equipment and storage medium, which can first retrieve text vectors matching the document file query information, and then use the file identifier and page identifier labeled by the text vector to query the corresponding page image, and then input the page image into the model for answer processing, which helps the model to understand the association between the text content and the image content through the page image, thereby improving the quality of the answer.

[0005] To solve the above technical problems, the present application provides a data retrieval method applied to a retrieval system, the retrieval system comprising a metadata database and a vector database, the metadata database storing abstract information of each document file and page images converted from each page of each document file, and the vector database storing text vectors generated from each page of each document file, the page images and the text vectors being labeled with file identifiers and page identifiers; the method comprising:

[0006] When receiving a user question text sent by a user terminal, abstract information of each document file is obtained from the metadata database, and the user question text and the abstract information are input into a multi-modal processing model, so that the multi-modal processing model outputs document file query information or answer text according to the user question text and the abstract information;

[0007] When receiving the document file query information, the document file query information is input into the vector database, so that the vector database matches the text vector with the document file query information, and outputs the file identifier and the page identifier corresponding to the matched text vector;

[0008] acquire a page image corresponding to the file identifier and the page identifier from the metadata database, and input the page image into the multi-modal processing model, so that the multi-modal processing model outputs the document file query information or the answer text according to the user question text, the summary information and the page image;

[0009] When the answer text is received, the answer text is sent to the user terminal.

[0010] Optionally, the document file query information is received, including:

[0011] detecting that the multi-modal processing model initiates a call based on a model context protocol, and receiving the document file query information based on the call.

[0012] Optionally, the vector database also outputs a matching degree of the matched text vector and the document file query information.

[0013] The vector database also outputs a matching degree of the matched text vector and the document file query information.

[0014] According to the matching degree, the file identifier and the page identifier are sorted.

[0015] According to the first group of file identifiers and page identifiers, the corresponding page images are acquired from the metadata database and input into the multi-modal processing model.

[0016] When it is detected that the multi-modal processing model initiates a call again based on the model context protocol, according to the next group of file identifiers and page identifiers, the corresponding page images are acquired from the metadata database and input into the multi-modal processing model.

[0017] Optionally, it further includes:

[0018] receiving the document file input by the user terminal, and setting the file identifier for the document file and the page identifier for each page of the document file;

[0019] converting each page of the document file into a page image, and marking the file identifier and the page identifier for the page image;

[0020] extracting the page text in each page of the document file, converting the page text into the text vector, marking the file identifier and the page identifier for the text vector, and saving the marked text vector to the vector database;

[0021] A summary information is generated for the document file, and the summary information, the file identifier, the page identifier, and the page image are set as the metadata of the document file, and the metadata is saved to the metadata database.

[0022] Optionally, generating summary information for the document file includes:

[0023] The document file and the preset type theme generated text are input into the multimodal processing model, so that the multimodal processing model generates text based on the preset type theme and the document file, and outputs the document theme and document type;

[0024] Extract a preset number of characters of text from the document file, and input the text and the preset summary into the multimodal processing model so that the multimodal processing model can generate a document summary based on the text and the preset summary.

[0025] The document topic, document type, and document summary are used to form the summary information of the document file.

[0026] Optionally, it also includes:

[0027] Obtain the tenant information of the user terminal, and create a tenant storage area corresponding to the tenant in the metadata database and the vector database based on the tenant information;

[0028] Saving the marked text vector to the vector database includes:

[0029] The tenant corresponding to the document file is determined, and the text vector is saved to the tenant storage area corresponding to the tenant in the vector database;

[0030] Saving the metadata to the metadata database includes:

[0031] The tenant corresponding to the document file is determined, and the metadata is saved to the tenant storage area corresponding to the tenant in the metadata database.

[0032] Optionally, obtaining the summary information of each document file from the metadata database includes:

[0033] Based on the tenant information of the user terminal, obtain the summary information of each document file from the corresponding tenant storage space in the metadata database;

[0034] The step of inputting the document file query information into a vector database, so that the vector database uses the document file query information to match text vectors, and outputs the file identifier and page identifier corresponding to the matched text vectors, includes:

[0035] The document file query information and tenant information are sent to the vector database, so that the vector database uses the document file query information to match text vectors in the corresponding tenant storage space, and outputs the file identifier and page identifier corresponding to the matched text vectors.

[0036] The present invention also provides a data retrieval device applied to a retrieval system, the retrieval system comprising a metadata database and a vector database, the metadata database storing summary information of each document file and page images obtained by converting each page of each document file, the vector database storing text vectors generated using each page of each document file, wherein both the page images and the text vectors are labeled with document identifiers and page identifiers; the device comprises:

[0037] The summary acquisition module is used to obtain summary information of each document file from the metadata database when it receives user question text sent by the user terminal, and input the user question text and the summary information into the multimodal processing model so that the multimodal processing model outputs document file query information or answer text based on the user question text and the summary information.

[0038] The vector query module is used to input the document file query information into the vector database when the document file query information is received, so that the vector database can use the document file query information to match text vectors and output the file identifier and page identifier corresponding to the matched text vectors;

[0039] The page image acquisition module is used to acquire the page image corresponding to the file identifier and the page identifier from the metadata database, and input the page image into the multimodal processing model so that the multimodal processing model outputs the document file query information or the answer text based on the user question text, the summary information, and the page image;

[0040] The response output module is used to send the response text to the user terminal when the response text is received.

[0041] The present invention also provides an electronic device, comprising:

[0042] Memory, used to store computer programs;

[0043] A processor is used to implement the data retrieval method described above when executing the computer program.

[0044] The present invention also provides a computer-readable storage medium storing computer-executable instructions, which, when loaded and executed by a processor, implement the above-described data retrieval method.

[0045] This invention provides a data retrieval method applied to a retrieval system. The retrieval system includes a metadata database and a vector database. The metadata database stores summary information of each document file and page images generated from each page of each document file. The vector database stores text vectors generated from each page of each document file. Both the page images and the text vectors are labeled with document identifiers and page identifiers. The method includes: when receiving user question text sent by a user terminal, retrieving summary information of each document file from the metadata database, and inputting the user question text and the summary information into a multimodal processing model, so that the multimodal processing model can perform data retrieval based on the user question text and the summary information. The system outputs document file query information or answer text; when the document file query information is received, it is input into a vector database so that the vector database uses the document file query information to match text vectors and outputs the file identifier and page identifier corresponding to the matched text vectors; it obtains the page image corresponding to the file identifier and page identifier from the metadata database and inputs the page image into the multimodal processing model so that the multimodal processing model outputs the document file query information or the answer text based on the user question text, the summary information, and the page image; when the answer text is received, it is sent to the user terminal.

[0046] The beneficial effects of this invention are as follows: This invention can introduce a metadata database and a vector database. The metadata database can store summary information of each document file and page images generated from each page of each document file, while the vector database can store text vectors generated from each page of each document file. Both page images and text vectors are labeled with document identifiers and page identifiers. It is worth noting that the page text simultaneously contains text information and image information within that page. Returning the page text to the model helps the model understand the relationship between text content and image content. When receiving user question text input from the user terminal, this invention can obtain summary information of each document file from the metadata database and input the user question text and the summary information into the multimodal processing model. This allows the multimodal processing model to output document file query information or answer text based on the user question text and the summary information, thus triggering the multimodal processing model to perform document file queries or answer user questions based on the user question text and summary information. Subsequently, if answer text is received, the answer text is sent to the user terminal. If a document query is received, this invention can input the document query information into a vector database, allowing the vector database to match text vectors using the document query information and output the file identifier and page identifier corresponding to the matched text vectors. Subsequently, it can retrieve the page image corresponding to the file identifier and page identifier from a metadata database and input the page image into a multimodal processing model. This allows the multimodal processing model to output the document query information or answer text based on the user's question text, summary information, and page image. In other words, this invention does not simply input the retrieved document text into the multimodal processing model, nor does it input separate document text and image content. Instead, it can input a page image containing both text and image content into the multimodal processing model. This helps the multimodal processing model understand the relationship between the text content and the image content, thereby improving the quality of the multimodal processing model's answer to the user's question.

[0047] The present invention also provides a data retrieval device, an electronic device, and a storage medium, which have the above-mentioned beneficial effects. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0049] Figure 1 A schematic diagram of a retrieval system provided in an embodiment of the present invention;

[0050] Figure 2A flowchart of a data retrieval method provided in an embodiment of the present invention;

[0051] Figure 3 A schematic diagram of another retrieval system provided in an embodiment of the present invention;

[0052] Figure 4 This is a structural block diagram of a data retrieval device provided in an embodiment of the present invention;

[0053] Figure 5 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0055] To improve the accuracy of pre-trained language models in answering user questions, Retrieval-augmented Generation (RAG) technology has emerged. Specifically, when a model needs to generate text or answer questions, this technology can retrieve relevant information from a large document collection and use the retrieved information to guide text generation, thereby improving the quality of the answer. However, traditional RAG technology often stores text and image content separately in document files, neglecting the correlation between the two, which can easily affect the quality of the model's answer.

[0056] In view of this, to address the technical problem of how to improve retrieval results and facilitate the model's effective understanding of the relationship between text content and image content, this invention provides a data retrieval method. This method first retrieves text vectors that match document file query information, then uses the file identifier and page identifier marked by the text vectors to query the corresponding page image, and finally inputs the page image into the model for response processing. This helps the model understand the relationship between text content and image content through the page image, thereby improving the quality of the response.

[0057] First, the retrieval system provided in this embodiment will be introduced. Please refer to... Figure 1 , Figure 1This is a schematic diagram of a retrieval system provided in an embodiment of the present invention. The retrieval system may include a metadata database 1, a vector database 2, and a retrieval module 3. The retrieval module 3 establishes communication connections with both the metadata database 1 and the vector database 2. The metadata database 1 stores summary information 11 for each document file and page images 12 generated from each page of each document file. Each page image 12 has annotation information 13, which includes the file identifier (file-id) of the document file to which the page image 12 belongs and the page identifier (page-id) of the page to which it belongs. The vector database 2 stores text vectors 21 generated from each page of each document file. Each text vector 21 has annotation information 22, which includes the file identifier of the document file to which the page image 21 belongs and the page identifier of the page to which it belongs. That is, the page image 12 and the text vector 21 can be associated through the file identifier and the page identifier. The retrieval module 3 is used to execute the data retrieval method provided in this embodiment.

[0058] In simple terms, this embodiment can set a file identifier for a document file and a page identifier for each page in the document file. Then, the text content of each page is converted into a text vector 21, the text vector 21 is labeled with the file identifier and page identifier, and saved in the vector database 2. Additionally, summary information 11 can be generated for the document file, and each page of the document file is converted into a page image 12, the page image 12 is labeled with the file identifier and page identifier, ensuring that the page image 12 simultaneously contains both the text and images on the page, and then the summary information 11 and page image 12 are saved in the metadata database 1.

[0059] Furthermore, to achieve a correlated understanding of text and image content, retrieval module 3 can also be connected to a multimodal processing model (the multimodal processing model is not included in...). Figure 1 (Drawn in the middle). The multimodal processing model is a pre-trained artificial intelligence model that can parse and process text and images, and can initiate text retrieval to the retrieval module 3. Unlike related technologies, when performing a retrieval according to the requirements of the multimodal processing model, the retrieval module 3 does not return the text from the document file to the multimodal processing model. Instead, it retrieves the text vector 21 from the vector database 2, and obtains the corresponding page image 12 from the metadata database 1 based on the annotation information 22 of the retrieved text vector 21, and then sends the page image 12 to the multimodal processing model. As mentioned above, since the page image 12 contains both text and images on the page, the multimodal processing model can effectively understand the text and images on the same page, avoid ignoring image information, and thus effectively improve the response quality of the multimodal processing model.

[0060] It should be noted that this embodiment does not limit the type of metadata database 1, as long as it can store summary information 11 and page images 12. It can be set according to actual application needs, for example, it can be a Postgres database. This embodiment does not limit the type of vector database 2, as long as it can store text vectors 21. It can be set according to actual application needs, for example, it can be a Qdrant database. This embodiment also does not limit the type of multimodal processing model, as long as it can simultaneously possess text parsing and image parsing capabilities, and has the ability to call the retrieval module 3.

[0061] It should also be noted that the retrieval module 3 may contain other sub-modules to provide different functions, as described in the following embodiments.

[0062] Based on the above description of the retrieval system, the data retrieval method provided in this embodiment will be described below. Please refer to... Figure 2 , Figure 2 A flowchart of a data retrieval method provided in an embodiment of the present invention, which is applied to the above-mentioned retrieval system, may include:

[0063] S101. When the user question text sent by the user terminal is received, the summary information of each document file is obtained from the metadata database, and the user question text and summary information are input into the multimodal processing model so that the multimodal processing model can query information or answer text from the document file based on the user question text and summary information.

[0064] In this step, the retrieval module of the retrieval system can receive user question text sent by the user terminal and pass it to the multimodal processing model. To facilitate the multimodal processing model's understanding of the basic information of the documents in the retrieval system and to initiate a search based on this information and user needs, the retrieval module can also obtain summary information of each document from the metadata database and input both the user question text and the summary information into the multimodal processing model. In this way, the multimodal processing model can parse the user question text and the summary information and decide whether to output document query information or answer text. The document query information is used to initiate a search with the retrieval module; this information may include keywords and other information required for the search. The answer text is used to answer the user question.

[0065] It should be noted that this embodiment does not limit the specific content of the summary information, which can be set according to actual application needs. For example, it can include a document summary (Overview) of the document file, or it can include the document topic (Topic) and document type (Types, such as official documents, bills, logs, etc.).

[0066] S102. When document file query information is received, the document file query information is input into the vector database so that the vector database can use the document file query information to match text vectors and output the file identifier and page identifier corresponding to the matched text vectors.

[0067] In this step, upon receiving document query information, the retrieval module of the retrieval system can input the document query information into a vector database, enabling the vector database to perform text vector matching using the document query information. For example, the vector database can vectorize (embedding) the keywords in the document query information and perform similarity matching between the keyword vectors and various text vectors. Finally, it determines the matched text vectors based on the similarity, such as selecting the text vector with the highest similarity as the matched text vector. It should be noted that this embodiment does not limit the specific steps of vector retrieval; relevant vector retrieval technologies can be referenced.

[0068] Subsequently, unlike related technologies, the vector database does not return the text vector or the text used to generate the text vector to the retrieval module. Instead, it can query the file identifier and page identifier corresponding to the matched text vector and return the file identifier and page identifier to the retrieval module to trigger the retrieval module to perform page image retrieval.

[0069] It should be noted that this embodiment does not limit how the multimodal processing model sends document query information, or how the retrieval module perceives that the information sent by the multimodal processing model is document query information. For example, the multimodal processing model can initiate a call to the retrieval module based on the Model Context Protocol (MCP), and when the retrieval module detects this call, it can receive document query information through this call.

[0070] In one implementation, receiving document file query information may include:

[0071] Step 11: Detect calls initiated by the multimodal processing model based on the model context protocol, and receive document file query information based on the calls.

[0072] Understandably, the retrieval module can execute step S102 each time a call initiated by the multimodal processing model is detected. However, considering the low efficiency of repeated retrieval, in another implementation, the vector database can retrieve multiple text vectors that match the document file query information, for example, the top-k text vectors that match the document file query information, and return the matching degree, file identifier, and page identifier of each text vector to the retrieval module. At this time, the retrieval module can sort the file identifier and page identifier pairs according to the matching degree to obtain the ranking result (Rank), and perform page image retrieval based on the first set of file identifiers and page identifiers. In subsequent interactions, if the multimodal processing model initiates the same call to the retrieval module, the retrieval module can perform page image retrieval based on the next set of file identifiers and page identifiers without initiating a text vector retrieval to the vector database again, thereby improving retrieval efficiency.

[0073] Based on this, the vector database also outputs the matching degree between the matched text vectors and the document file query information; it retrieves the page image corresponding to the file identifier and page identifier from the metadata database, and inputs the page image into the multimodal processing model, which may include:

[0074] Step 21: Sort the file identifier and page identifier pairs according to their matching degree.

[0075] Step 22: Obtain the corresponding page image from the metadata database based on the first set of file identifiers and page identifiers, and input the page image into the multimodal processing model.

[0076] Step 23: When the multimodal processing model is detected to initiate another call based on the model context protocol, the corresponding page image is obtained from the metadata database according to the next set of file identifiers and page identifiers, and the page image is input into the multimodal processing model.

[0077] S103. Obtain the page image corresponding to the file identifier and page identifier from the metadata database, and input the page image into the multimodal processing model so that the multimodal processing model can output document file query information or answer text based on the user question text, summary information, and page image.

[0078] In this step, the retrieval module retrieves page images corresponding to file and page identifiers from the metadata database and inputs these images into the multimodal processing model. As mentioned above, the page image contains both text and image content from the corresponding page; therefore, the multimodal processing model can understand the relationship between text and image content based on the page image. Thus, by parsing the user's question text, summary information, and page images, the multimodal processing model can further improve the retrieval quality of document file queries and also enhance the quality of the response text.

[0079] It should be noted that, since the multimodal processing model decides to initiate the call or output the response text in this embodiment, the multimodal processing model can initiate multiple calls, and the retrieval module can execute steps S102 and S103 multiple times until the multimodal processing model outputs the response text.

[0080] S104. When the reply text is received, the reply text is sent to the user terminal.

[0081] In this step, the retrieval module receives the answer text output by the multimodal processing model and outputs the question-and-answer text to the user, thus completing a round of question-and-answer initiated by the user.

[0082] Based on the above embodiments, this invention can introduce a metadata database and a vector database. The metadata database can store summary information of each document file and page images generated from each page of each document file, while the vector database can store text vectors generated from each page of each document file. Both the page images and text vectors are labeled with document identifiers and page identifiers. It is worth noting that the page text contains both text information and image information within that page. Returning the page text to the model helps the model understand the relationship between text content and image content. When receiving user question text input from the user terminal, this invention can obtain summary information of each document file from the metadata database and input the user question text and the summary information into the multimodal processing model. This allows the multimodal processing model to output document file query information or answer text based on the user question text and the summary information, thus triggering the multimodal processing model to perform document file queries or answer user questions based on the user question text and summary information. Subsequently, if answer text is received, the answer text is sent to the user terminal. If a document query is received, this invention can input the document query information into a vector database, allowing the vector database to match text vectors using the document query information and output the file identifier and page identifier corresponding to the matched text vectors. Subsequently, it can retrieve the page image corresponding to the file identifier and page identifier from a metadata database and input the page image into a multimodal processing model. This allows the multimodal processing model to output the document query information or answer text based on the user's question text, summary information, and page image. In other words, this invention does not simply input the retrieved document text into the multimodal processing model, nor does it input separate document text and image content. Instead, it can input a page image containing both text and image content into the multimodal processing model. This helps the multimodal processing model understand the relationship between the text content and the image content, thereby improving the quality of the multimodal processing model's answer to the user's question.

[0083] The following describes the construction method of the retrieval system. In one implementation, this method may further include:

[0084] S201. Receive the document file input by the user terminal, set the file identifier for the document file, and set the page identifier for each page of the document file.

[0085] In this step, when the retrieval module receives a document file to be stored input by the user, it can set a file identifier for the document file and a page identifier for each page of the document file. This embodiment does not limit the setting method of file identifier and page identifier, as long as it can distinguish different document files and different pages within the same document file.

[0086] It should be noted that this embodiment does not limit the file format of the document file, and it can be any format such as txt, csv, ppt, pdf, docx, excel, etc.

[0087] S202. Convert each page of the document file into a page image, and mark the page image with file identifier and page identifier.

[0088] In this step, the retrieval module converts each page of the document file into a page image and labels the page image with a file identifier and a page identifier. Performing this step ensures that the page image simultaneously contains both text and image content from the page, facilitating parsing by the multimodal processing model.

[0089] It should be noted that this embodiment does not limit how each page of the document file is converted into a page image. It can be set according to the actual application requirements, such as taking screenshots.

[0090] S203. Extract the page text from each page of the document file, convert the page text into a text vector, label the text vector with file identifier and page identifier, and save the labeled text vector to the vector database.

[0091] In this step, the retrieval module extracts the page text from each page of the document file and performs vector embedding on the page text, thereby converting the page text into text vectors. Subsequently, the retrieval module labels the text vectors with file identifiers and page identifiers, and saves the labeled text vectors to the vector database.

[0092] It should be noted that this embodiment does not limit how to convert page text into text vectors; relevant vector conversion technologies can be referenced.

[0093] Furthermore, for situations where there is too much text on a single page, this embodiment can either split the text on the single page into multiple parts for vector conversion and label them with the same file identifier and page identifier; or it can extract the first preset number of characters from the page as the page text and perform vector conversion. This embodiment does not limit the specific value of the preset number and can be set according to actual application needs.

[0094] S204. Generate summary information for the document file, set the summary information, file identifier, page identifier, and page image as metadata of the document file, and save the metadata to the metadata database.

[0095] In this step, the retrieval module can generate summary information for the document file and set the summary information, file identifier, page identifier, and page image as metadata for the document file, saving them to the metadata database. As mentioned above, the summary information can be the document topic, document type, and document summary. To improve the efficiency of summary information generation, this embodiment can utilize a multimodal processing model for summary generation.

[0096] The process of abstract generation is described below. In one implementation, generating abstract information for a document file may include:

[0097] Step 31: Input the document file and the preset type topic into the multimodal processing model so that the multimodal processing model can generate text and document files according to the preset type topic and output the document topic and document type.

[0098] The preset topic-generated text indicates that the input document file needs to be analyzed to determine its document topic and document type. This text can contain the specified document topic and document type for the multimodal processing model to select; additionally, it can include the output format for the document topic and document type for the model's reference. In this step, the retrieval module inputs the document file and the preset topic-generated text into the multimodal processing model. The model then analyzes the document file based on the preset topic-generated text to obtain its document topic and document type.

[0099] Step 32: Extract the text of a preset word count from the document file, and input the text and the preset summary into the multimodal processing model so that the multimodal processing model can generate a document summary based on the text and the preset summary.

[0100] The preset summary generation text indicates that the input text needs to be analyzed to generate a text summary. In this step, the retrieval module can extract a preset number of words (e.g., the first 1000 words) from the document file and input this text along with the preset summary generation text into the multimodal processing model. The multimodal processing model then generates the text based on the preset summary, analyzes the input text, and outputs a document summary.

[0101] Step 33: Combine the document topic, document type, and document summary to form the summary information of the document file.

[0102] In this step, the retrieval module can generate summary information for a document file from the document topic, document type, and document summary. Then, the summary information, file identifier, page identifier, and page image can be set as metadata for the document file and saved to the metadata database.

[0103] Furthermore, to facilitate multi-user use of the retrieval system, this embodiment can also configure multi-tenancy functionality for the retrieval system. Specifically, the retrieval module can obtain tenant information from the user's end (such as tenant ID, tenant API token, etc.) and create tenant storage areas corresponding to the tenant in the metadata database and vector database based on the tenant information. For example, it can create data tables (Tables) corresponding to the tenant information in the metadata database and vector database. Consequently, the user's end can only access the corresponding tenant storage space in the metadata database and vector database by providing tenant information, enabling multi-tenant access to the retrieval system under tenant isolation conditions.

[0104] In one embodiment, the method may further include:

[0105] Step 41: Obtain the tenant information from the user's end, and create a tenant storage area corresponding to the tenant in the metadata database and vector database based on the tenant information;

[0106] Save the marked text vectors to a vector database, including:

[0107] Step 51: Determine the tenant corresponding to the document file and save the text vector to the tenant storage area corresponding to the tenant in the vector database;

[0108] Save metadata to the metadata database, including:

[0109] Step 61: Determine the tenant corresponding to the document file and save the metadata to the tenant storage area corresponding to the tenant in the metadata database.

[0110] The following describes a data retrieval method for multi-tenant scenarios. In one implementation, this method may further include:

[0111] S301. When receiving user question text sent by the user terminal, according to the tenant information of the user terminal, obtain the summary information of each document file from the corresponding tenant storage space in the metadata database, and input the user question text and summary information into the multimodal processing model so that the multimodal processing model can output document file query information or answer text based on the user question text and summary information.

[0112] In this embodiment, before obtaining the summary information of the document file, the retrieval module needs to obtain the tenant information of the user terminal, and only obtain the summary information in the tenant storage space corresponding to the tenant information in the metadata database.

[0113] S302. When receiving document file query information, send the document file query information and tenant information to the vector database so that the vector database can use the document file query information to match text vectors in the corresponding tenant storage space and output the file identifier and page identifier corresponding to the matched text vectors.

[0114] In this embodiment, when the retrieval module receives document file query information, it needs to send the document file query information and tenant information to the vector database together, so that the vector database can use the document file query information to match the text vector in the corresponding tenant storage space.

[0115] S303. Obtain the page image corresponding to the file identifier and page identifier from the metadata database, and input the page image into the multimodal processing model so that the multimodal processing model outputs document file query information or answer text based on the user question text, summary information, and page image.

[0116] In this embodiment, the retrieval module retrieves the page image only from the tenant storage space corresponding to the tenant information in the metadata database.

[0117] S304. When the response text is received, the response text is sent to the user terminal.

[0118] Based on the above embodiments, the retrieval system and data retrieval method described below are based on the complete accompanying drawings. Please refer to... Figure 3 , Figure 3 This is a schematic diagram of another retrieval system provided in an embodiment of the present invention. This system may include a metadata database 1, a vector database 2, an interface service module 31, a coordination module 32, a vector query module 33, a vector transformation module 34, an extractor module 35, and a tenant management module 36. The coordination module 32, vector transformation module 34, and extractor module 35 can interact with the multimodal processing model 4. The interaction process between the above modules will be described below from the file storage stage and the file retrieval stage respectively.

[0119] File storage stage:

[0120] 1. The tenant management module 36 obtains tenant information (such as tenant identifier and tenant API token) and creates the corresponding data table for the tenant in the metadata database 1 and the vector database 2 based on the tenant information.

[0121] 2. The tenant management module 36 receives the input source file and passes it to the extractor module 35.

[0122] 3. The extractor module 35 splits and converts the source file to obtain document files, page text and page images, marks the page text and page images with file identifiers and page identifiers, and saves the document files, page text and page images to the output directory.

[0123] 4. The extractor module 35 sends the source file to the multimodal processing model 4 to generate document type and document topic, forms metadata by including file name, file storage location, file type and file topic, and saves the metadata and page images to the metadata database 1.

[0124] 5. The vector conversion module 34 obtains the document file and page text, sends a portion of the text in the document file to the multimodal processing model 4 to obtain a document summary, and saves the document summary as metadata to the metadata database; and converts the page text into text vectors, annotates the text vectors with text identifiers and page identifiers, and saves them to the vector database 2.

[0125] Document retrieval stage:

[0126] 1. The interface service module 31 verifies the tenant information recorded in the metadata database 1, and after successful verification, receives the user problem text sent by the user and transmits the user problem text to the coordination module 32.

[0127] 2. The coordination module 32 obtains the document summary of each document file from the metadata database and sends the document summary and user question text to the multimodal processing model 4;

[0128] 3. When the coordination module 32 receives the function call from the multimodal processing model 4, it sends the document file query information sent by the multimodal processing model 4 to the vector query module 33.

[0129] 4. The vector query module 33 inputs the document file query information into the vector database 2, so that the vector database 2 uses the document file query information to match text vectors and outputs the file identifier, page identifier and matching degree (Top-k information) corresponding to the matched text vectors.

[0130] 5. The vector query module 33 sends the file identifier, page identifier, and matching degree sorting level to the coordination module 32.

[0131] 6. The coordination module 32 sorts the matching degree levels, uses the file identifier and page identifier to obtain the page image from the metadata database 1 in sequence, and sends the page image to the multimodal processing model 4.

[0132] 7. When the coordination module 32 receives the response text from the multimodal processing model 4, it transmits the response text to the interface service module 31, which then transmits the response text to the user terminal.

[0133] It is worth noting that the present invention has the following characteristics:

[0134] 1. Supports multimodal processing: Multimodal processing models can perform multimodal processing on text and images;

[0135] 2. Supports multi-tenancy: Multi-tenant isolation can be implemented in metadata databases and vector databases;

[0136] 3. Multilingual support: The text vector generation and multimodal processing models support various languages ​​for text parsing and processing;

[0137] 4. Supports multiple file formats: It can process various file types;

[0138] 5. Supports multiple model selections: It can interface with various multimodal processing models;

[0139] 6. Supports model context invocation.

[0140] The data retrieval device, electronic device, computer program product, and computer-readable storage medium provided in the embodiments of the present invention will be described below. The data retrieval device, electronic device, computer program product, and computer-readable storage medium described below can be referred to in correspondence with the data retrieval method described above.

[0141] Please refer to Figure 4 , Figure 4 This is a structural block diagram of a data retrieval device provided in an embodiment of the present invention. The device can be applied to a retrieval system, which includes a metadata database and a vector database. The metadata database stores summary information of each document file and page images generated from each page of each document file. The vector database stores text vectors generated from each page of each document file. Both the page images and text vectors are labeled with document identifiers and page identifiers. The device may include:

[0142] The summary acquisition module 401 is used to obtain summary information of each document file from the metadata database when it receives the user question text sent by the user terminal, and input the user question text and summary information into the multimodal processing model so that the multimodal processing model can query information or answer text from the document file output by the user question text and summary information.

[0143] The vector query module 402 is used to input the document file query information into the vector database when it receives the document file query information, so that the vector database can use the document file query information to match text vectors and output the file identifier and page identifier corresponding to the matched text vectors;

[0144] The page image acquisition module 403 is used to obtain the page image corresponding to the file identifier and page identifier from the metadata database, and input the page image into the multimodal processing model so that the multimodal processing model can output document file query information or answer text based on the user question text, summary information and page image;

[0145] The response output module 404 is used to send the response text to the user terminal when a response text is received.

[0146] Optionally, the summary acquisition module 401 includes:

[0147] The call detection module is used to detect calls initiated by the multimodal processing model based on the model context protocol, and to receive document file query information based on the calls.

[0148] Optionally, the vector database also outputs the matching degree between the matched text vectors and the document file query information;

[0149] Page image acquisition module 403 includes:

[0150] The sorting submodule is used to sort file identifiers and page identifiers based on their matching degree.

[0151] The page image acquisition submodule is used to retrieve the corresponding page image from the metadata database based on the first set of file identifiers and page identifiers, and input the page image into the multimodal processing model; when it is detected that the multimodal processing model initiates another call based on the model context protocol, it retrieves the corresponding page image from the metadata database based on the next set of file identifiers and page identifiers, and inputs the page image into the multimodal processing model.

[0152] Optionally, the device may further include:

[0153] The identifier setting module is used to receive document files input by the user and set file identifiers and page identifiers for each page of the document file.

[0154] The page image conversion module is used to convert each page of a document file into a page image and to mark the page image with file identifiers and page identifiers;

[0155] The text vector conversion module is used to extract the page text from each page of the document file, convert the page text into text vectors, mark the text vectors with file identifiers and page identifiers, and save the marked text vectors to the vector database;

[0156] The metadata generation module is used to generate summary information for document files, set the summary information, file identifier, page identifier, and page image as the metadata of the document file, and save the metadata to the metadata database.

[0157] Optionally, the metadata generation module may include:

[0158] The topic type generation submodule is used to input document files and preset topic type text into the multimodal processing model, so that the multimodal processing model can generate text and document files according to the preset topic type and output document topic and document type;

[0159] The summary generation submodule is used to extract text of a preset word count from the document file, input the text and the preset summary generation text into the multimodal processing model, so that the multimodal processing model can generate text based on the text and the preset summary and output a document summary.

[0160] The summary merging submodule is used to combine document topic, document type, and document summary into a summary information for a document file.

[0161] Optionally, the device may further include:

[0162] The tenant management module is used to obtain tenant information from the user's end and create tenant storage areas corresponding to the tenant in the metadata database and vector database based on the tenant information.

[0163] The text vector conversion module may include:

[0164] The vector database storage submodule is used to determine the tenant corresponding to the document file and save the text vector to the tenant storage area corresponding to the tenant in the vector database;

[0165] Metadata generation module may include:

[0166] The metadata storage submodule is used to determine the tenant corresponding to the document file and save the metadata to the tenant storage area corresponding to the tenant in the metadata database.

[0167] Optionally, the summary acquisition module 401 can be used for:

[0168] Based on the tenant information on the user's end, retrieve the summary information of each document file from the corresponding tenant storage space in the metadata database;

[0169] Vector query module 402 can be used for:

[0170] The document file query information and tenant information are sent to the vector database, so that the vector database can use the document file query information to match text vectors in the corresponding tenant storage space, and output the file identifier and page identifier corresponding to the matched text vectors.

[0171] Please refer to Figure 5 , Figure 5 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. The electronic device 50 provided in this embodiment includes a processor 51 and a memory 52; wherein, the memory 52 is used to store a computer program; and the processor 51 is used to execute the data retrieval method provided in the foregoing embodiment when executing the computer program.

[0172] For details regarding the specific process of the above data retrieval method, please refer to the relevant content provided in the foregoing embodiments, which will not be repeated here.

[0173] Furthermore, the memory 52, as a carrier for resource storage, can be a read-only memory, random access memory, disk, or optical disk, and the storage method can be temporary storage or permanent storage.

[0174] In addition, the electronic device 50 also includes a power supply 53, a communication interface 54, an input / output interface 55, and a communication bus 56; wherein, the power supply 53 is used to provide operating voltage for the various hardware devices on the electronic device 50; the communication interface 54 can create a data transmission channel between the electronic device 50 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this invention, and is not specifically limited here; the input / output interface 55 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0175] This invention also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the data retrieval method described in the above embodiments.

[0176] Since the embodiments of the computer program product portion correspond to the embodiments of the data retrieval method portion, please refer to the description of the embodiments of the data retrieval method portion for the embodiments of the computer program product portion, and will not be repeated here.

[0177] This invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the data retrieval method described in the above embodiments.

[0178] Since the embodiments of the computer-readable storage medium portion correspond to the embodiments of the data retrieval method portion, the embodiments of the storage medium portion are described in the description of the embodiments of the data retrieval method portion, and will not be repeated here.

[0179] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0180] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0181] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0182] The data retrieval method, apparatus, electronic device, and storage medium provided by this invention have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of this invention. It should be noted that those skilled in the art can make various improvements and modifications to this invention without departing from its principles, and these improvements and modifications also fall within the protection scope of this invention.

Claims

1. A data retrieval method, characterized in that, The method is applied to a retrieval system, which includes a metadata database and a vector database. The metadata database stores summary information of each document file and page images generated from each page of each document file. The vector database stores text vectors generated from each page of each document file. Both the page images and the text vectors are labeled with document and page identifiers. The method includes: When a user question text is received from a user terminal, the summary information of each document file is obtained from the metadata database, and the user question text and the summary information are input into the multimodal processing model so that the multimodal processing model outputs document file query information or answer text based on the user question text and the summary information. When the document file query information is received, the document file query information is input into the vector database so that the vector database uses the document file query information to match text vectors and outputs the file identifier and page identifier corresponding to the matched text vectors; Obtain the page image corresponding to the file identifier and the page identifier from the metadata database, and input the page image into the multimodal processing model so that the multimodal processing model outputs the document file query information or the answer text based on the user question text, the summary information, and the page image; When the response text is received, it is sent to the user's terminal.

2. The data retrieval method according to claim 1, characterized in that, Receive the document file query information, including: The system detects calls initiated by the multimodal processing model based on the model context protocol and receives document file query information based on these calls.

3. The data retrieval method according to claim 2, characterized in that, The vector database also outputs the matching degree between the matched text vectors and the document file query information; The step of obtaining the page image corresponding to the file identifier and the page identifier from the metadata database and inputting the page image into the multimodal processing model includes: The file identifier and page identifier pairs are sorted according to the matching degree. The corresponding page image is obtained from the metadata database based on the first set of file identifiers and page identifiers, and the page image is input into the multimodal processing model; When the multimodal processing model is detected to initiate another call based on the model context protocol, the corresponding page image is obtained from the metadata database according to the next set of file identifiers and page identifiers, and the page image is input into the multimodal processing model.

4. The data retrieval method according to any one of claims 1 to 3, characterized in that, Also includes: Receive the document file input by the user terminal, and set the file identifier for the document file and set the page identifier for each page of the document file; Each page of the document file is converted into a page image, and the page image is labeled with the file identifier and the page identifier; Extract the page text from each page of the document file, convert the page text into the text vector, label the text vector with the file identifier and the page identifier, and save the labeled text vector to the vector database; A summary information is generated for the document file, and the summary information, the file identifier, the page identifier, and the page image are set as the metadata of the document file, and the metadata is saved to the metadata database.

5. The data retrieval method according to claim 4, characterized in that, The step of generating summary information for the document file includes: The document file and the preset type theme generated text are input into the multimodal processing model, so that the multimodal processing model outputs the document theme and document type according to the preset type theme generated text and the document file; Extract a preset number of characters of text from the document file, and input the text and the preset summary into the multimodal processing model so that the multimodal processing model can generate a document summary based on the text and the preset summary. The document topic, document type, and document summary are used to form the summary information of the document file.

6. The data retrieval method according to claim 4, characterized in that, Also includes: Obtain the tenant information of the user terminal, and create a tenant storage area corresponding to the tenant in the metadata database and the vector database based on the tenant information; Saving the marked text vector to the vector database includes: The tenant corresponding to the document file is determined, and the text vector is saved to the tenant storage area corresponding to the tenant in the vector database; Saving the metadata to the metadata database includes: The tenant corresponding to the document file is determined, and the metadata is saved to the tenant storage area corresponding to the tenant in the metadata database.

7. The data retrieval method according to claim 6, characterized in that, The step of obtaining summary information of each document file from the metadata database includes: Based on the tenant information of the user terminal, obtain the summary information of each document file from the corresponding tenant storage space in the metadata database; The step of inputting the document file query information into a vector database, so that the vector database uses the document file query information to match text vectors, and outputs the file identifier and page identifier corresponding to the matched text vectors, includes: The document file query information and tenant information are sent to the vector database, so that the vector database uses the document file query information to match text vectors in the corresponding tenant storage space, and outputs the file identifier and page identifier corresponding to the matched text vectors.

8. A data retrieval device, characterized in that, An apparatus is used in a retrieval system, the retrieval system comprising a metadata database and a vector database. The metadata database stores summary information of each document file and page images generated from each page of each document file. The vector database stores text vectors generated from each page of each document file. Both the page images and the text vectors are labeled with document identifiers and page identifiers. The apparatus includes: The summary acquisition module is used to obtain summary information of each document file from the metadata database when it receives user question text sent by the user terminal, and input the user question text and the summary information into the multimodal processing model so that the multimodal processing model outputs document file query information or answer text based on the user question text and the summary information. The vector query module is used to input the document file query information into the vector database when the document file query information is received, so that the vector database can use the document file query information to match text vectors and output the file identifier and page identifier corresponding to the matched text vectors; The page image acquisition module is used to acquire the page image corresponding to the file identifier and the page identifier from the metadata database, and input the page image into the multimodal processing model so that the multimodal processing model outputs the document file query information or the answer text based on the user question text, the summary information, and the page image; The response output module is used to send the response text to the user terminal when the response text is received.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the data retrieval method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when loaded and executed by a processor, implement the data retrieval method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image-text related recommendation method and device, electronic equipment and storage medium

    CN117235370A

  • Large language model knowledge retrieval method and system based on semantic vectorization

    CN119336864A

  • Text retrieval enhancement generation method, device and equipment and readable storage medium

    CN119398168A

  • Multi-modal document retrieval enhancement generation method based on large model

    CN119988588A

  • Cross-modal data search method and device based on Surreal DB

    CN120277255A

Cited By

  • Document retrieval method and device, equipment, storage medium and computer program product

    CN121658634A