Intelligent reply method, system and device under user query information and medium

By performing structured processing of documents and decomposition and storage of multimodal information, and combining user query information for multimodal search, the problem of difficulty in parsing and utilizing multimodal information in existing systems is solved, and efficient and accurate information retrieval and question-and-answer services are achieved.

CN120104772APending Publication Date: 2025-06-06DARK MATTER ARTIFICIAL INTELLIGENT (BEIJING) TECHNOLOGY CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510159920.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing multimodal information retrieval system is difficult to achieve comprehensive analysis and efficient utilization of multimodal information in documents, and cannot support users to query in text, pictures or combinations of the two. It lacks effective integration and sorting strategies, resulting in the relevance and accuracy of the search results need to be improved.

Method used

By structuring the document, a layout knowledge block containing the document logical structure and multimodal information is constructed, and it is decomposed into text knowledge blocks and picture knowledge blocks, an embedded vector is generated and stored in a vector database, and multimodal search is performed in combination with user query information, and finally the reply result corresponding to the user query information is output through the big model.

Benefits of technology

It realizes comprehensive analysis and efficient utilization of multimodal information in the document, supports multiple query methods for users, provides accurate and efficient search results and question-and-answer services, and significantly improves the breadth and depth of information processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104772A_ABST
    Figure CN120104772A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent reply method, system and device under user query information and a medium, and the method comprises the following steps: carrying out structured processing on a document provided by a user, and constructing a layout knowledge block containing a document logic structure and multi-modal information; decomposing the layout knowledge blocks into text knowledge blocks and picture knowledge blocks, generating embedded vectors for the text knowledge blocks and the picture knowledge blocks, and storing the embedded vectors in a vector database; in combination with user query information, performing multi-modal retrieval on the vector database to obtain a multi-modal retrieval result; and based on the multi-modal retrieval result, outputting a reply result corresponding to the user query information through the large model. According to the method, by supporting cross-modal retrieval, optimizing picture abstract extraction and utilizing structured document information, document content can be comprehensively understood and efficiently utilized, and more accurate and comprehensive search results and high-quality question and answer services are provided for users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and information processing, and more specifically to an intelligent reply method, system, device and medium for user query information. Background Art

[0002] With the rapid development of big data and artificial intelligence technologies, users' needs for information acquisition are becoming increasingly diverse in the field of information retrieval and question-answering systems. Especially in document processing and knowledge retrieval, how to efficiently process multimodal information in documents (including text, images, tables, formulas, etc.) and achieve accurate information retrieval and high-quality question-answering generation has become a hot topic and difficulty in current research.

[0003] Traditional document parsing and information retrieval technologies mainly rely on single-modal information processing, such as retrieving only text information or recognizing only images, which makes it difficult to fully parse and efficiently utilize multimodal information in documents. In addition, existing multimodal information retrieval systems often lack in-depth understanding and structured representation of document structures, resulting in inaccurate retrieval results that cannot meet users' needs for high-quality question-answer generation.

[0004] In addition, when processing user queries, existing multimodal information retrieval systems can often only accept input of a single modality (such as text only or image only), and cannot effectively support mixed queries of text and images. This limits the flexibility and diversity of user queries and reduces the practicality of the system. At the same time, when processing multimodal retrieval results, existing systems lack effective integration and sorting strategies, resulting in the need to improve the relevance and accuracy of retrieval results.

[0005] Therefore, how to design an intelligent reply method for user query information that can achieve comprehensive analysis and structured representation of multimodal information in documents, support users to query in the form of text, pictures, or a combination of both, and provide accurate and efficient retrieval results and question-and-answer services is an issue that technical personnel in this field urgently need to solve. Summary of the invention

[0006] In view of this, the present invention provides an intelligent reply method under user query information, which constructs a layout knowledge block containing the document logical structure and multimodal information by structurally processing the document, and decomposes it into text knowledge blocks and image knowledge blocks for storage and retrieval. Combined with user query information, the method can realize multimodal retrieval, and finally output the reply result corresponding to the user query information through a large model, thereby providing users with high-quality intelligent question-and-answer services.

[0007] In order to achieve the above object, the present invention adopts the following technical solution:

[0008] In a first aspect, the present invention provides a method for intelligent replying to user query information, comprising the following steps:

[0009] S1. Structural processing of documents provided by users to construct layout knowledge blocks containing document logical structure and multimodal information;

[0010] S2, decomposing the layout knowledge block into a text knowledge block and a picture knowledge block, generating embedding vectors for the text knowledge block and the picture knowledge block, and storing them in a vector database;

[0011] S3. Perform a multimodal search on the vector database in combination with the user query information to obtain a multimodal search result;

[0012] S4. Based on the multimodal retrieval results, a reply result corresponding to the user query information is output through a large model.

[0013] Furthermore, the S1 includes:

[0014] S11, standardizing the format and converting the document provided by the user into an image to obtain an image PDF document;

[0015] S12, using an optical character recognition algorithm to recognize text in the imaged PDF document and determine a text box position;

[0016] S13, performing layout analysis on the imaged PDF document in which the text box position is determined, identifying and classifying layout elements, and determining the position of the layout elements in the document;

[0017] S14, processing the pictures, tables, and formulas in the layout elements respectively, and aggregating the text boxes and text type layout elements to form paragraphs and titles;

[0018] S15. Based on the reading order of the document, the layout elements are restored in order to ensure that the presentation order of the information is consistent with the original document, and the corresponding layout knowledge blocks are output.

[0019] Furthermore, in S14, the pictures, tables, and formulas in the layout elements are processed separately, including:

[0020] Gain in-depth understanding of the image and obtain the title, description, and context corresponding to each image layout element through a greedy algorithm;

[0021] Restore the structure of the table, restore its original row and column structure, and output it in markdown format;

[0022] Identify formulas and extract their structure and content.

[0023] Furthermore, the S2 includes:

[0024] S21, decomposing the layout knowledge block into two types of information: text and picture;

[0025] S22, dividing the text information into blocks to generate text knowledge blocks; wherein the length of each text knowledge block is 512 to 1024;

[0026] S23, extracting image information, constructing a prompt input multimodal large model, and generating image knowledge blocks for retrieval and question-answer generation; wherein each image knowledge block includes summary text information and image semantic information;

[0027] S24: Generate an embedding vector based on the text knowledge block and the image knowledge block, store it in a vector database, and establish a corresponding index to support similarity retrieval.

[0028] Furthermore, in S24, generating an embedding vector based on the text knowledge block and the image knowledge block includes:

[0029] Use the bge-m3 model to generate a 1024-dimensional embedding vector for the text knowledge block;

[0030] The bge-m3 model is used to generate a semantic embedding vector containing the summary text information, and the bge-m3 model and the visualized-m3 model are used to generate a 1024-dimensional embedding vector containing image semantic information for the image semantic information.

[0031] Furthermore, the S3 includes:

[0032] S31, receiving query information input by a user; the query information is text information, picture information, or a combination of text and picture information;

[0033] S32, if the query information is text information, convert the text information into a 1024-dimensional embedding vector; if the query information is image information, convert the image information into a 1024-dimensional embedding vector; if the query information is a combination of text and image information, convert the text information and image information into a mixed text-image embedding vector;

[0034] S33, sending the corresponding embedded vector to a vector database for multimodal retrieval to obtain the knowledge block with the highest similarity;

[0035] S34, downgrading the knowledge blocks, and sorting the downgraded search knowledge blocks according to cosine similarity;

[0036] S35. Remove duplicate items from the sorted results, and re-sort the results after removing duplicates to obtain multimodal retrieval results.

[0037] Furthermore, the S4 includes:

[0038] S41. If the multimodal search result is an empty search result, a prompt instruction corresponding to the query information without additionally citing knowledge is generated through the chat generator; if the multimodal search result is a non-empty search result, the non-empty search results of the traceability reference generator are combined to generate a prompt instruction for a traceability task;

[0039] S42. Based on the prompt instruction, select a language large model or a multimodal large model to generate a reply result corresponding to the user query information.

[0040] In a second aspect, the present invention provides an intelligent reply system for user query information, comprising:

[0041] Document parsing module: used to perform structural processing on documents provided by users and construct layout knowledge blocks containing document logical structure and multimodal information;

[0042] Vector database construction module: used for decomposing the layout knowledge block into text knowledge block and picture knowledge block, generating embedding vectors for the text knowledge block and the picture knowledge block, and storing them in the vector database;

[0043] Multimodal retrieval module: used to perform multimodal retrieval on the vector database in combination with user query information to obtain multimodal retrieval results;

[0044] A reply output module is used to output a reply result corresponding to the user query information through a large model based on the multimodal retrieval result.

[0045] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned intelligent reply method under user query information when executing the computer program.

[0046] In a fourth aspect, the present invention provides a computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the intelligent reply method under the user query information described above is implemented.

[0047] The descriptions of the second to fourth aspects of the present invention can refer to the detailed description of the first aspect; and the beneficial effects of the descriptions of the second to fourth aspects can refer to the analysis of the beneficial effects of the first aspect, which will not be repeated here.

[0048] It can be seen from the above technical solution that compared with the prior art, the present invention has the following beneficial effects:

[0049] 1. This method can simultaneously process information in multiple modalities, such as text, images, tables, formulas, etc., through deep analysis and structured processing of documents, thus achieving a comprehensive understanding and efficient use of document content, and significantly improving the breadth and depth of information processing.

[0050] 2. By decomposing document knowledge blocks into two types of information, text and images, and generating embedded vectors for each type of information and storing them in the vector database, the rapid storage and efficient retrieval of documents are achieved. At the same time, the multimodal retrieval strategy is adopted to handle complex queries of text, images and their combination, providing users with more accurate and comprehensive search results.

[0051] 3. It can intelligently process multiple types of user input, including text, images, or a combination of the two. Through question extraction, retrieval, and source reference generation, combined with large model processing, it can generate response results that are highly matched with user query information. In addition, when the search results are empty or the user's question is more suitable for small talk, it can automatically switch to the small talk generator to provide a humanized interactive experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0053] Figure 1 A flow chart of an intelligent reply method under user query information provided by an embodiment of the present invention;

[0054] Figure 2 A schematic diagram of a process for constructing a layout knowledge block including a document logical structure and multimodal information provided by an embodiment of the present invention;

[0055] Figure 3 A schematic diagram of a process for generating embedding vectors for text knowledge blocks and image knowledge blocks provided in an embodiment of the present invention;

[0056] Figure 4 A schematic diagram of a process for obtaining multimodal search results provided by an embodiment of the present invention;

[0057] Figure 5 A schematic diagram of a process of outputting a reply result corresponding to user query information through a large model provided by an embodiment of the present invention;

[0058] Figure 6 A structural framework diagram of an intelligent reply system for user query information provided by an embodiment of the present invention;

[0059] Figure 7 A schematic diagram of the structure of an electronic device is provided for an embodiment of the present invention. DETAILED DESCRIPTION

[0060] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0061] The detection method provided in the embodiment of the present application can be applied to a detection server. The above-mentioned detection server can be hardware or software. When the detection server is hardware, it can be implemented as a distributed server cluster that provides detection services, or it can be implemented as a single server. When the detection server is software, it can be installed in the servers listed above. It can be implemented as multiple software or software modules, or it can be implemented as a single software or software module, which is not specifically limited here.

[0062] Embodiment 1;

[0063] like Figure 1 As shown, the present embodiment provides an intelligent reply method for user query information, comprising the following steps:

[0064] S1. Structural processing of documents provided by users to construct layout knowledge blocks containing document logical structure and multimodal information;

[0065] S2, decomposing the layout knowledge block into a text knowledge block and a picture knowledge block, generating embedding vectors for the text knowledge block and the picture knowledge block, and storing them in a vector database;

[0066] S3. Perform a multimodal search on the vector database in combination with the user query information to obtain a multimodal search result;

[0067] S4. Based on the multimodal retrieval results, a reply result corresponding to the user query information is output through a large model.

[0068] This method provides more comprehensive information services by supporting cross-modal retrieval capabilities to process multiple data types; adopts an optimized image summary extraction strategy to increase the amount of information and retrieval accuracy, especially in scenarios with rich image information; and uses structured parsed document information and advanced embedding vector generation technology to achieve high-quality question and answer generation, enhancing the user experience while providing specific citation basis, significantly improving the credibility and accuracy of the answers.

[0069] The following is a further detailed description of each step in the above method:

[0070] In this embodiment S1, Figure 2 As shown, the document provided by the user is structured to construct a layout knowledge block containing the document logical structure and multimodal information; specifically, it includes:

[0071] S11, standardizing the format and converting the document provided by the user into an image to obtain an image PDF document;

[0072] S12, using an optical character recognition algorithm to recognize text in the imaged PDF document and determine a text box position;

[0073] S13, performing layout analysis on the imaged PDF document in which the text box position is determined, identifying and classifying layout elements, and determining the position of the layout elements in the document;

[0074] S14, processing the pictures, tables, and formulas in the layout elements respectively, and aggregating the text boxes and text type layout elements to form paragraphs and titles;

[0075] S15. Based on the reading order of the document, the layout elements are restored in order to ensure that the presentation order of the information is consistent with the original document, and the corresponding layout knowledge blocks are output.

[0076] Furthermore, in S14, the pictures, tables, and formulas in the layout elements are processed respectively, including:

[0077] Gain in-depth understanding of the image and obtain the title, description, and context corresponding to each image layout element through a greedy algorithm;

[0078] Restore the structure of the table, restore its original row and column structure, and output it in markdown format;

[0079] Identify formulas and extract their structure and content.

[0080] This step converts the document provided by the user into a layout knowledge block containing rich logical structure and multimodal information. It not only focuses on the accuracy and completeness of the information, but also fully considers the characteristics and processing methods of different layout elements (pictures, tables, formulas). Combined with the application of greedy algorithms, it can deeply mine the content of pictures, restore tables in a structured manner, and accurately identify formulas, thereby providing users with more comprehensive and accurate information support.

[0081] In this embodiment S2, Figure 3 As shown, the layout knowledge block is decomposed into a text knowledge block and an image knowledge block, and embedding vectors are generated for the text knowledge block and the image knowledge block, and stored in a vector database; specifically, it includes:

[0082] S21, decomposing the layout knowledge block into two types of information: text and picture;

[0083] S22, dividing the text information into blocks to generate text knowledge blocks; wherein the length of each text knowledge block is 512 to 1024;

[0084] S23, extracting image information, constructing a prompt input multimodal large model, and generating image knowledge blocks for retrieval and question-answer generation; wherein each image knowledge block includes summary text information and image semantic information;

[0085] S24: Generate an embedding vector based on the text knowledge block and the image knowledge block, store it in a vector database, and establish a corresponding index to support similarity retrieval.

[0086] Furthermore, in S24, an embedding vector is generated based on the text knowledge block and the image knowledge block, including:

[0087] Use the bge-m3 model to generate a 1024-dimensional embedding vector for the text knowledge block;

[0088] The bge-m3 model is used to generate a semantic embedding vector containing the summary text information, and the bge-m3 model and the visualized-m3 model are used to generate a 1024-dimensional embedding vector containing image semantic information for the image semantic information.

[0089] In this step, the layout knowledge block is decomposed into two types of information: text and picture, and knowledge blocks and embedding vectors are generated for text and pictures respectively, so as to achieve deep analysis and efficient storage of document content. The bge-m3 model is used to generate 1024-dimensional embedding vectors for text knowledge blocks and summary text information, and the bge-m3 and visualized-m3 models are combined to generate embedding vectors for image semantic information, which improves the accuracy and efficiency of information retrieval.

[0090] In this embodiment S3, Figure 4 As shown, in combination with the user query information, a multimodal search is performed on the vector database to obtain a multimodal search result; specifically including:

[0091] S31, receiving query information input by a user; the query information is text information, picture information, or a combination of text and picture information;

[0092] S32, if the query information is text information, convert the text information into a 1024-dimensional embedding vector; if the query information is image information, convert the image information into a 1024-dimensional embedding vector; if the query information is a combination of text and image information, convert the text information and image information into a mixed text-image embedding vector;

[0093] S33, sending the corresponding embedded vector to a vector database for multimodal retrieval to obtain the knowledge block with the highest similarity;

[0094] S34, downgrading the knowledge blocks, and sorting the downgraded search knowledge blocks according to cosine similarity;

[0095] S35. Remove duplicate items from the sorted results, and re-sort the results after removing duplicates to obtain multimodal retrieval results.

[0096] This step receives the query information of text, image or image-text combination input by the user and converts it into a unified 1024-dimensional embedding vector to achieve efficient multimodal retrieval of the vector database. Through down-weighting, cosine similarity sorting and de-duplication optimization, the accuracy and relevance of the search results are improved, providing users with a richer and more accurate search experience.

[0097] In this embodiment S4, Figure 5 As shown, based on the multimodal retrieval results, a reply result corresponding to the user query information is output through a large model. Specifically including:

[0098] S41. If the multimodal search result is an empty search result, a prompt instruction corresponding to the query information without additionally citing knowledge is generated through the chat generator; if the multimodal search result is a non-empty search result, the non-empty search results of the traceability reference generator are combined to generate a prompt instruction for a traceability task;

[0099] S42, based on the prompt instruction, select a language large model or a multimodal large model to generate a reply result corresponding to the user query information. Specifically, the language large model can select the Qwen272B Instruct model, and the multimodal large model can select the Qwen2 VL 72B Instruct model;

[0100] When the multimodal retrieval results are empty, the chat generator is used to provide replies that do not require additional knowledge, thereby enhancing the friendliness and interactivity of the system. When the retrieval results are not empty, the tracing reference generator is used to combine the retrieval results into prompt instructions for the tracing task, providing detailed knowledge sources, and using the large model to generate accurate replies based on this knowledge, providing users with more accurate and comprehensive information services.

[0101] Embodiment 2;

[0102] like Figure 6 As shown, this embodiment provides an intelligent reply system for user query information, including:

[0103] Document parsing module: used to perform structural processing on documents provided by users and construct layout knowledge blocks containing document logical structure and multimodal information;

[0104] Vector database construction module: used for decomposing the layout knowledge block into text knowledge block and picture knowledge block, generating embedding vectors for the text knowledge block and the picture knowledge block, and storing them in the vector database;

[0105] Multimodal retrieval module: used to perform multimodal retrieval on the vector database in combination with user query information to obtain multimodal retrieval results;

[0106] A reply output module is used to output a reply result corresponding to the user query information through a large model based on the multimodal retrieval result.

[0107] The intelligent reply system for user query information provided in this embodiment realizes deep analysis and structured processing of documents, as well as efficient multimodal retrieval and intelligent reply through four major modules: document parsing, vector database construction, multimodal retrieval and reply output. It improves the accuracy of information retrieval and the quality of question and answer generation, and provides users with a richer and more accurate search and question and answer experience.

[0108] Embodiment 3;

[0109] like Figure 7 As shown, this embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the intelligent reply method under the user query information in the above embodiment when executing the computer program.

[0110] Embodiment 4;

[0111] This embodiment provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the intelligent reply method in the above embodiment is implemented.

[0112] In the above embodiments provided in the present application, it should be understood that the disclosed methods, systems, devices and media can be implemented in other ways. The above described methods, systems, devices and media embodiments are only schematic. For example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation. Each functional unit can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0113] The units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0114] The computer program includes computer program code, which may be in source code form, object code form, executable file or some intermediate form, etc. Computer readable media may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal and software distribution medium, etc.

[0115] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for intelligent reply to user query information, characterized in that: The following steps are involved: S1. Structural processing of documents provided by users to construct layout knowledge blocks containing document logical structure and multimodal information; S2, decomposing the layout knowledge block into a text knowledge block and a picture knowledge block, generating embedding vectors for the text knowledge block and the picture knowledge block, and storing them in a vector database; S3. Perform a multimodal search on the vector database in combination with the user query information to obtain a multimodal search result; S4. Based on the multimodal retrieval results, a reply result corresponding to the user query information is output through a large model.

2. According to claim 1, a smart reply method for user query information is characterized in that: Said S1 comprises: S11, standardizing the format and converting the document provided by the user into an image to obtain an image PDF document; S12, using an optical character recognition algorithm to recognize text in the imaged PDF document and determine a text box position; S13, performing layout analysis on the imaged PDF document in which the text box position is determined, identifying and classifying layout elements, and determining the position of the layout elements in the document; S14, processing the pictures, tables, and formulas in the layout elements respectively, and aggregating the text boxes and text type layout elements to form paragraphs and titles; S15. Based on the reading order of the document, the layout elements are restored in order to ensure that the presentation order of the information is consistent with the original document, and the corresponding layout knowledge blocks are output.

3. The intelligent reply method for user query information according to claim 3 is characterized in that: In S14, the pictures, tables, and formulas in the layout elements are processed respectively, including: Gain in-depth understanding of the image and obtain the title, description, and context corresponding to each image layout element through a greedy algorithm; Restore the structure of the table, restore its original row and column structure, and output it in markdown format; Identify formulas and extract their structure and content.

4. The intelligent reply method for user query information according to claim 1, characterized in that: The S2 comprises: S21, decomposing the layout knowledge block into two types of information: text and picture; S22, dividing the text information into blocks to generate text knowledge blocks; wherein the length of each text knowledge block is 512 to 1024; S23, extracting image information, constructing a prompt input multimodal large model, and generating image knowledge blocks for retrieval and question-answer generation; wherein each image knowledge block includes summary text information and image semantic information; S24: Generate an embedding vector based on the text knowledge block and the image knowledge block, store it in a vector database, and establish a corresponding index to support similarity retrieval.

5. The intelligent reply method for user query information according to claim 4, characterized in that: In S24, generating an embedding vector based on the text knowledge block and the image knowledge block includes: Use the bge-m3 model to generate a 1024-dimensional embedding vector for the text knowledge block; The bge-m3 model is used to generate a semantic embedding vector containing the summary text information, and the bge-m3 model and the visualized-m3 model are used to generate a 1024-dimensional embedding vector containing image semantic information for the image semantic information.

6. The intelligent reply method for user query information according to claim 1, characterized in that: The S3 includes: S31, receiving query information input by a user; the query information is text information, picture information, or a combination of text and picture information; S32, if the query information is text information, convert the text information into a 1024-dimensional embedding vector; if the query information is image information, convert the image information into a 1024-dimensional embedding vector; if the query information is a combination of text and image information, convert the text information and image information into a mixed text-image embedding vector; S33, sending the corresponding embedded vector to a vector database for multimodal retrieval to obtain the knowledge block with the highest similarity; S34, downgrading the knowledge blocks, and sorting the downgraded search knowledge blocks according to cosine similarity; S35. Remove duplicate items from the sorted results, and re-sort the results after removing duplicates to obtain multimodal retrieval results.

7. The intelligent reply method for user query information according to claim 1, characterized in that: The S4 comprises: S41. If the multimodal search result is an empty search result, a prompt instruction corresponding to the query information without additionally citing knowledge is generated through the chat generator; if the multimodal search result is a non-empty search result, the non-empty search results of the traceability reference generator are combined to generate a prompt instruction for a traceability task; S42. Based on the prompt instruction, select a language large model or a multimodal large model to generate a reply result corresponding to the user query information.

8. An intelligent reply system for user query information, characterized in that: include: Document parsing module: used to perform structural processing on documents provided by users and construct layout knowledge blocks containing document logical structure and multimodal information; Vector database construction module: used for decomposing the layout knowledge block into text knowledge block and picture knowledge block, generating embedding vectors for the text knowledge block and the picture knowledge block, and storing them in the vector database; Multimodal retrieval module: used to perform multimodal retrieval on the vector database in combination with user query information to obtain multimodal retrieval results; Reply output module; Used to output a reply result corresponding to the user query information through a large model based on the multimodal retrieval result.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the intelligent reply method under user query information as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the intelligent reply method under user query information as described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Document processing method, document question and answer method and computer program product

    CN121561092A