File material retrieval method and device, electronic equipment and storage medium
By structuring the target file and logical structure division, combining the semantic matching of the double tower model and the large language model, the problem of inaccurate search results in the existing technology is solved, and more efficient and accurate file retrieval is achieved.
Patent Information
- Application Number
- CN202510321709.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-18
AI Technical Summary
The existing technology has difficulty in understanding the deep intention and organizational orientation behind the documents in depth, resulting in inaccurate and targeted search results.
By structuring the target file, extracting meta information, and dividing knowledge blocks according to the logical structure of the file content, using the double tower model and the large language model for semantic matching and encoding processing, the accuracy of the search results is improved.
It enhances the semantic understanding ability during the search process, can more accurately grasp the correlation between user intentions and knowledge blocks, and improves the accuracy and efficiency of search results.
Smart Images

Figure CN120336265A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, and particularly relates to a method, apparatus, electronic device, and storage medium for retrieving document materials. Background Art
[0002] In the field of document management and retrieval, information retrieval technology is the cornerstone of the work of various organizations and institutions. Accurate retrieval capabilities can not only improve the standardization of work, avoid risks, but also assist in the design of activity plans, providing a basis for decision-making and activities.
[0003] Retrieval technology based on Artificial Intelligence (AI) can quickly and accurately screen a large number of normative articles, intelligently analyze their relevance and scope of application, greatly improving retrieval efficiency and accuracy. However, due to the strong professionalism and particularity of some documents, involving many complex management, requirements, and work processes, etc. Related retrieval technologies are difficult to deeply understand the underlying intentions, requirements, and organizational orientations behind documents like humans. When processing specific retrieval requests, they may only mechanically match keywords and cannot comprehensively consider various factors to provide more in-depth and targeted retrieval results. Summary of the Invention
[0004] This application aims to solve at least one of the technical problems existing in the prior art. For this purpose, this application proposes a method, apparatus, electronic device, and storage medium for retrieving document materials to enhance the semantic understanding ability during retrieval and improve the accuracy of retrieval results.
[0005] In a first aspect, this application provides a method for retrieving document materials, including:
[0006] Structuring the basic information of the target document to obtain structured meta-information;
[0007] Dividing the document content according to the logical structure of the document content in the target document and associating it with the meta-information to obtain multiple knowledge blocks;
[0008] Retrieving a target knowledge block associated with the question from multiple knowledge blocks according to the similarity between the question input by the user and the multiple knowledge blocks.
[0009] The retrieval method for document materials provided by the embodiments of the present application obtains structured meta-information by performing structured processing on the basic information of the target document; divides the document content according to the logical structure of the document content in the target document and associates it with the meta-information to obtain a plurality of knowledge blocks; and retrieves a target knowledge block associated with the question from the plurality of knowledge blocks according to the similarity between the question input by the user and the plurality of knowledge blocks. Based on the characteristics of the target document, the embodiments of the present application perform structured processing on the basic information of the target document, extract meta-information, and divide the document content according to the logical structure of the document content to obtain a plurality of knowledge blocks, so that the knowledge blocks contain rich semantic information, enhancing the semantic understanding ability in the retrieval process, and thus enabling the retrieval process to more accurately grasp the relevance between the user's intention and the semantics of the knowledge blocks, improving the accuracy of the retrieval results.
[0010] According to an embodiment of the present application, the retrieving a target knowledge block associated with the question from the plurality of knowledge blocks according to the similarity between the question input by the user and the plurality of knowledge blocks includes:
[0011] Obtaining a first feature vector of the question and second feature vectors of the plurality of knowledge blocks respectively;
[0012] Calculating the similarity between the first feature vector and the second feature vectors;
[0013] Determining the knowledge block corresponding to the second feature vector whose similarity to the first feature vector is greater than a preset threshold as the target knowledge block.
[0014] In this embodiment, through the extraction of feature vectors, the knowledge blocks and the question can be transformed into a quantifiable and comparable mathematical expression form, enabling the structured processing of semantic information. Based on the similarity calculation, the knowledge block closest to the semantics of the user's question can be quickly screened out, improving the accuracy of the retrieval results.
[0015] According to an embodiment of the present application, the retrieving a target knowledge block associated with the question from the plurality of knowledge blocks according to the similarity between the question input by the user and the plurality of knowledge blocks includes:
[0016] Inputting the question and the plurality of knowledge blocks into a preset dual-tower model; wherein, the dual-tower model is used to extract a first feature vector of the question and second feature vectors of the plurality of knowledge blocks respectively, calculate the similarity between the first feature vector and the second feature vectors, and output a target knowledge block based on the similarity.
[0017] In this embodiment, through the semantic matching of the user input question and the knowledge chunks by the dual - tower model, the semantic features of the question and the knowledge chunks can be fully captured, and the similarity of the feature vectors can be calculated in the shared semantic space to determine the association between the question and the knowledge chunks, reducing the feature confusion problem caused by a single model, enhancing the semantic understanding ability of the retrieval, and improving the accuracy of the retrieval results.
[0018] According to an embodiment of the present application, the method further includes:
[0019] Storing the second feature vectors of the multiple knowledge chunks extracted by the dual - tower model into a preset vector database.
[0020] In this embodiment, by storing the second feature vectors of the knowledge chunks extracted by the dual - tower model into a preset vector database, fast retrieval and multiple utilization of a large number of knowledge chunks can be realized, reducing the overhead of recalculating feature vectors during each retrieval, and significantly enhancing the retrieval efficiency.
[0021] According to an embodiment of the present application, retrieving the target knowledge chunk associated with the question from the multiple knowledge chunks according to the similarity between the question input by the user and the multiple knowledge chunks includes:
[0022] Screening out the first candidate knowledge chunks from the multiple knowledge chunks according to the similarity between the question input by the user and the multiple knowledge chunks;
[0023] Inputting the first candidate knowledge chunks into a preset large - language model; wherein, the large - language model screens out the target knowledge chunks from the first candidate knowledge chunks according to the degree of association between the question and the first candidate knowledge chunks.
[0024] In this embodiment, through preliminary screening according to the similarity between the question input by the user and the knowledge chunks, the first candidate knowledge chunks are obtained from a large number of knowledge chunks, and the first candidate knowledge chunks are input into a preset large - language model. By utilizing the powerful language understanding and generation ability of the large - language model, the degree of association between the question and the candidate knowledge chunks can be deeply analyzed, enabling the retrieval results to meet the user's needs, further improving the accuracy of the retrieval results, and being applicable to complex text retrieval tasks.
[0025] According to an embodiment of the present application, inputting the candidate knowledge chunks into a preset large - language model includes:
[0026] Extracting the first encoding corresponding to the first candidate knowledge chunks from a preset knowledge base; wherein, the knowledge base stores multiple knowledge chunks and multiple first encodings corresponding to the multiple knowledge chunks, and each knowledge chunk corresponds to a first encoding as a unique identifier;
[0027] Construct a first prompt message based on the first encoding corresponding to the first candidate knowledge block;
[0028] Input the first prompt message and the first candidate knowledge block into the large language model; wherein, based on the guidance of the first prompt message, the large language model screens out the target knowledge block from the first candidate knowledge blocks according to the degree of association between the question and the first candidate knowledge blocks, and outputs the target encoding corresponding to the target knowledge block.
[0029] In this embodiment, a unique identifier is assigned to each knowledge block through an encoding mechanism, and the prompt message guides the large language model to output the encoding instead of directly outputting the original text of the target file, reducing the problem of inappropriate addition, deletion, and modification of the original text of the target file caused by hallucinations that the large language model may generate. Through the encoding output by the large language model, the original text corresponding to the encoding can be further matched from the knowledge blocks, improving the normativity of the content output by the retrieval results.
[0030] According to an embodiment of the present application, the inputting the first candidate knowledge block into a preset large language model includes:
[0031] Divide the first candidate knowledge blocks into multiple knowledge block groups based on the number of concurrent inference paths of the large language model;
[0032] Input the multiple knowledge block groups into the respective inference paths of the large language model to obtain second candidate knowledge blocks screened by each inference path from each knowledge block group;
[0033] Input the second candidate knowledge blocks into the large language model; wherein, the large language model screens out the target knowledge block from the second candidate knowledge blocks according to the degree of association between the question and the second candidate knowledge blocks.
[0034] In this embodiment, by dividing the first candidate knowledge blocks into multiple knowledge block groups and using the concurrent inference paths of the large language model for multi-stage screening, the parallel processing ability of the model is fully utilized, reducing the efficiency bottleneck problem caused by processing a large number of knowledge blocks at one time. It not only reduces the burden on a single inference path but also can significantly shorten the retrieval time through parallel computing, greatly improving the retrieval efficiency.
[0035] According to an embodiment of the present application, the inputting the second candidate knowledge blocks into the large language model includes:
[0036] Construct a second encoding for the second candidate knowledge blocks; wherein the second encoding is the unique identifier of the second candidate knowledge blocks;
[0037] Construct a second prompt message according to the second encoding corresponding to the second candidate knowledge blocks;
[0038] Input the second prompt information and the second candidate knowledge block into the large language model; wherein, based on the guidance of the second prompt information, the large language model screens out the target knowledge block from the second candidate knowledge blocks according to the degree of association between the question and the second candidate knowledge blocks, and outputs the target encoding corresponding to the target knowledge block.
[0039] In this embodiment, by re-encoding the second candidate knowledge blocks output by multiple inference paths, the continuity of the encoding can be improved, making the input of the large language model more in line with the inference habit of the large language model, and improving the accuracy of the retrieval results.
[0040] Second, the present application provides a retrieval method for document materials, including:
[0041] Input the question input by the user into a preset two-tower model; wherein, the two-tower model is used to extract the first feature vector of the question, extract the second feature vectors of multiple knowledge blocks from a preset vector database, and screen out the first candidate knowledge blocks from the multiple knowledge blocks according to the similarity between the first feature vector and the second feature vectors; the multiple knowledge blocks are obtained by dividing the document content according to the logical structure of the document content in the target document and associating with meta-information, and the meta-information is obtained by structuring the basic information of the target document;
[0042] Extract the first encoding corresponding to the first candidate knowledge block from a preset knowledge base; wherein, the knowledge base stores multiple knowledge blocks and multiple first encodings corresponding to the multiple knowledge blocks, and each knowledge block corresponds to a first encoding as a unique identifier;
[0043] Construct the first prompt information according to the first encoding corresponding to the first candidate knowledge block;
[0044] Input the first prompt information and the first candidate knowledge block into a preset large language model; wherein, based on the guidance of the first prompt information, the large language model screens out the target knowledge block from the first candidate knowledge blocks according to the degree of association between the question and the first candidate knowledge blocks, and outputs the target encoding corresponding to the target knowledge block.
[0045] The retrieval method for document materials provided by the embodiments of the present application, based on the characteristics of the target document, structures the basic information of the target document, extracts meta-information, and divides the document content according to the logical structure of the document content to obtain multiple knowledge blocks, so that the knowledge blocks contain rich semantic information, enhancing the semantic understanding ability in the retrieval process, thereby enabling the retrieval process to more accurately grasp the relevance between the user's intention and the semantics of the knowledge blocks, and improving the accuracy of the retrieval results.
[0046] According to an embodiment of the present application, inputting the first prompt information and the first candidate knowledge block into a preset large language model includes:
[0047] Dividing the first candidate knowledge block into multiple knowledge block groups based on the number of inference paths concurred by the large language model;
[0048] Inputting the multiple knowledge block groups into respective inference paths of the large language model to obtain second candidate knowledge blocks screened by each inference path from each knowledge block group;
[0049] Constructing a second encoding for the second candidate knowledge block; wherein, the second encoding is the unique identifier of the second candidate knowledge block;
[0050] Constructing second prompt information according to the second encoding corresponding to the second candidate knowledge block;
[0051] Inputting the second prompt information and the second candidate knowledge block into the large language model; wherein, based on the guidance of the second prompt information, the large language model screens out a target knowledge block from the second candidate knowledge blocks according to the degree of association between the question and the second candidate knowledge blocks, and outputs the target encoding corresponding to the target knowledge block.
[0052] In a third aspect, the present application provides a retrieval device for document materials, and the device includes:
[0053] A processing module, configured to perform structured processing on basic information of a target document to obtain structured meta-information;
[0054] A partitioning module, configured to partition the document content according to the logical structure of the document content in the target document and associate it with the meta-information to obtain multiple knowledge blocks;
[0055] A retrieval module, configured to retrieve a target knowledge block associated with the question from the multiple knowledge blocks according to the similarity between the question input by the user and the multiple knowledge blocks.
[0056] The retrieval device for document materials provided by the embodiments of the present application processes the basic information of the target document to obtain structured meta-information; divides the document content according to the logical structure of the document content in the target document and associates it with the meta-information to obtain multiple knowledge blocks; and retrieves the target knowledge block associated with the question from the multiple knowledge blocks according to the similarity between the question input by the user and the multiple knowledge blocks. Based on the characteristics of the target document, the embodiments of the present application structurally process the basic information of the target document, extract meta-information, and divide the document content according to the logical structure of the document content to obtain multiple knowledge blocks, so that the knowledge blocks contain rich semantic information, enhancing the semantic understanding ability in the retrieval process, thereby enabling the retrieval process to more accurately grasp the relevance between the user's intention and the semantics of the knowledge blocks and improving the accuracy of the retrieval results.
[0057] In a fourth aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the retrieval method for document materials as described in the first aspect above.
[0058] In a fifth aspect, the present application provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the retrieval method for document materials as described in the first aspect above.
[0059] In a sixth aspect, the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the retrieval method for document materials as described in the first aspect above.
[0060] One or more of the above technical solutions in the embodiments of the present application have at least one of the following technical effects:
[0061] The retrieval method for document materials provided by the embodiments of the present application processes the basic information of the target document to obtain structured meta-information; divides the document content according to the logical structure of the document content in the target document and associates it with the meta-information to obtain multiple knowledge blocks; and retrieves the target knowledge block associated with the question from the multiple knowledge blocks according to the similarity between the question input by the user and the multiple knowledge blocks. Based on the characteristics of the target document, the embodiments of the present application structurally process the basic information of the target document, extract meta-information, and divide the document content according to the logical structure of the document content to obtain multiple knowledge blocks, so that the knowledge blocks contain rich semantic information, enhancing the semantic understanding ability in the retrieval process, thereby enabling the retrieval process to more accurately grasp the relevance between the user's intention and the semantics of the knowledge blocks and improving the accuracy of the retrieval results.
[0062] Furthermore, through the extraction of feature vectors, knowledge chunks and questions can be transformed into quantifiable and comparable mathematical expression forms, enabling semantic information to be processed structurally. Based on similarity calculation, knowledge chunks with the closest semantics to the user's question can be quickly screened out, improving the accuracy of retrieval results.
[0063] Furthermore, through the semantic matching of the user's input question and knowledge chunks by the dual-tower model, the semantic features of the question and knowledge chunks can be fully captured, and the similarity of feature vectors can be calculated in a shared semantic space to determine the association between the question and knowledge chunks, reducing the problem of feature confusion caused by a single model, enhancing the semantic understanding ability of retrieval, and improving the accuracy of retrieval results.
[0064] Furthermore, by storing the second feature vectors of the knowledge chunks extracted by the dual-tower model in a preset vector database, fast retrieval and multiple utilization of a large number of knowledge chunks can be realized, reducing the overhead of recalculating feature vectors during each retrieval and significantly improving the retrieval efficiency.
[0065] Furthermore, by performing a preliminary screening based on the similarity between the user's input question and knowledge chunks to obtain the first candidate knowledge chunks from a large number of knowledge chunks, and inputting the first candidate knowledge chunks into a preset large language model, the strong language understanding and generation ability of the large language model can be utilized to deeply analyze the degree of association between the question and the candidate knowledge chunks, enabling the retrieval results to meet the user's needs, further improving the accuracy of retrieval results, and being applicable to complex text retrieval tasks.
[0066] Furthermore, through the encoding mechanism, a unique identifier is assigned to each knowledge chunk, and the prompt information guides the large language model to output the encoding instead of directly outputting the original text of the target file, reducing the problem of inappropriate addition, deletion, and modification of the original text of the target file caused by hallucinations that may occur in the large language model. Through the encoding output by the large language model, the original text corresponding to the encoding can be further matched from the knowledge chunks, improving the normality of the content output by the retrieval results.
[0067] Furthermore, by dividing the first candidate knowledge chunks into multiple knowledge chunk groups and using the concurrent inference paths of the large language model for multi-stage screening, the parallel processing ability of the model is fully utilized to reduce the efficiency bottleneck problem caused by processing a large number of knowledge chunks at one time. This not only reduces the burden on a single inference path but also can significantly shorten the retrieval time through parallel computing, greatly improving the retrieval efficiency.
[0068] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other relevant drawings can also be obtained based on these drawings.
[0070] Figure 1 is one of the schematic flowcharts of the retrieval method for document materials provided by the embodiments of the present application;
[0071] Figure 2 is the schematic architecture diagram of the dual-tower model provided by the embodiments of the present application;
[0072] Figure 3 is the schematic diagram of the scenario example provided by the embodiments of the present application;
[0073] Figure 4 is the second schematic flowchart of the retrieval method for document materials provided by the embodiments of the present application;
[0074] Figure 5 is the schematic structural diagram of the retrieval device for document materials provided by the embodiments of the present application;
[0075] Figure 6 is the schematic structural diagram of the electronic device provided by the embodiments of the present application. Detailed Embodiments
[0076] The following will clearly describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application fall within the scope of protection of the present application.
[0077] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects and are not used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0078] The following will, with reference to the accompanying drawings, detail the retrieval method, device, electronic device, and storage medium for document materials provided by the embodiments of the present application through specific embodiments and their application scenarios.
[0079] Among them, the retrieval method of document materials can be applied to a terminal, and can be specifically executed by hardware or software in the terminal.
[0080] The terminal includes, but is not limited to, portable communication devices such as mobile phones or tablet computers having a touch-sensitive surface (e.g., a touch screen display and / or a touchpad). It should also be understood that in some embodiments, the terminal may not be a portable communication device, but a desktop computer having a touch-sensitive surface (e.g., a touch screen display and / or a touchpad).
[0081] In the following various embodiments, a terminal including a display and a touch-sensitive surface is described. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, a mouse, and a joystick.
[0082] The retrieval method of document materials provided by the embodiments of the present application may be executed by an electronic device or a functional module or functional entity capable of implementing the method in the electronic device. The electronic devices mentioned in the embodiments of the present application include, but are not limited to, mobile phones, tablet computers, computers, cameras, and wearable devices, etc. Hereinafter, taking the electronic device as the execution subject, the retrieval method of document materials provided by the embodiments of the present application will be described.
[0083] As Figure 1 shown, the retrieval method of document materials includes: Step 110, Step 120, and Step 130.
[0084] Step 110: Structurally process the basic information of the target file to obtain structured meta-information.
[0085] In the embodiments of the present application, the target file may be an object of information retrieval, and the target file may be a specific document or data set. The target file may include important information related to an organization, a project, or an activity, such as a normative document (standard, process specification, etc.), a report, a meeting record, a statute, etc. Of course, the target file may also be other types of files, and the present application does not limit this.
[0086] In the embodiments of the present application, the basic information of the target file refers to the metadata used to describe its own attributes, and the basic information may include: file name, issuance time, issuing agency, business type, status, etc. Among them, the business type can be used to classify the business fields involved in the target file. For example, in the investment banking field, the business type may include securities-related, listing-related, asset-related, investment-related, etc. The status is used to identify the current status of the target file, such as abolished, revised, trial implementation, and current.
[0087] Meta-information is structured information that is processed by structuring basic information. It is presented in a structured form to facilitate parsing and indexing by computer programs. In the process of structuring basic information, the structure of basic information needs to be clarified. For example, the issuance time can be in the format of "YYYY year MM month DD day".
[0088] Step 120: Divide the file content according to the logical structure of the file content in the target file, and associate it with the meta-information to obtain multiple knowledge blocks.
[0089] In an embodiment of the present application, the logical structure of the file content may be the internal logic and hierarchical relationship followed by the target file when organizing and expressing its clauses, requirements, instructions, etc. For example, the logical structure may include chapters, sections, clauses, and other levels.
[0090] In some embodiments, since the target file may contain tables, pictures, and other contents in addition to text, for tables, the table contents may be converted into Markdown text format to maintain the readability and structure of the information; for pictures, AI technology and manual description may be combined to generate text descriptions to ensure the integrity and accessibility of the information.
[0091] In the embodiment of the present application, the chapter, section, and other levels in the file content can be identified, and the text content can be divided into multiple sub-texts based on these levels, and the sub-texts are associated with meta-information to obtain multiple knowledge blocks. The following is an example template of a knowledge block:
[0092] “×××× Law (issuing date: YYYY-MM-DD; issuing body: ××; business type: ××; status: ××): ×× Chapter: ×× Section: ×× Article: Content.”
[0093] Here is a real example of a knowledge block:
[0094] "XX Company Employee Handbook" (Issuing date: December 19, 2023; Issuing body: XX Company; Business area: Company rules and regulations; Status: Current): Chapter 1 General Provisions: Employees should abide by laws, regulations and company rules and regulations, be honest and trustworthy, and eliminate any form of fraud, bribery or improper behavior.
[0095] In an embodiment of the present application, a knowledge block is constructed in the above manner, the basic information of the target file and the file content are structured, and the file content is divided by a logical structure, so that the semantic information of the knowledge block is richer, and the semantics of the file content can be better understood in the subsequent retrieval process, thereby improving the accuracy of the retrieval.
[0096] Step 130: Retrieve the target knowledge chunks associated with the question from multiple knowledge chunks according to the similarity between the question input by the user and the multiple knowledge chunks.
[0097] In some embodiments, the question input by the user in the user interface can be obtained. The question can be a natural language question, reflecting the user's query need for the content of the target file. For example, the question can be "What are the penalty measures for violations in the management regulations of the XX industry?", "How to understand the'management measures' clause in the 'XX Regulations'", etc.
[0098] In some embodiments, feature extraction can be performed on the question input by the user and multiple knowledge chunks respectively, and the similarity between the feature vectors of the question input by the user and the multiple knowledge chunks can be compared through similarity measurement methods, such as cosine similarity, Euclidean distance, semantic similarity, etc. According to the similarity calculation results, the top K knowledge chunks with the highest similarity can be selected as the target knowledge chunks.
[0099] The retrieval method of document materials provided by the embodiments of the present application obtains structured meta-information by performing structured processing on the basic information of the target document; divides the document content according to the logical structure of the document content in the target document and associates it with the meta-information to obtain multiple knowledge chunks; retrieves the target knowledge chunks associated with the question from the multiple knowledge chunks according to the similarity between the question input by the user and the multiple knowledge chunks. The embodiments of the present application, based on the characteristics of the target document, perform structured processing on the basic information of the target document to extract meta-information, and divide the document content in combination with the logical structure of the document content to obtain multiple knowledge chunks, so that the knowledge chunks contain rich semantic information, enhance the semantic understanding ability in the retrieval process, and thus enable the retrieval process to more accurately grasp the relevance between the user's intention and the semantics of the knowledge chunks, improving the accuracy of the retrieval results.
[0100] In some embodiments, retrieving the target knowledge chunks associated with the question from the multiple knowledge chunks according to the similarity between the question input by the user and the multiple knowledge chunks includes:
[0101] Obtain the first feature vector of the question and the second feature vectors of the multiple knowledge chunks respectively;
[0102] Calculate the similarity between the first feature vector and the second feature vectors;
[0103] Determine the knowledge chunks corresponding to the second feature vectors whose similarity to the first feature vector is greater than a preset threshold as the target knowledge chunks.
[0104] In this embodiment, natural language processing technology can be used to perform feature extraction on the question input by the user and the multiple knowledge chunks, use the feature vector corresponding to the question as the first feature vector, and use the feature vector corresponding to the knowledge chunk as the second feature vector.
[0105] Specifically, a language model such as BERT (Bidirectional Encoder Representations from Transformer), Sentence-BERT, etc. can be used to extract the feature vectors of the question and multiple knowledge chunks. Other methods such as TF-IDF (Term Frequency-Inverse Document Frequency), Word Embedding, etc. can also be used to extract the feature vectors of the question and multiple knowledge chunks, which are not limited in the embodiments of the present application.
[0106] In this embodiment, any similarity measurement method such as cosine similarity, Euclidean distance, semantic similarity, etc. can be used to calculate the similarity between the first feature vector and the second feature vector. Taking the calculation of the similarity between the first feature vector and the second feature vector by cosine similarity as an example, the value range of cosine similarity is between -1 and 1. The closer the value is to 1, the closer the directions of the two feature vectors are, and the higher the similarity. By calculating the cosine similarity between the first feature vector and each second feature vector, the semantic association degree between the user question and each knowledge chunk can be quantified. For example, if the content of a certain knowledge chunk is highly semantically relevant to the user question, then their feature vectors will be very close in the high-dimensional space, and the cosine similarity value will also be high.
[0107] After calculating the similarity between the user question and each knowledge chunk, the target knowledge chunks can be screened according to a preset threshold. Among them, the setting of the threshold is to improve the quality of the retrieval results. When the similarity between the knowledge chunk and the user question is higher than this threshold, it will be recognized as the target knowledge chunk related to the question.
[0108] Specifically, the preset threshold can be adjusted according to the actual application scenario and user requirements. For example, in a scenario with high requirements for the accuracy of retrieval results, the threshold can be set higher so that the returned target knowledge chunks are highly relevant to the user question; while in a scenario with more requirements for the number of retrieval results, the threshold can be appropriately reduced to increase the number of retrieval results. Of course, in order to control the number of retrieval results, a threshold for the number of retrieval results can also be set. For example, among the knowledge chunks corresponding to the second feature vectors with similarity greater than the preset threshold, the first K knowledge chunks can be determined as the target knowledge chunks, where K is the threshold for the number of retrieval results.
[0109] In this embodiment, through the extraction of feature vectors, the knowledge chunks and questions can be transformed into a quantifiable and comparable mathematical expression form, enabling the semantic information to be structurally processed. Based on the similarity calculation, the knowledge chunks with the closest semantics to the user question can be quickly screened out, improving the accuracy of the retrieval results.
[0110] In some embodiments, according to the similarity between the problem input by the user and multiple knowledge chunks, retrieving the target knowledge chunk associated with the problem from the multiple knowledge chunks includes:
[0111] Inputting the problem and the multiple knowledge chunks into a preset two-tower model; wherein, the two-tower model is used to respectively extract the first feature vector of the problem and the second feature vectors of the multiple knowledge chunks, calculate the similarity between the first feature vector and the second feature vectors, and output the target knowledge chunk based on the similarity.
[0112] The two-tower model is an efficient deep learning architecture. Its core idea is to respectively extract features of two different types of inputs (such as users and items, queries and documents) through two independent neural networks (i.e., "two towers"), and finally map them into the same semantic space. In this embodiment, the problem input by the user and the multiple knowledge chunks are respectively used as different types of inputs and input into the preset two-tower model.
[0113] The design of the two-tower model enables the features of the problem and the knowledge chunks to be extracted independently, which not only improves the computing efficiency but also enables the two-tower model to capture deeper semantic information.
[0114] In this embodiment, the architecture of the two-tower model is as Figure 3 shown. The problem (Question) input by the user can be input into the BERT model on the left. Through the BERT model, the problem is converted into a high-dimensional feature vector (Embedding Vector), that is, the first feature vector. The multiple knowledge chunks (Chunks) can be input into the BERT model on the right. Through the BERT model, the multiple knowledge chunks are converted into high-dimensional feature vectors (Embedding Vectors), that is, the second feature vectors. Among them, the BERT model on the left and the BERT model on the right are the same model.
[0115] In some embodiments, the second feature vectors of the multiple knowledge chunks extracted by the two-tower model can be stored in a preset vector database. As Figure 2 shown, VectorDB is the vector database, which can be used to store and retrieve high-dimensional vectors and support fast similarity search.
[0116] In this embodiment, by storing the second feature vectors of the knowledge chunks extracted by the two-tower model in the preset vector database, fast retrieval and multiple utilization of a large number of knowledge chunks can be realized, the overhead of recalculating feature vectors during each retrieval is reduced, and the retrieval efficiency is significantly improved.
[0117] In this embodiment, efficient similarity calculation tools such as Elasticsearch (ES), Faiss, or Milvus can be used to calculate the similarity between the first feature vector and the second feature vector, retrieve the second feature vectors of the top K knowledge chunks that are most similar to the first feature vector corresponding to the question, and output these K knowledge chunks as the target knowledge chunks.
[0118] In some embodiments, thanks to the powerful performance of the pre-trained BERT model, the two-tower model can be directly put into use. Of course, in order to make the two-tower model better adapt to the dataset related to the target document and thus obtain a more accurate and customized retrieval effect, the BERT model can be further trained. The format of the training data is: {"Q":"×××","Pos_Ans":"×××","Neg_Ans":"×××"}. Where "Q" represents the business question related to the target document, "Pos_Ans" is the correct knowledge chunks in the target document; "Neg_Ans" contains knowledge chunks irrelevant to the question and knowledge chunks related to the question but unable to answer the question correctly, and the ratio between the two can be 3:1 or other ratios to improve the training effect.
[0119] In this embodiment, through the semantic matching of the user input question and the knowledge chunks by the two-tower model, the semantic features of the question and the knowledge chunks can be fully captured, and the similarity of the feature vectors can be calculated in the shared semantic space to determine the association between the question and the knowledge chunks, reducing the feature confusion problem caused by a single model, improving the semantic understanding ability of the retrieval, and enhancing the accuracy of the retrieval results.
[0120] In some embodiments, according to the similarity between the user input question and multiple knowledge chunks, retrieving the target knowledge chunks associated with the question from the multiple knowledge chunks includes:
[0121] Screening out the first candidate knowledge chunks from the multiple knowledge chunks according to the similarity between the user input question and the multiple knowledge chunks;
[0122] Inputting the first candidate knowledge chunks into a preset large language model; wherein, the large language model screens out the target knowledge chunks from the first candidate knowledge chunks according to the degree of association between the question and the first candidate knowledge chunks.
[0123] Large Language Model (LLM) is an artificial intelligence technology based on deep learning. By analyzing a large amount of text data, the large language model learns the structure, grammar, semantics, and context of language, enabling it to perform natural language processing tasks such as text generation, translation, question answering, and dialogue systems. In this embodiment, pre-trained large language models such as the GPT (Generative Pre-trained Transformer) series, T5 (Text-to-Text Transfer Transformer), iFLYTEK Spark, etc. can be used. These large language models have been trained on a large amount of text data and can understand complex language patterns and semantic relationships.
[0124] In this embodiment, based on similarity, the first candidate knowledge chunks can be screened from multiple knowledge chunks. The number of the first candidate knowledge chunks is usually multiple. To further determine which knowledge chunks in the first candidate knowledge chunks are more in line with the user's needs, the first candidate knowledge chunks can be input into the large language model, and the large language model is used to analyze and determine the knowledge chunks most relevant to the user's question.
[0125] Specifically, after receiving the input, the large language model will conduct in-depth semantic analysis on the first candidate knowledge chunks. For example, the large language model will consider the relationship between the content and structure of the first candidate knowledge chunks and the user's question, and evaluate the degree of association between each knowledge chunk in the first candidate knowledge chunks and the user's question, so as to screen out the target knowledge chunks that the large language model believes are most likely to answer the user's question and output the target knowledge chunks to the user. It should be noted that the target knowledge chunks can be one knowledge chunk or multiple knowledge chunks.
[0126] In this embodiment, by initially screening based on the similarity between the user input question and the knowledge chunks, the first candidate knowledge chunks are obtained from a large number of knowledge chunks. The first candidate knowledge chunks are input into a preset large language model. Utilizing the powerful language understanding and generation capabilities of the large language model, the degree of association between the question and the candidate knowledge chunks can be deeply analyzed, enabling the retrieval results to meet the user's needs and further improving the accuracy of the retrieval results, which is applicable to complex text retrieval tasks.
[0127] In some embodiments, inputting the candidate knowledge chunks into a preset large language model includes:
[0128] extracting the first encoding corresponding to the first candidate knowledge chunks from a preset knowledge base; wherein, the knowledge base stores multiple knowledge chunks and multiple first encodings corresponding to the multiple knowledge chunks, and each knowledge chunk corresponds to a first encoding as a unique identifier;
[0129] Construct the first hint message according to the first encoding corresponding to the first candidate knowledge block;
[0130] Input the first hint message and the first candidate knowledge block into the large language model; wherein, based on the guidance of the first hint message, the large language model screens out the target knowledge block from the first candidate knowledge blocks according to the relevance between the question and the first candidate knowledge blocks, and outputs the target encoding corresponding to the target knowledge block.
[0131] The streaming output technology of the large language model undoubtedly brings a smoother experience to users, almost eliminating the perception of waiting. However, in certain specific scenarios, it is often necessary to wait for the large language model to complete the output to obtain the complete semantic information, and then perform subsequent operations. The output efficiency of the large language model is affected by various factors, including the scale of model parameters, the length of the input prompt, the scale of batch processing, and GPU resources.
[0132] Taking the iFlytek Spark 3.5 large model on the public cloud as an example, it takes about 2.5 seconds to output each token. In the rigorous business scenario of industry standard retrieval, it is necessary to ensure the accuracy of the industry standard content of the original text without any deviation. If the large language model is directly allowed to output the original text of the industry standard related to the industry standard field, due to the hallucinations that the model may generate, it may lead to inappropriate additions, deletions, and modifications to the original text of the target document. In order to reduce the hallucination problem while maintaining the high-efficiency output of the large language model, the embodiments of the present application adopt an innovative method: encoding the knowledge blocks of the target document. The large language model can directly output the encoding instead of the original text, and each encoding is very short relative to the original text, such as "[1]", "[3]", "[4]". In this way, the output efficiency of the large language model will be greatly improved. According to the encoding output by the large language model, the original text is retrieved from the key-value pair database or the KV structure memory according to these encodings, thereby effectively reducing the hallucination problem of the large language model.
[0133] Specifically, the encoding can be represented in any way, such as letters, numbers, symbols, or any combination of letters, numbers, and symbols. Taking the encoding in the form of "[*]" as an example, where "*" represents a number, the following are examples of encoding and knowledge blocks:
[0134] [1]: "Employee Handbook of XX Company" (Revised and implemented on July 1, 2024): Chapter 2 New Employee Induction Training: Article 9: After new employees join the company, the company will arrange a one-week induction training, including company culture, business processes, and job skills training.
[0135] [2]: "Employee Handbook of XX Company" (Revised and implemented on July 1, 2024): Chapter 5 Dress Code: Article 1: The company implements business casual dress. From Monday to Friday, employees need to wear neat and appropriate clothing, avoiding overly casual or revealing outfits.
[0136] In this embodiment, the encoding can be associated with the knowledge block and stored in the knowledge base. Each knowledge block corresponds to a unique encoding, and this unique encoding serves as the unique identifier for the corresponding knowledge block.
[0137] After obtaining the first candidate knowledge block, the knowledge base can be queried to find the entry that matches the first candidate knowledge block and obtain the corresponding encoding. Based on the encoding corresponding to the first candidate knowledge block, a prompt can be constructed. The prompt can be a structured text, including the user's question, the content of the first candidate knowledge block, and the encoding of the first candidate knowledge block. Through the prompt, a context environment can be provided for the large language model to help the large language model understand the relationship between the user's question and the first candidate knowledge block, and guide the large language model to output the encoding instead of the original text of the knowledge block. The construction of the prompt is as follows, where "{0}" represents the knowledge block and the corresponding encoding, and "{1}" represents the user's question:
[0138] partyPrompt = r
[0139] {0}
[0140] Answer the question based on the above information: "{1}"
[0141] Output requirements:
[0142] 1. Each sentence in the answer needs to be supported by the information.
[0143] 2. Do not output the original text of the information. Directly give the serial numbers [*], such as [1], [3], [2], etc.;
[0144] 3. For information irrelevant to the question, do not use it directly;
[0145] 4. If all the information is irrelevant to the question, directly output Unable to answer;
[0146] 5. There are two examples below. Please output strictly according to the following example format.
[0147] Output Example 1:
[0148] Working hours regulations: According to the "Employee Handbook of XX Company", the company's standard working hours are from Monday to Friday, from 9:00 am to 6:00 pm, and the lunch break is from 12:00 noon to 1:00 pm [1].
[0149] Attendance System: The company adopts an electronic attendance system, and employees are required to punch in and out on time. Late arrivals, early departures, or absenteeism will be dealt with accordingly in accordance with the company's attendance policy [3].
[0150] Leave and Vacation: Employees need to apply for leave in advance through the company's leave system and obtain approval from their direct supervisor. The company provides statutory holidays such as paid annual leave, sick leave, marriage leave, maternity leave, etc., and the specific number of days is implemented according to national laws and regulations and company policies [3][6].
[0151] Output Example 2:
[0152] According to the provisions of the "Employee Handbook of XX Company", employees' salaries include basic salary, performance bonuses, and year-end bonuses. The company determines the salary standards based on market levels and the value of employees' positions [2].
[0153] The company implements a quarterly performance evaluation system for employees. The evaluation results will affect employees' salary adjustments and promotion opportunities. Performance evaluations include multiple dimensions such as work results, work attitudes, and teamwork [2][5].
[0154] After the prompt information is constructed, the prompt information and the first candidate knowledge block can be input into the large language model together. Based on the guidance of the prompt information, the large language model analyzes the correlation between the user's question and the first candidate knowledge block, and filters out the most relevant target knowledge block from the first candidate knowledge block. After determining the target knowledge block, the large language model outputs the target code corresponding to the target knowledge block, rather than the target knowledge block itself. After obtaining the target code, if the target code is the first code, the target knowledge block corresponding to the target code can be matched from the knowledge base to obtain the original information of the relevant target file.
[0155] In this embodiment, a unique identifier is assigned to each knowledge block through the encoding mechanism, and the prompt information guides the large language model to output the code instead of directly outputting the original text of the target file, reducing the problem of inappropriate addition, deletion, and modification of the original text of the target file caused by hallucinations that may occur in the large language model. Through the code output by the large language model, the original text corresponding to the code can be further matched from the knowledge block, improving the standardization of the content output by the retrieval results.
[0156] In some embodiments, inputting the first candidate knowledge block into a preset large language model includes:
[0157] Dividing the first candidate knowledge block into multiple knowledge block groups based on the number of concurrent inference paths of the large language model;
[0158] Inputting the multiple knowledge block groups into each inference path of the large language model respectively to obtain second candidate knowledge blocks screened by each inference path from each knowledge block group;
[0159] Input the second candidate knowledge block into the large language model; among them, the large language model screens out the target knowledge block from the second candidate knowledge blocks according to the degree of association between the question and the second candidate knowledge blocks.
[0160] In this embodiment, parallel processing refers to the ability to execute multiple computing tasks simultaneously. Large language models, such as BERT or GPT series based on the Transformer architecture, are usually designed to be able to process multiple inputs in parallel to improve computing efficiency and processing speed. Each inference path can be regarded as an independent processing flow inside the large language model. In these inference paths, the large language model can independently analyze and process the input data to obtain their respective output results.
[0161] In this embodiment, the first candidate knowledge blocks can be divided into multiple knowledge block groups according to the number of inference paths. For example, if the number of the first candidate knowledge blocks is m and the number of inference paths is n, then the number of the first candidate knowledge blocks in each knowledge block group is m / n. Each knowledge block group can be independently input into an inference path of the large language model, thereby improving the overall processing speed.
[0162] Each inference path can output a set of second candidate knowledge blocks. The second candidate knowledge blocks can be one knowledge block or multiple knowledge blocks. The second candidate knowledge blocks output by each inference path can be aggregated and input into the large language model again, and the target knowledge block is screened out from the second candidate knowledge blocks.
[0163] In some embodiments, by querying the knowledge base, an entry matching the second candidate knowledge block can be found and the corresponding encoding can be obtained. According to the encoding corresponding to the second candidate knowledge block, a prompt can be constructed. The prompt can be a structured text, including the user's question, the content of the second candidate knowledge block, and the encoding of the first candidate knowledge block. Through the prompt, a context environment can be provided for the large language model to help the large language model understand the relationship between the user's question and the second candidate knowledge block, and guide the large language model to output the encoding instead of the original text of the knowledge block.
[0164] After the prompt is constructed, the prompt and the second candidate knowledge block can be input into the large language model together. Based on the guidance of the prompt, the large language model analyzes the degree of association between the user's question and the second candidate knowledge block, and screens out the most relevant target knowledge block from the second candidate knowledge blocks. When the target knowledge block is determined, the large language model outputs the target encoding corresponding to the target knowledge block, rather than the target knowledge block itself.
[0165] In this embodiment, by dividing the first candidate knowledge block into multiple knowledge block groups and using the concurrent inference paths of the large language model for multi-stage screening, the parallel processing ability of the model is fully utilized, and the efficiency bottleneck problem caused by processing a large number of knowledge blocks at one time is reduced. This not only alleviates the burden on a single inference path but also can significantly shorten the retrieval time through parallel computing, greatly improving the retrieval efficiency.
[0166] In some embodiments, inputting the second candidate knowledge block into the large language model includes:
[0167] Constructing a second encoding for the second candidate knowledge block; wherein, the second encoding is the unique identifier of the second candidate knowledge block;
[0168] Constructing a second prompt message according to the second encoding corresponding to the second candidate knowledge block;
[0169] Inputting the second prompt message and the second candidate knowledge block into the large language model; wherein, based on the guidance of the second prompt message, the large language model screens out the target knowledge block from the second candidate knowledge blocks according to the degree of association between the question and the second candidate knowledge blocks, and outputs the target encoding corresponding to the target knowledge block.
[0170] In this embodiment, by querying the knowledge base, the entries matching the second candidate knowledge block can be found, and the corresponding first encoding can be obtained. However, since each inference path can output a set of second candidate knowledge blocks, the first encodings corresponding to multiple sets of second candidate knowledge blocks may not be continuous. For example, the first encodings corresponding to the second candidate knowledge blocks output by one inference path are [2][3][6], the first encodings corresponding to the second candidate knowledge blocks output by one inference path are [2][8], and the first encodings corresponding to the second candidate knowledge blocks output by one inference path are [3][8][9]. After integrating the second candidate knowledge blocks output by each inference path, the corresponding first encodings are [2][3][6][8][9]. These encodings are not continuous. If the prompt message is constructed based on the first encoding and input into the large language model, according to the inference characteristics of the large language model, the large language model may think that there is a typo in the input encoding and will adjust the order of the input encoding, resulting in inaccurate results output by the large language model.
[0171] To solve the above technical problems, in some embodiments, the second candidate knowledge blocks output by each inference path can be integrated and re-encoded. For example, a second encoding can be constructed for each second candidate module, and each second encoding is the unique identifier of the corresponding second candidate knowledge block, and the corresponding relationship between the second candidate knowledge block and the second encoding is stored in the cache database.
[0172] After constructing the second encoding, the second prompt information can be constructed based on the second encoding. Among them, the construction method of the second prompt information can refer to the construction method of the first prompt information, which will not be elaborated one by one in the embodiments of the present application.
[0173] After the second prompt information is constructed, the second prompt information and the second candidate knowledge blocks can be input into the large language model together. Based on the guidance of the second prompt information, the large language model analyzes the correlation between the user's question and the second candidate knowledge blocks, and screens out the most relevant target knowledge blocks from the second candidate knowledge blocks. After determining the target knowledge blocks, the large language model outputs the target encoding corresponding to the target knowledge blocks, rather than the target knowledge blocks themselves. After obtaining the target encoding, if the target encoding is the second encoding, the target knowledge block corresponding to the target encoding can be matched from the cache database to obtain the original text information of the relevant target file.
[0174] In this embodiment, by re-encoding the second candidate knowledge blocks output by multiple inference paths, the continuity of the encoding can be improved, making the input of the large language model more in line with the inference habits of the large language model and improving the accuracy of the retrieval results.
[0175] The following introduces the retrieval method of the document materials in the embodiments of the present application through a scenario example. As Figure 3 shown,
[0176] The editing module is used to structurally process the basic information of the target file to obtain structured meta-information; divide the file content according to the logical structure of the file content in the target file and associate it with the meta-information to obtain multiple knowledge blocks.
[0177] Through the dual tower model, the knowledge blocks can be vectorized, and the feature vectors of the knowledge blocks are stored in the vector database. According to the question input by the user, the relevant knowledge blocks, that is, the first candidate knowledge blocks, can be quickly recalled.
[0178] In the first stage, through the dual tower model, the first candidate knowledge blocks can be recalled from the vector database. If the number of the recalled first candidate knowledge blocks is m and the number of inference paths of the large language model is n, then the number of knowledge blocks processed by each inference path is m / n, so as to realize the n-way concurrent inference of the large language model. Based on the guidance of the prompt information, the first encoding of the second candidate knowledge blocks output by each inference path is obtained, and the knowledge blocks corresponding to the encoding can be matched from the knowledge base through these first encodings. Then, the second candidate knowledge blocks are re-encoded, that is, the second encoding is constructed, and the corresponding relationship between the second candidate knowledge blocks and the second encoding is stored in the cache database.
[0179] In the second stage, the prompt information can be reconstructed according to the second encoding of the second candidate knowledge block. Since the number of the second candidate knowledge blocks has been screened and has been greatly reduced compared to the number of the first candidate knowledge blocks, a single inference path can be used for inference, and the corresponding knowledge blocks can be retrieved from the knowledge base according to the encoding output by the large language model.
[0180] Through the above process, the retrieved results are highly relevant to the user's question and the output speed is relatively fast. In addition, the output results also carry structured meta-information, which is convenient for tracing back to the source file, improving the accuracy and reliability of the retrieval structure.
[0181] The embodiment of the present application also provides a retrieval method for document materials, as Figure 4 shown, the retrieval method for document materials includes: step 410, step 420, step 430 and step 440.
[0182] Step 410: Input the question entered by the user into a preset dual-tower model; wherein, the dual-tower model is used to extract the first feature vector of the question, extract the second feature vectors of multiple knowledge blocks from a preset vector database, and screen the first candidate knowledge blocks from the multiple knowledge blocks according to the similarity between the first feature vector and the second feature vectors; the multiple knowledge blocks are obtained by dividing the document content according to the logical structure of the document content in the target document and associating with meta-information, and the meta-information is obtained by structuring the basic information of the target document;
[0183] Step 420: Extract the first encoding corresponding to the first candidate knowledge block from a preset knowledge base; wherein, the knowledge base stores multiple knowledge blocks and multiple first encodings corresponding to the multiple knowledge blocks, and each knowledge block corresponds to a first encoding as a unique identifier;
[0184] Step 430: Construct the first prompt information according to the first encoding corresponding to the first candidate knowledge block;
[0185] Step 440: Input the first prompt information and the first candidate knowledge block into a preset large language model; wherein, based on the guidance of the first prompt information, the large language model screens out the target knowledge block from the first candidate knowledge blocks according to the correlation degree between the question and the first candidate knowledge blocks, and outputs the target encoding corresponding to the target knowledge block.
[0186] The retrieval method for document materials provided by the embodiments of the present application obtains structured meta-information by performing structured processing on the basic information of the target document; divides the document content according to the logical structure of the document content in the target document and associates it with the meta-information to obtain multiple knowledge blocks; and retrieves the target knowledge block associated with the question from the multiple knowledge blocks according to the similarity between the question input by the user and the multiple knowledge blocks. Based on the characteristics of the target document, the embodiments of the present application perform structured processing on the basic information of the target document, extract meta-information, and divide the document content according to the logical structure of the document content to obtain multiple knowledge blocks, so that the knowledge blocks contain rich semantic information, enhance the semantic understanding ability in the retrieval process, and thus enable the retrieval process to more accurately grasp the relevance between the user's intention and the semantics of the knowledge blocks, improving the accuracy of the retrieval results.
[0187] In some embodiments, inputting the first prompt information and the first candidate knowledge block into a preset large language model includes:
[0188] Dividing the first candidate knowledge block into multiple knowledge block groups based on the number of concurrent inference paths of the large language model;
[0189] Inputting the multiple knowledge block groups into each inference path of the large language model respectively to obtain second candidate knowledge blocks screened by each inference path from each knowledge block group;
[0190] Constructing a second encoding for the second candidate knowledge block; wherein, the second encoding is the unique identifier of the second candidate knowledge block;
[0191] Constructing second prompt information according to the second encoding corresponding to the second candidate knowledge block;
[0192] Inputting the second prompt information and the second candidate knowledge block into the large language model; wherein, based on the guidance of the second prompt information, the large language model screens out the target knowledge block from the second candidate knowledge blocks according to the degree of association between the question and the second candidate knowledge blocks, and outputs the target encoding corresponding to the target knowledge block.
[0193] The retrieval method for document materials provided by the embodiments of the present application may have a retrieval device for document materials as the execution subject. In the embodiments of the present application, taking the retrieval device for document materials to execute the retrieval method for document materials as an example, the retrieval device for document materials provided by the embodiments of the present application is described.
[0194] The embodiments of the present application further provide a retrieval device for document materials.
[0195] As Figure 5 shown, the retrieval device for document materials includes:
[0196] A processing module 510, configured to perform structured processing on the basic information of the target document to obtain structured meta-information;
[0197] A partitioning module 520, configured to partition the file content according to the logical structure of the file content in the target file, and associate it with the meta information to obtain a plurality of knowledge chunks;
[0198] A retrieval module 530, configured to retrieve a target knowledge chunk associated with the question from the plurality of knowledge chunks according to the similarity between the question input by the user and the plurality of knowledge chunks.
[0199] The retrieval device for document materials provided by the embodiments of the present application obtains structured meta information by performing structured processing on the basic information of the target document; partitions the file content according to the logical structure of the file content in the target document, and associates it with the meta information to obtain a plurality of knowledge chunks; and retrieves a target knowledge chunk associated with the question from the plurality of knowledge chunks according to the similarity between the question input by the user and the plurality of knowledge chunks. Based on the characteristics of the target document, the embodiments of the present application perform structured processing on the basic information of the target document, extract meta information, and partition the file content in combination with the logical structure of the file content to obtain a plurality of knowledge chunks, so that the knowledge chunks contain rich semantic information, enhance the semantic understanding ability in the retrieval process, and thus enable the retrieval process to more accurately grasp the relevance between the user's intention and the semantics of the knowledge chunks, improving the accuracy of the retrieval results.
[0200] In some embodiments, the retrieval module 530 is further configured to:
[0201] Obtain a first feature vector of the question and second feature vectors of the plurality of knowledge chunks respectively;
[0202] Calculate the similarity between the first feature vector and the second feature vectors;
[0203] Determine the knowledge chunk corresponding to the second feature vector whose similarity to the first feature vector is greater than a preset threshold as the target knowledge chunk.
[0204] In some embodiments, the retrieval module 530 is further configured to:
[0205] Input the question and the plurality of knowledge chunks into a preset two-tower model; wherein, the two-tower model is configured to extract a first feature vector of the question and second feature vectors of the plurality of knowledge chunks respectively, calculate the similarity between the first feature vector and the second feature vectors, and output the target knowledge chunk based on the similarity.
[0206] In some embodiments, the retrieval module 530 is further configured to:
[0207] Store the second feature vectors of the plurality of knowledge chunks extracted by the two-tower model into a preset vector database.
[0208] In some embodiments, the retrieval module 530 is further configured to:
[0209] Select the first candidate knowledge block from multiple knowledge blocks according to the similarity between the question input by the user and the multiple knowledge blocks;
[0210] Input the first candidate knowledge block into a preset large language model; wherein, the large language model screens out the target knowledge block from the first candidate knowledge blocks according to the degree of association between the question and the first candidate knowledge blocks.
[0211] In some embodiments, the retrieval module 530 is further configured to:
[0212] Extract the first encoding corresponding to the first candidate knowledge block from a preset knowledge base; wherein, multiple knowledge blocks and multiple first encodings corresponding to the multiple knowledge blocks are stored in the knowledge base, and each knowledge block corresponds to a first encoding as a unique identifier;
[0213] Construct the first prompt message according to the first encoding corresponding to the first candidate knowledge block;
[0214] Input the first prompt message and the first candidate knowledge block into the large language model; wherein, the large language model, guided by the first prompt message, screens out the target knowledge block from the first candidate knowledge blocks according to the degree of association between the question and the first candidate knowledge blocks, and outputs the target encoding corresponding to the target knowledge block.
[0215] In some embodiments, the retrieval module 530 is further configured to:
[0216] Divide the first candidate knowledge blocks into multiple knowledge block groups based on the number of concurrent inference paths of the large language model;
[0217] Input the multiple knowledge block groups into each inference path of the large language model respectively, and obtain the second candidate knowledge blocks screened by each inference path from each knowledge block group;
[0218] Input the second candidate knowledge blocks into the large language model; wherein, the large language model screens out the target knowledge block from the second candidate knowledge blocks according to the degree of association between the question and the second candidate knowledge blocks.
[0219] In some embodiments, the retrieval module 530 is further configured to:
[0220] Construct a second encoding for the second candidate knowledge blocks; wherein, the second encoding is the unique identifier of the second candidate knowledge blocks;
[0221] Construct the second prompt message according to the second encoding corresponding to the second candidate knowledge blocks;
[0222] Input the second prompt message and the second candidate knowledge block into the large language model; wherein, based on the guidance of the second prompt message, the large language model screens out the target knowledge block from the second candidate knowledge blocks according to the relevance between the question and the second candidate knowledge blocks, and outputs the target encoding corresponding to the target knowledge block.
[0223] The retrieval device for document materials in the embodiments of the present application can be an electronic device or a component in the electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices other than the terminal. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and can also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.
[0224] The retrieval device for document materials in the embodiments of the present application can be a device with an operating system. The operating system can be a Microsoft (Windows) operating system, an Android operating system, an IOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.
[0225] In some embodiments, as Figure 6 shown, the embodiments of the present application further provide an electronic device 600, including a processor 601, a memory 602, and a computer program stored on the memory 602 and executable on the processor 601. When the program is executed by the processor 601, it implements each process of the above-mentioned embodiment of the retrieval method for document materials and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0226] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.
[0227] An embodiment of the present application further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above-described embodiment of the method for retrieving document materials and can achieve the same technical effects. To avoid repetition, details are not described herein again.
[0228] Wherein, the processor is the processor in the electronic device in the above embodiment. The readable storage medium includes computer-readable storage media such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disc.
[0229] An embodiment of the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the above-described method for retrieving document materials.
[0230] Wherein, the processor is the processor in the electronic device in the above embodiment. The readable storage medium includes computer-readable storage media such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disc.
[0231] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run programs or instructions to implement each process of the above-described embodiment of the method for retrieving document materials and can achieve the same technical effects. To avoid repetition, details are not described herein again.
[0232] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip.
[0233] It should be noted that in this document, the terms "include", "comprise", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device including a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including that element. In addition, it should be pointed out that the methods and devices in the embodiments of the present application are not limited to performing functions in the order shown or discussed. They may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described method may be executed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0234] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application.
[0235] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
[0236] In the description of this specification, the description of reference terms such as "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0237] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and purposes of the present application. The scope of the present application is defined by the claims and their equivalents.
Claims
1. A retrieval method for document materials, characterized in that, Including: Structurally process the basic information of the target file to obtain structured meta-information; Divide the file content according to the logical structure of the file content in the target file and associate it with the meta-information to obtain multiple knowledge chunks; Retrieve the target knowledge chunk associated with the question from multiple knowledge chunks according to the similarity between the question input by the user and the multiple knowledge chunks.
2. The method according to claim 1, wherein The retrieving the target knowledge chunk associated with the question from multiple knowledge chunks according to the similarity between the question input by the user and the multiple knowledge chunks includes: Obtain the first feature vector of the question and the second feature vectors of the multiple knowledge chunks respectively; Calculate the similarity between the first feature vector and the second feature vectors; Determine the knowledge chunk corresponding to the second feature vector with a similarity greater than a preset threshold to the first feature vector as the target knowledge chunk.
3. The method according to claim 1, characterized in that The retrieving the target knowledge chunk associated with the question from multiple knowledge chunks according to the similarity between the question input by the user and the multiple knowledge chunks includes: Input the question and the multiple knowledge chunks into a preset dual-tower model; wherein, the dual-tower model is used to extract the first feature vector of the question and the second feature vectors of the multiple knowledge chunks respectively, calculate the similarity between the first feature vector and the second feature vectors, and output the target knowledge chunk based on the similarity.
4. The method according to claim 3, characterized in that The method further includes: Store the second feature vectors of the multiple knowledge chunks extracted by the dual-tower model into a preset vector database.
5. The method according to claim 1, characterized in that, The retrieving the target knowledge chunk associated with the question from multiple knowledge chunks according to the similarity between the question input by the user and the multiple knowledge chunks includes: Screen and obtain the first candidate knowledge chunks from multiple knowledge chunks according to the similarity between the question input by the user and the multiple knowledge chunks; Input the first candidate knowledge chunks into a preset large language model; wherein, the large language model screens out the target knowledge chunks from the first candidate knowledge chunks according to the degree of association between the question and the first candidate knowledge chunks.
6. The method according to claim 5, wherein The inputting the candidate knowledge chunks into a preset large language model includes: Extract the first encoding corresponding to the first candidate knowledge chunk from a preset knowledge base; wherein, the knowledge base stores multiple knowledge chunks and multiple first encodings corresponding to the multiple knowledge chunks, and each knowledge chunk corresponds to a first encoding as a unique identifier; Construct the first prompt information according to the first encoding corresponding to the first candidate knowledge chunk; Input the first prompt information and the first candidate knowledge chunks into the large language model; wherein, the large language model, guided by the first prompt information, screens out the target knowledge chunks from the first candidate knowledge chunks according to the degree of association between the question and the first candidate knowledge chunks, and outputs the target encoding corresponding to the target knowledge chunks.
7. The method according to claim 5, wherein The inputting the first candidate knowledge chunks into a preset large language model includes: Divide the first candidate knowledge chunks into multiple knowledge chunk groups based on the number of concurrent inference paths of the large language model; Input each of the multiple groups of knowledge chunks into each inference path of the large language model to obtain second candidate knowledge chunks screened by each inference path from each group of knowledge chunks; Input the second candidate knowledge chunks into the large language model; wherein, the large language model screens out target knowledge chunks from the second candidate knowledge chunks according to the degree of association between the question and the second candidate knowledge chunks.
8. The method according to claim 7, wherein The step of inputting the second candidate knowledge chunks into the large language model includes: Construct a second encoding for the second candidate knowledge chunks; wherein, the second encoding is the unique identifier of the second candidate knowledge chunks; Construct second prompt information according to the second encoding corresponding to the second candidate knowledge chunks; Input the second prompt information and the second candidate knowledge chunks into the large language model; wherein, the large language model, guided by the second prompt information, screens out target knowledge chunks from the second candidate knowledge chunks according to the degree of association between the question and the second candidate knowledge chunks, and outputs the target encoding corresponding to the target knowledge chunks.
9. A retrieval method for document materials, characterized in that, It includes: Input the question input by the user into a preset two-tower model; wherein, the two-tower model is used to extract the first feature vector of the question, extract the second feature vectors of multiple knowledge chunks from a preset vector database, and screen out first candidate knowledge chunks from the multiple knowledge chunks according to the similarity between the first feature vector and the second feature vectors; the multiple knowledge chunks are obtained by dividing the file content according to the logical structure of the file content in the target file and associating with meta-information, and the meta-information is obtained by structuring the basic information of the target file; Extract the first encoding corresponding to the first candidate knowledge chunks from a preset knowledge base; wherein, the knowledge base stores multiple knowledge chunks and multiple first encodings corresponding to the multiple knowledge chunks, and each knowledge chunk corresponds to a first encoding as a unique identifier; Construct first prompt information according to the first encoding corresponding to the first candidate knowledge chunks; Input the first prompt information and the first candidate knowledge chunks into a preset large language model; wherein, the large language model, guided by the first prompt information, screens out target knowledge chunks from the first candidate knowledge chunks according to the degree of association between the question and the first candidate knowledge chunks, and outputs the target encoding corresponding to the target knowledge chunks.
10. The method according to claim 9, wherein The step of inputting the first prompt information and the first candidate knowledge chunks into a preset large language model includes: Divide the first candidate knowledge chunks into multiple groups of knowledge chunks based on the number of concurrent inference paths of the large language model; Input each of the multiple groups of knowledge chunks into each inference path of the large language model to obtain second candidate knowledge chunks screened by each inference path from each group of knowledge chunks; Construct a second encoding for the second candidate knowledge chunks; wherein, the second encoding is the unique identifier of the second candidate knowledge chunks; Construct second prompt information according to the second encoding corresponding to the second candidate knowledge chunks; Input the second prompt information and the second candidate knowledge block into the large language model; wherein, based on the guidance of the second prompt information, the large language model screens out the target knowledge block from the second candidate knowledge blocks according to the degree of association between the question and the second candidate knowledge blocks, and outputs the target encoding corresponding to the target knowledge block.
11. A retrieval device for document materials, characterized in that, Including: A processing module, configured to perform structured processing on the basic information of the target file to obtain structured meta-information; A partitioning module, configured to partition the file content according to the logical structure of the file content in the target file and associate it with the meta-information to obtain multiple knowledge blocks; A retrieval module, configured to retrieve the target knowledge block associated with the question from multiple knowledge blocks according to the similarity between the question input by the user and the multiple knowledge blocks.
12. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the method described in any one of claims 1-10 is implemented.
13. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1-10 is implemented.