Large model retrieval question and answer method and device based on document structure tree
By adopting a large-scale model retrieval method based on document structure tree in the intelligent customer service system, the efficiency and cost problems of existing systems when dealing with complex business processes and diversified problems are solved, and accurate, efficient and low-cost question responses are achieved.
Patent Information
- Application Number
- CN202510050486.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-01-13
AI Technical Summary
When dealing with complex business processes and diversified problems, existing intelligent customer service question-and-answer issues such as missing content, wrong sorting, integration errors, and answering non-questions. The graph construction cost is high, making it difficult to deal with complex processes and cannot achieve targeted rhetorical questions.
The big model search question and answer method based on the document structure tree is adopted, and the document content is organized according to hierarchical relationships by building a business tree; the vector model is used to filter the candidate node collection, and hierarchically match it through the large language model to generate answers.
It realizes quick responses in the user's known business scenarios, and supports heuristic questions and answers to users who do not understand the business process, providing accurate, efficient and low-cost question answers.
Smart Images

Figure CN120104727A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of document intelligence, and in particular relates to a large model retrieval question-answering method and device based on a document structure tree. Background Art
[0002] As an intelligent system that can handle various common user questions online, the customer service question and answer system can greatly reduce the workload of manual customer service and improve service efficiency. It has always had great application value in scenarios such as public services and consumer consultation.
[0003] With the development of NLP technology, intelligent customer service systems are also constantly evolving. Early intelligent customer service question-and-answer systems were mainly based on FAQs based on common question-and-answer pairs and KBQA based on knowledge graphs. In essence, they both answered user questions by searching the local knowledge base. However, the scope of their answers is limited to the content that is clearly organized in the knowledge base. The quality of the knowledge base construction directly determines the effect of the question-and-answer system, and it often requires a lot of manpower and material resources to continuously polish the knowledge base. At the same time, due to the limited problem parsing capabilities of small models, mismatches or failures to recall often occur.
[0004] The retrieval enhancement system based on the large language model has brought new opportunities to the intelligent question-answering system. On the one hand, the effect of using the large language model to understand user questions is far better than that of the traditional small model, and it can be truly realized at the semantic level, greatly improving the recall rate of the question-answering system; on the other hand, the paragraphs, chapters and even documents related to the answer can be directly handed over to the large language model for understanding and answering, reducing the process of manually constructing fine-grained knowledge bases in early question-answering systems. The GraphRAG system, which combines the knowledge graph with the large language model, further performs graph analysis on the document objects to be understood, improving the quality of the system output content.
[0005] However, basic RAG systems often have problems such as missing content, incorrect sorting, incorrect integration, irrelevant answers, etc., and cannot accurately answer questions. GraghRAG systems also still have problems such as high cost of graph construction, difficulty in handling complex processes, and inability to implement targeted counter-questions.
[0006] Therefore, for online customer service scenarios, especially those with diverse business categories and complicated processes, there is an urgent need for an accurate and low-cost method for answering questions. Summary of the invention
[0007] The present invention is made to solve the above-mentioned problems, and its purpose is to provide a large model retrieval question answering method and device based on a document structure tree.
[0008] The invention provides a large model retrieval question-answering method based on a document structure tree, which is used for generating answers corresponding to user questions according to multiple documents, and has the following characteristics, including the following steps: step S1, respectively parsing each document to obtain corresponding business trees respectively; step S2, for each business tree, converting each complete path from the root node to the leaf node in the business tree into a corresponding path vector; step S3, according to the user question and all the path vectors, combined with the vector model, obtaining a candidate node set; step S4, judging whether to generate an answer according to the candidate node set, if so, generating an answer according to the large language model and the candidate node set, if not, executing step S5; step S5, according to the user question, hierarchically matching all business trees through the large language model to obtain a hierarchical matching result; step S6, generating an answer through the large language model combined with the hierarchical matching result, wherein the root node of the business tree is the main title of the corresponding document, and each leaf node of the business tree is each specific content description in the document, and each title of the document is used as a node of the business tree according to the hierarchical relationship, and when the title is used as the parent node, each subtitle corresponding to the title is used as each child node corresponding to the parent node.
[0009] In the large model retrieval question and answer method based on the document structure tree provided by the present invention, it can also have the following characteristics: wherein, step S3 includes the following sub-steps: step S3-1, vectorizing the user question to obtain the user question vector; step S3-2, calculating the similarity between the user question vector and each path vector through the vector model; step S3-3, screening the leaf nodes corresponding to the path vector corresponding to the similarity greater than the preset similarity threshold, and taking all the screened leaf nodes as the candidate node set.
[0010] In the large model retrieval question and answer method based on the document structure tree provided by the present invention, it can also have the following characteristics: wherein, step S4 includes the following sub-steps: step S4-1, judging whether the candidate node set meets the preset judgment condition, if so, executing step S4-2, if not, executing step S5; step S4-2, for each leaf node in the candidate node set, combining the leaf node with the user question and inputting the large language model to obtain the confidence corresponding to the leaf node; step S4-3, selecting the leaf node corresponding to the highest confidence and combining it with the polishing prompt word and inputting the large language model to obtain the answer, and the preset judgment condition includes at least one preset condition that the total number of leaf nodes in the candidate node set is not greater than the preset number of nodes and the candidate node set is not an empty set.
[0011] In the large model retrieval question and answer method based on the document structure tree provided by the present invention, it can also have the following characteristics: wherein, step S5 includes the following sub-steps: step S5-1, using the large language model to determine in turn whether each root node is relevant to the user question, and taking the root node with relevance as the relevant node; step S5-2, using the large language model to determine in turn whether each child node corresponding to each relevant node is relevant to the user question, and taking the child node with relevance as the relevant node; step S5-3, repeating step S5-2 until a hierarchical matching result is obtained, and the hierarchical matching result includes: a single leaf node is relevant; all child nodes of a single parent node are relevant; all nodes are not relevant.
[0012] The large model retrieval question-answering method based on the document structure tree provided by the present invention may also have the following feature: when the hierarchical matching result is that a single leaf node has relevance, the leaf node is combined with the polishing prompt word and input into the large language model to obtain the answer.
[0013] The large model retrieval question-answering method based on the document structure tree provided by the present invention may also have the following feature: wherein the neighbor nodes of the leaf nodes are combined with the recommended prompt words and input into the large language model to generate recommended information for the user's question.
[0014] In the large model retrieval question and answer method based on the document structure tree provided by the present invention, it can also have the following characteristics: wherein, when the hierarchical matching result is that all child nodes of a single parent node are relevant, the large language model is used to determine whether the user question corresponds to a single child node or to multiple child nodes. When the user question corresponds to a single child node, all child nodes are combined with rhetorical question prompt words and input into the large language model, rhetorical question information is generated and displayed to the user as a target node to obtain the user's answer, all leaf nodes corresponding to the target node are combined with polishing prompt words and input into the large language model to obtain the answer. When the user question corresponds to multiple child nodes, all leaf nodes corresponding to all child nodes are combined with polishing prompt words and input into the large language model to obtain the answer.
[0015] The large-model retrieval question-answering method based on the document structure tree provided by the present invention may also have the following feature: when the hierarchical matching result is that all nodes are not relevant, the answer is a preset reply content.
[0016] The large model retrieval question and answer method based on the document structure tree provided by the present invention may also have the following features: wherein, in step S5-2, when a single intermediate node has relevance, each child node corresponding to the intermediate node is combined with a rhetorical question prompt word and input into the large language model, and rhetorical question information is generated and displayed to the user. The large language model is used to determine in turn whether each child node corresponding to the intermediate node has relevance to the user's reply to the rhetorical question information, and the child nodes with relevance are used as relevant nodes.
[0017] The present invention also provides a large-model retrieval question-answering device based on a document structure tree, which is used to generate answers corresponding to user questions based on multiple documents, and has the following characteristics: a business tree construction module, a path vector generation module, a candidate node screening module, a judgment module, a hierarchical matching module and an answer generation module, wherein the business tree construction module is used to parse each document separately to obtain the corresponding business tree respectively, the path vector generation module is used to convert each complete path from the root node to the leaf node in the business tree into a corresponding path vector for each business tree, the candidate node screening module contains a vector model for obtaining a candidate node set based on the user question and all path vectors in combination with the vector model, and the judgment module is used to obtain a candidate node set based on the user question and all path vectors. The diagnosis module is used to determine whether to generate an answer based on the candidate node set. If so, the answer is generated based on the large language model and the candidate node set. If not, the hierarchical matching module is executed. The hierarchical matching module is used to hierarchically match all business trees through the large language model according to the user question to obtain the hierarchical matching result. The answer generation module is used to generate the answer through the large language model combined with the hierarchical matching result. The root node of the business tree is the main title of the corresponding document. The leaf nodes of the business tree are the specific content descriptions in the document. According to the hierarchical relationship, the titles of the document are used as nodes of the business tree. When the title is used as the parent node, the sub-titles corresponding to the title are used as the child nodes corresponding to the parent node.
[0018] Functions and Effects of the Invention
[0019] According to the large model retrieval question-answering method and device based on the document structure tree involved in the present invention, on the one hand, by constructing a business tree that conforms to the original logical order of the document and has no loops, the business corresponding to the document is guided; on the other hand, by screening the candidate node set and hierarchical matching, a quick reply is achieved in the business scenario where the user knows what he is doing, and heuristic question-answering is also supported in the scenario where the user does not understand the actual business process. Therefore, the large model retrieval question-answering method and device based on the document structure tree of the present invention can achieve accurate, efficient and low-cost question answers. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a block diagram of a large model retrieval question-answering device in an embodiment of the present invention;
[0021] Figure 2 is a schematic diagram of a service tree in an embodiment of the present invention;
[0022] Figure 3 is a schematic diagram of a process of generating a candidate node set by a candidate node screening module in an embodiment of the present invention;
[0023] Figure 4 Schematic diagram of the working process of the judgment module in an embodiment of the present invention;
[0024] Figure 5 is a schematic diagram of a process of generating a hierarchical matching result by a hierarchical matching module in an embodiment of the present invention;
[0025] Figure 6 It is a flowchart of a large model retrieval question-answering method based on a document structure tree in an embodiment of the present invention. DETAILED DESCRIPTION
[0026] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the following embodiments and the accompanying drawings specifically illustrate the large model retrieval question-answering method and device based on the document structure tree of the present invention.
[0027] This embodiment provides a large model retrieval question and answer device based on a document structure tree, hereinafter referred to as a large model retrieval question and answer device, which is used to generate answers corresponding to user questions based on multiple documents.
[0028] Figure 1 It is a block diagram of a large model retrieval question and answer device in an embodiment of the present invention.
[0029] like Figure 1 As shown, the large model retrieval question and answer device 100 includes a business tree construction module 11, a path vector generation module 12, a candidate node screening module 13, a judgment module 14, a hierarchical matching module 15, an answer generation module 16 and a general control module 17 for controlling the operation of the above modules.
[0030] The service tree construction module 11 is used to parse each document to obtain the corresponding service tree. In this embodiment, all service trees constitute a service forest.
[0031] The root node of the business tree is the main title of the corresponding document. The leaf nodes of the business tree are the specific content descriptions in the document. According to the hierarchical relationship, the titles of the document are used as nodes of the business tree. When the title is used as the parent node, the sub-titles corresponding to the title are used as the child nodes corresponding to the parent node.
[0032] In this embodiment, the business tree construction module 11 first analyzes the directory structure of the document and extracts the main title, title and hierarchical relationship therein. Then, the business tree construction module 11 inserts the child node into the corresponding parent node according to the chapter level. If a title level is found to be deeper than the current title, it means that it is a sub-chapter of the current chapter and needs to be inserted into the current node; if the level becomes shallow, it should be inserted into the node of the previous level. Finally, the business tree construction module 11 stores the entire business tree in json format.
[0033] Furthermore, in this embodiment, each fragmented common question that is not within the scope of the business tree content is used as a root node, and the corresponding preset answer is used as a leaf node corresponding to the root node, thereby forming a business tree for each common question.
[0034] Figure 2 It is a schematic diagram of a service tree in an embodiment of the present invention.
[0035] like Figure 2 As shown, each box contains the content of the corresponding node, and the parent node and its corresponding child node are connected by arrows. The main title of the document is "Pension Insurance", and its corresponding child nodes include "Employee Pension Insurance", "Pension Insurance for Rural-to-Non-agricultural Personnel", etc. The child nodes corresponding to "Employee Pension Insurance" include "Participation and Payment", "Account Management", etc. The child nodes corresponding to "Participation and Payment" include "Individual", "Unit", etc. The child nodes corresponding to "Individual" include "Individual", "Individual Business Owner", etc. The child nodes corresponding to "Individual" include "Scope of Participation", "Location of Participation", "Payment Ratio", etc., among which "Scope of Participation", "Location of Participation", and "Payment Ratio" contain specific content descriptions, that is, these nodes are leaf nodes.
[0036] The path vector generation module 12 is used to convert each complete path from the root node to the leaf node in each business tree into a corresponding path vector. For example, the complete path from the root node "pension insurance" to the leaf node "insurance coverage" is "pension insurance-employee pension insurance-insurance and payment-individual-individual-insurance coverage". In this embodiment, the path vector is used as the id of the corresponding leaf node and is stored uniformly with the business forest.
[0037] The candidate node screening module 13 includes a vector model for obtaining a candidate node set based on the user question and all path vectors in combination with the vector model.
[0038] Figure 3 It is a schematic diagram of a flow chart of a candidate node screening module generating a candidate node set in an embodiment of the present invention.
[0039] like Figure 3 As shown, the candidate node screening module 13 generates a candidate node set including the following steps:
[0040] Step S3-1, vectorize the user question to obtain a user question vector.
[0041] Step S3-2, calculating the similarity between the user question vector and each path vector through a vector model. In this embodiment, the vector model is a BGE-M3 model.
[0042] Step S3-3, filter the leaf nodes corresponding to the path vectors corresponding to the similarities greater than a preset similarity threshold, and use all the leaf nodes obtained by the filter as the candidate node set. In this embodiment, the preset similarity threshold is 0.4.
[0043] The judgment module 14 is used to judge whether to generate an answer based on the candidate node set. If so, the answer is generated based on the large language model and the candidate node set. If not, the hierarchical matching module 15 is executed.
[0044] Figure 4 It is a schematic diagram of the working process of the judgment module in the embodiment of the present invention.
[0045] like Figure 4 As shown, the workflow of the judgment module 14 includes the following steps:
[0046] Step S4 - 1 , determining whether the candidate node set meets a preset determination condition, if so, executing step S4 - 2 , if not, executing the hierarchical matching module 15 .
[0047] The preset judgment condition includes at least one of the following preset conditions: the total number of leaf nodes in the candidate node set is not greater than the preset number of nodes, and the candidate node set is not an empty set. In this embodiment, the preset judgment condition is that the total number of leaf nodes in the candidate node set is not greater than the preset number of nodes and the candidate node set is not an empty set.
[0048] Step S4-2: for each leaf node in the candidate node set, the leaf node is combined with the user question and input into the large language model to obtain the confidence corresponding to the leaf node.
[0049] Step S4-3, select the leaf node corresponding to the highest confidence and the polishing prompt word combination and input them into the large language model to obtain the answer.
[0050] The hierarchical matching module 15 is used to perform hierarchical matching on all business trees according to user questions through a large language model to obtain hierarchical matching results.
[0051] Figure 5 It is a schematic diagram of a process of generating a hierarchical matching result by a hierarchical matching module in an embodiment of the present invention.
[0052] like Figure 5As shown, the hierarchical matching module 15 generates a hierarchical matching result, including the following steps:
[0053] Step S5-1, using the large language model to determine in turn whether each root node is relevant to the user's question, and taking the root node with relevance as the relevant node.
[0054] Step S5-2: Use the large language model to determine in turn whether each child node corresponding to each relevant node is relevant to the user's question, and use the child nodes with relevance as relevant nodes.
[0055] Among them, in step S5-2, when a single intermediate node has relevance, that is, when each child node corresponding to the intermediate node has no relevance, the hierarchical matching module 15 combines each child node corresponding to the intermediate node with the rhetorical question prompt word and inputs the combination into the large language model, generates rhetorical question information and displays it to the user. Then, the hierarchical matching module 15 sequentially determines whether each child node corresponding to the intermediate node has relevance to the user's reply to the rhetorical question information through the large language model, and uses the child nodes with relevance as relevant nodes.
[0056] For example, the user question "How to pay pension insurance" is hierarchically matched by the hierarchical matching module 15, and an intermediate node is matched. When the complete path of the intermediate node is "pension insurance-employee pension insurance-participation and payment", the hierarchical matching module 15 generates a counter-question information "Is the business handled by the individual or the unit" according to the sub-nodes "individual" and "unit", and continues to generate counter-question information according to the user's reply "individual handling" until a leaf node is matched or there is no matching leaf node.
[0057] Step S5-3, repeat step S5-2 until a hierarchical matching result is obtained.
[0058] The hierarchical matching results include: a single leaf node is relevant; all child nodes of a single parent node are relevant; and all nodes are not relevant.
[0059] The answer generation module 16 is used to generate answers by combining the hierarchical matching results with the large language model.
[0060] When the hierarchical matching result shows that a single leaf node is relevant, the answer generation module 16 combines the leaf node with the polishing prompt word and inputs it into the large language model to obtain the answer. Furthermore, the answer generation module 16 combines the neighboring nodes of the leaf node with the recommendation prompt word and inputs it into the large language model to generate recommendation information for the user's question, that is, to provide the user with the answer and related recommendation information at the same time.
[0061] For example, when the hierarchical matching result corresponding to the user question "What are the conditions for individuals to participate in pension insurance?" is a single leaf node, and the complete path of the leaf node is "pension insurance-employee pension insurance-participation and payment-individual-individual-coverage of insurance", the answer generation module 16 combines the specific description content of the leaf node "coverage of insurance" with the polishing prompt words and inputs them into the large language model to obtain the answer.
[0062] When the hierarchical matching result is that all child nodes of a single parent node are relevant, the answer generation module 16 determines whether the user question corresponds to a single child node or to multiple child nodes through the large language model.
[0063] When the user question corresponds to a single sub-node, the answer generation module 16 combines all the sub-nodes with the rhetorical question prompt words and inputs them into the large language model, generates rhetorical question information and displays the target node for the user to get the user's answer, combines all the leaf nodes corresponding to the target node with the polishing prompt words and inputs them into the large language model to get the answer.
[0064] For example, the user asks "I was unemployed and was planning to open a supermarket. How should I pay for my pension insurance?", and the complete paths of the matched sub-nodes include "pension insurance-employee pension insurance-participation and payment-individual-individual" and "pension insurance-employee pension insurance-participation and payment-individual-individual business owner", then the answer generation module 16 generates the counter-question "Are you an individual or an individual business owner?", so that the user can select one of the sub-nodes "individual" and "individual business owner" as the target node.
[0065] When the user question corresponds to multiple sub-nodes, the answer generation module 16 combines all leaf nodes corresponding to all sub-nodes with the polishing prompt words and inputs them into the large language model to obtain the answer.
[0066] For example, the user question "What is the difference between individual and self-employed business pension insurance?" The complete paths of the matched sub-nodes include "pension insurance-employee pension insurance-insurance and payment-individual-individual" and "pension insurance-employee pension insurance-insurance and payment-individual-self-employed business". Obviously, the user question needs to be answered by comprehensively considering the insurance scope, insurance location, payment ratio and other factors of the two. In this regard, the answer generation module 16 inputs the specific content descriptions described by all leaf nodes into the large language model together with the user question and the polishing prompt words to obtain a comprehensive answer.
[0067] When the hierarchical matching result is that all nodes are not relevant, the answer output by the answer generation module 16 is the preset answer content. In this embodiment, the preset answer content is empty, that is, the answer generation module 16 does not respond to questions that are not relevant to the document.
[0068] The master control module 17 stores a control program for controlling the operation of each module.
[0069] The following describes the process of using the large model search and answer device 100 to perform a large model search and answer method based on a document structure tree in conjunction with the accompanying drawings.
[0070] Figure 6 It is a flowchart of a large model retrieval question-answering method based on a document structure tree in an embodiment of the present invention.
[0071] like Figure 6 As shown in FIG. 1 , the large model retrieval question answering method based on the document structure tree includes the following steps:
[0072] Step S1, using the business tree construction module 11 to parse each document respectively to obtain the corresponding business tree.
[0073] Step S2: using the path vector generation module 12 to convert each complete path from the root node to the leaf node in each service tree into a corresponding path vector.
[0074] Step S3, using the candidate node screening module 13 to obtain a candidate node set based on the user question and all path vectors in combination with the vector model.
[0075] Step S4, using the judgment module 14 to judge whether to generate an answer based on the candidate node set, if so, generate an answer based on the large language model and the candidate node set, if not, execute step S5.
[0076] Step S5, using the hierarchical matching module 15 to perform hierarchical matching on all business trees according to the user question through the large language model to obtain a hierarchical matching result.
[0077] Step S6, using the answer generation module 16 to generate an answer through the large language model combined with the hierarchical matching results.
[0078] Functions and Effects of the Embodiments
[0079] According to the large model retrieval question-answering method and device based on the document structure tree involved in this embodiment, on the one hand, by constructing a business tree that conforms to the original logical order of the document and has no loops, the business corresponding to the document is guided; on the other hand, by screening the candidate node set and hierarchical matching, a quick reply is achieved in the scenario where the user knows the business he is handling, and heuristic question-answering is also supported in the scenario where the user does not understand the actual business process. In short, this method can achieve accurate, efficient and low-cost question answers.
[0080] Furthermore, the business tree is obtained by reconstructing the directory-level data, and there is no need to consider the segmentation or reorganization issues at the chapter level from a semantic level, thereby achieving lightweight data processing.
[0081] Furthermore, through matching nodes, rhetorical question prompt words and a large language model, rhetorical question information is generated, thereby proactively solving all user-related business questions.
[0082] Those skilled in the art should understand that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A large model retrieval question answering method based on a document structure tree, used to generate answers corresponding to user questions based on multiple documents, characterized in that: The following steps are involved: Step S1, parsing each of the documents to obtain corresponding business trees; Step S2, for each of the service trees, converting each complete path from the root node to the leaf node in the service tree into a corresponding path vector; Step S3, obtaining a candidate node set according to the user question and all the path vectors in combination with a vector model; Step S4, determining whether the answer is generated according to the candidate node set, if so, generating the answer according to the large language model and the candidate node set, if not, executing step S5; Step S5, performing hierarchical matching on all the service trees through the large language model according to the user question to obtain a hierarchical matching result; Step S6, generating the answer by combining the large language model with the hierarchical matching result, The root node of the business tree is the main title of the corresponding document. Each leaf node of the business tree is a description of each specific content in the document. According to the hierarchical relationship, each title of the document is used as a node of the business tree. When the title is used as a parent node, each sub-title corresponding to the title is used as a sub-node corresponding to the parent node.
2. The large model retrieval question answering method based on the document structure tree according to claim 1 is characterized in that: in, The step S3 comprises the following sub-steps: Step S3-1, vectorizing the user question to obtain a user question vector; Step S3-2, calculating the similarity between the user question vector and each of the path vectors through the vector model; Step S3-3, screening the leaf nodes corresponding to the path vectors corresponding to the similarities greater than a preset similarity threshold, and taking all the leaf nodes screened as the candidate node set.
3. The large model retrieval question answering method based on the document structure tree according to claim 1, Features: Wherein, the step S4 includes the following sub-steps: Step S4-1, determine whether the candidate node set meets the preset judgment condition, if yes, execute step S4-2, if no, execute step S5; Step S4-2, for each leaf node in the candidate node set, combine the leaf node with the user question and input the result into a large language model to obtain the confidence corresponding to the leaf node; Step S4-3, selecting the leaf node corresponding to the highest confidence and the polishing prompt word combination and inputting them into the large language model to obtain the answer, The preset judgment condition includes at least one preset condition that the total number of leaf nodes in the candidate node set is not greater than a preset number of nodes and the candidate node set is not an empty set.
4. The large model retrieval question answering method based on the document structure tree according to claim 1, Features: Wherein, the step S5 includes the following sub-steps: Step S5-1, using the large language model to determine in turn whether each of the root nodes is relevant to the user question, and taking the root nodes with relevance as relevant nodes; Step S5-2, using the large language model to sequentially determine whether each of the sub-nodes corresponding to each of the relevant nodes is relevant to the user question, and using the sub-nodes with relevance as the relevant nodes; Step S5-3, repeat step S5-2 until the hierarchical matching result is obtained. The hierarchical matching results include: A single leaf node has relevance; All the child nodes of a single parent node are related; All nodes are not related.
5. The large model retrieval question answering method based on the document structure tree according to claim 4 is characterized in that: in, When the hierarchical matching result shows that a single leaf node has relevance, the leaf node is combined with the polishing prompt word and input into the large language model to obtain the answer.
6. The large model retrieval question answering method based on the document structure tree according to claim 5 is characterized by: in, The neighbor nodes of the leaf nodes are combined with the recommended prompt words and input into the large language model to generate recommendation information for the user question.
7. The large model retrieval question answering method based on the document structure tree according to claim 4 is characterized in that: in, When the hierarchical matching result is that all the child nodes of a single parent node are relevant, the large language model is used to determine whether the user question corresponds to a single child node or to multiple child nodes. When the user question corresponds to a single sub-node, all the sub-nodes are combined with rhetorical question prompt words and input into the large language model, rhetorical question information is generated and displayed to the user as a target node to obtain the user's answer, and all the leaf nodes corresponding to the target node are combined with polishing prompt words and input into the large language model to obtain the answer, When the user question corresponds to a plurality of the sub-nodes, all the leaf nodes corresponding to all the sub-nodes are combined with the polishing prompt words and input into the large language model to obtain the answer.
8. The large model retrieval question answering method based on the document structure tree according to claim 4 is characterized by: in, When the hierarchical matching result is that all nodes are not relevant, the answer is a preset reply content.
9. The large model retrieval question answering method based on the document structure tree according to claim 4 is characterized in that: in, In step S5-2, when a single intermediate node has relevance, each of the subnodes corresponding to the intermediate node is combined with a rhetorical question prompt word and input into the large language model to generate rhetorical question information and display it to the user. The large language model is used to sequentially determine whether each of the sub-nodes corresponding to the intermediate node is relevant to the user's reply to the rhetorical question information, and the sub-nodes with relevance are used as the relevant nodes.
10. A large model retrieval question-answering device based on a document structure tree, used to generate answers corresponding to user questions based on multiple documents, characterized in that: include: Business tree construction module, path vector generation module, candidate node screening module, judgment module, hierarchical matching module and answer generation module, The business tree construction module is used to parse each of the documents to obtain the corresponding business trees. The path vector generation module is used to convert each complete path from the root node to the leaf node in the service tree into a corresponding path vector for each service tree. The candidate node screening module includes a vector model for obtaining a candidate node set according to the user question and all the path vectors, combined with the vector model, The judgment module is used to judge whether the answer is generated according to the candidate node set. If so, the answer is generated according to the large language model and the candidate node set. If not, the hierarchical matching module is executed. The hierarchical matching module is used to perform hierarchical matching on all the business trees through the large language model according to the user question to obtain a hierarchical matching result. The answer generation module is used to generate the answer by combining the large language model with the hierarchical matching result. The root node of the business tree is the main title of the corresponding document. Each leaf node of the business tree is a description of each specific content in the document. According to the hierarchical relationship, each title of the document is used as a node of the business tree. When the title is used as a parent node, each sub-title corresponding to the title is used as a sub-node corresponding to the parent node.
Citation Information
Patent Citations
Document question and answer method based on multi-way tree and large-scale language model and related equipment
CN116932730A
Generative question answering method and device based on knowledge graph and storage medium
CN118069817A
Knowledge graph retrieval method and device, medium and product
CN118227668A
Academic conference question-answering system based on large language model
CN118377867A
Large model knowledge question-answering method and device for power field
CN118585626A
Cited By
Tree document retrieval method and device, equipment and storage medium
CN121166842A
Tree document retrieval method, device and equipment and storage medium
CN121166842B
Code library question and answer method and device based on large model
CN121303374A