Question and answer method based on structured document retrieval enhancement

By constructing a tree-structured network, the problems of lost hierarchical structure of normative documents and insufficient cross-document reference management in existing technologies are solved, achieving efficient and accurate question-and-answer support, which is particularly suitable for structured documents such as policy documents.

CN120873137APending Publication Date: 2025-10-31GRG BANKING EQUIPMENT CO LTD

Patent Information

Application Number
CN202510981334.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing technologies struggle to preserve the hierarchical structure of highly structured normative documents with complex dependencies between chapters and clauses. This results in a lack of contextual continuity in search results, an inability to effectively manage potential cross-document connections and references, and an impact on the consistency and traceability of responses.

Method used

By constructing a tree-structured network, the hierarchical structure of documents is preserved. Leaf nodes are used to represent the semantic content of chunked processing, and summary information is generated in non-leaf nodes and root nodes. Combined with cross-document association relationships, a multi-document tree structure network is formed, which supports the tracing of logical connections between multiple documents.

Benefits of technology

It improves the accuracy and traceability of question and answer generation, ensures the comprehensiveness and consistency of information, and provides detailed and multi-level contextual summary information when users ask questions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873137A_ABST
    Figure CN120873137A_ABST
Patent Text Reader

Abstract

The invention discloses a question answering method based on structured document retrieval enhancement, and belongs to the technical field of artificial intelligence. The method comprises the following steps: performing structured information analysis and blocking processing on a structured document to construct a tree structure reflecting a hierarchical relationship of the document; based on a cross-document association relationship, combining the tree structures corresponding to the plurality of structured documents to form a multi-document tree structure network; and when a question request is received, generating corresponding retrieval enhancement information through the tree structure network, and inputting the retrieval enhancement information into the question and answer model to obtain answer information matched with the question request. According to the method, the multi-document tree structure network which takes the tree hierarchical relationship of the structured document as a core and fuses cross-document topic clustering and reference mapping is constructed, so that efficient retrieval enhanced questions and answers oriented to the structured document are realized, and the accuracy, context consistency and traceability of answers are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, and in particular relates to a question-answering method based on structured document retrieval enhancement. Background Technology

[0002] With the development of artificial intelligence, large-scale language models, and information retrieval technologies, users expect to quickly and accurately obtain information that closely matches their query intent from normative documents containing a large amount of structured information (such as policy documents, standards and regulations, and industry rules). This is especially true in government governance, legal compliance, and corporate management scenarios, where these documents often possess rigorous logical structures, complex hierarchical systems, and rich internal references, making traditional information retrieval models insufficient to meet the demands for intelligent interpretation and question-and-answer interactive experiences.

[0003] In related technologies, normative documents are typically input into information retrieval systems or generative language models in full-text or segmented form. Segmentation based on vector similarity or inverted index retrieval based on keyword matching is used to obtain fragments that match the user's query. Then, a large-scale pre-trained language model generates answers from the search results.

[0004] However, when faced with highly structured normative documents with complex dependencies between chapters and clauses, the relevant technologies still have the following shortcomings: First, the commonly used segmentation and vectorization strategies can destroy the hierarchical structure of the original document during the segmentation process, resulting in a lack of contextual continuity in the search results; second, potential connections and references across documents are difficult to manage uniformly during the search, resulting in fragmented answers; and third, some solutions cannot take into account the global logic and contextual consistency of the document when generating answers, affecting the explanatory power and traceability of user queries. Summary of the Invention

[0005] This application aims to address at least one of the technical problems existing in related technologies. To this end, this application proposes a question-answering method based on structured document retrieval enhancement to adapt to complex text scenarios, including policy documents, normative documents, or other hierarchical structured documents, and to achieve intelligent interpretation and efficient question-answering support for structured documents.

[0006] Firstly, this application provides a question-answering method based on structured document retrieval enhancement, the method comprising:

[0007] Structured information parsing and chunking are performed on structured documents to construct a tree structure that reflects the hierarchical relationship of the documents. In the tree structure, leaf nodes are used to represent semantic content based on chunking, and non-leaf nodes and root nodes are used to represent semantic content that summarizes the child nodes they include.

[0008] Based on cross-document relationships, the tree structures corresponding to multiple structured documents are combined to form a multi-document tree structure network.

[0009] Upon receiving a question request, corresponding retrieval enhancement information is generated through the tree structure network, and the retrieval enhancement information is input into the question answering model to obtain answer information that matches the question request.

[0010] Secondly, this application provides a question-answering device based on structured document retrieval enhancement, the device comprising:

[0011] The construction module is used to perform structured information parsing and chunking processing on structured documents to construct a tree structure that reflects the hierarchical relationship of the documents. In the tree structure, leaf nodes are used to represent semantic content based on chunking processing, and non-leaf nodes and root nodes are used to represent semantic content that summarizes the child nodes they include.

[0012] The union module is used to unite the tree structures corresponding to multiple structured documents based on cross-document relationships to form a multi-document tree structure network.

[0013] The question-answering module is used to generate corresponding retrieval enhancement information through the tree structure network when a question request is received, and input the retrieval enhancement information into the question-answering model to obtain answer information that matches the question request.

[0014] Thirdly, this application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the question-answering method based on structured document retrieval enhancement as described in the first aspect above.

[0015] Fourthly, this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the question-answering method based on structured document retrieval enhancement as described in the first aspect above.

[0016] Fifthly, this application provides a chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run computer programs or instructions to implement the question-answering method based on structured document retrieval enhancement as described in the first aspect above.

[0017] Sixthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the question-answering method based on structured document retrieval enhancement as described in the first aspect above.

[0018] This application proposes a question-answering method, apparatus, computer device, non-transitory computer-readable storage medium, chip, and computer program product based on structured document retrieval enhancement. By performing structured information parsing and block processing on structured documents, it can retain chapter numbers, clause numbers, and title information during the document preprocessing stage, ensuring the consistency of the document's hierarchical structure in subsequent processing, thereby avoiding the context loss problem caused by traditional block splitting. Specifically, by setting leaf nodes in the tree structure to carry original semantic units and setting non-leaf nodes and root nodes to generate summary information for their subordinate child nodes, it can form a layer-by-layer semantic summary within the document. Hierarchical information representation allows subsequent retrieval to obtain a combination of details and summaries as needed, thereby improving the comprehensiveness of information. Furthermore, the tree structure of multiple structured documents is combined based on cross-document relationships, which supports the tracing of logical connections between multiple documents, improves the coverage of cross-document queries and the completeness of document interpretation, and prevents one-sided answers caused by isolated interpretations. Moreover, when a user's question request is received, the retrieval enhancement information returned through the tree structure network can provide refined detailed evidence and multi-level contextual summary information in the question-answering model, thereby significantly improving the accuracy and traceability of question-answer generation.

[0019] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0020] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0021] Figure 1 This is a flowchart illustrating the question-answering method based on structured document retrieval enhancement provided in some embodiments of this application;

[0022] Figure 2 This is a schematic diagram of the retrieval information enhancement process provided in some other embodiments of this application;

[0023] Figure 3 This is a schematic diagram of the tree structure provided in some embodiments of this application;

[0024] Figure 4 This is a schematic diagram of the structure of a question-answering device based on structured document retrieval enhancement provided in some embodiments of this application;

[0025] Figure 5 This is a schematic diagram of the structure of a computer device provided in some embodiments of this application. Detailed Implementation

[0026] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0027] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used in the description of this application is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms "comprising" and "having," and any variations thereof, in the description, claims, and accompanying drawings of this application are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the description, claims, or accompanying drawings of this application are used to distinguish different objects, not to describe a specific order or hierarchy.

[0028] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments.

[0029] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," "linking," and "attachment" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0030] In this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, in this application, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0031] In this application, "multiple" refers to two or more (including two), and similarly, "multiple groups" refers to two or more (including two), and "multiple pieces" refers to two or more (including two).

[0032] Policy documents, as core documents for governments, enterprises, and organizations to formulate rules and guide actions, are characterized by rigorous logic and clear hierarchy. With the advancement of digital government construction, government staff and the public have an increasingly urgent need for intelligent retrieval and interpretation of policy documents, which are complex and diverse in scope. Furthermore, policy provisions are highly interconnected and hierarchical, and users often need to obtain highly customized information in specific business scenarios. Traditional title and keyword search methods are no longer sufficient to meet users' needs for accurate and efficient information retrieval.

[0033] With the rapid development of artificial intelligence technology, language models with strong semantic understanding capabilities and retrieval-augmented generation (RAG) technology have provided new technical pathways for policy analysis. However, these technologies face significant technical bottlenecks in their application:

[0034] (1) Existing RAG technology is difficult to capture the semantic and structural relationships between chapters. Its block division method often destroys the original tree structure of policy documents, resulting in the loss of its hierarchical features in the preprocessing stage.

[0035] (2) There are usually explicit reference networks between policy documents. Traditional RAG methods have difficulty capturing cross-document correlations. Information between different documents cannot be effectively connected and integrated, resulting in poor completeness and consistency of information retrieval.

[0036] (3) The text blocks generated by the existing RAG technology are too fragmented and the information density is insufficient, resulting in the model lacking efficient summarization ability and failing to meet users' needs for accurate and concise information interpretation.

[0037] In view of this, this application provides a tree-structured policy document customized large model RAG technology, which aims to achieve comprehensive and intelligent interpretation of policy documents, taking into account their highly structured and highly correlated characteristics. By constructing a technical logic framework with "parsing - segmentation - tree building - retrieval" as the core, it maximizes the retention of key information and realizes hierarchical automatic segmentation and integration of policy content to meet users' needs for multi-document queries.

[0038] It should be noted that in all specific embodiments of this application, any information or data related to policy documents will be authorized or agreed to in advance, and the collection, use and processing of such information or data will comply with relevant laws, regulations and standards.

[0039] The question-answering method based on structured document retrieval enhancement provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0040] The question-answering method based on structured document retrieval enhancement provided in this application embodiment can be executed by a computer device or a functional module or entity in a computer device that can implement the method.

[0041] The following uses a computer device as the execution subject to illustrate the question-answering method based on structured document retrieval enhancement provided in this application embodiment.

[0042] Figure 1 This is a flowchart illustrating a question-answering method based on structured document retrieval enhancement provided in some embodiments of this application. For example... Figure 1 As shown, the method includes steps 210 to 230.

[0043] Step 210: Perform structured information parsing and chunking on the structured document to construct a tree structure that reflects the hierarchical relationship of the document. In the tree structure, leaf nodes are used to represent semantic content based on chunking, and non-leaf nodes and root nodes are used to represent semantic content that summarizes the child nodes they include.

[0044] Structured documents refer to electronic documents that contain clearly defined chapters, clauses, titles, paragraph numbers, and other structural information, and have hierarchical and parsable characteristics, such as contracts, regulations, and policy documents.

[0045] Structured information parsing refers to the process of identifying and extracting hierarchical identifiers such as chapter numbers, clause numbers, headings, and paragraph numbers contained in structured documents using predefined rules or text analysis techniques, in order to facilitate subsequent structured modeling.

[0046] Chunking refers to the process of splitting the content in a document according to logical boundaries or semantic units based on the results of structured information parsing, and dividing it into several independently representable minimum information units.

[0047] A tree structure refers to a hierarchical structure model where semantic units are divided into blocks, and nodes are organized into parent-child hierarchical relationships based on the structured information. There is only one root node, and all other nodes are uniquely connected through parent-child relationships. Leaf nodes represent the smallest information units, non-leaf nodes represent summary information of their subordinate information units, and the root node represents the summary of the entire document or group of information.

[0048] When a computer receives a structured document, it first performs structured information parsing and segmentation processing on the document, including identifying hierarchical identifiers such as chapter numbers, clause numbers, paragraph levels, and heading text. Based on this, the document content is split into multiple independently understandable semantic units.

[0049] Based on the extracted hierarchical identifiers, the computer device splits the document into multiple independently representable semantic units, each of which is a chunk result; the computer device uses each semantic unit as a leaf node of a tree structure to form the smallest indexable information unit.

[0050] Subsequently, the computer device constructs a tree structure based on the semantic units contained in each leaf node and their hierarchical relationships. Non-leaf nodes generate corresponding summary information based on the information of their subordinate child nodes, while the root node summarizes the information of all subordinate non-leaf nodes and leaf nodes to generate an overview summary, thus forming multi-level node content with summaries in the tree structure.

[0051] Step 220: Based on the cross-document association, combine the tree structures corresponding to multiple structured documents to form a multi-document tree structure network.

[0052] A tree-structured network refers to a complex information network formed by connecting multiple tree-like structures through cross-document relationships (such as references or thematic similarity).

[0053] Specifically, cross-document relationships can include: relationships between different policy documents, such as the citation relationship or subject matter similarity between Law A and Law B; and relationships between different versions of the same policy, such as the similarity of clauses or modifications between the 2015 and 2020 versions of Law C.

[0054] For multiple structured documents, computer devices each construct a tree structure. Then, by utilizing the cross-document relationships between documents, such as reference identifiers or topic relevance, the nodes in the multiple tree structures are connected to form a cross-document traceable information network structure, namely a multi-document tree structure network.

[0055] For example, between different versions of the same policy document, computer equipment generates cross-version citation maps by comparing the revisions of the same clauses in different versions, ensuring that content differences between versions can be traced and integrated. Furthermore, the computer equipment also groups clauses belonging to the same topic in different versions into the same node based on the thematic relevance of the clauses, forming cross-version thematic aggregations. In this way, whether it's citations between different policy documents or version comparisons of the same policy document, information aggregation and tracing can be effectively performed through a tree-structured network.

[0056] Step 230: Upon receiving a question request, generate corresponding retrieval enhancement information through the tree structure network, and input the retrieval enhancement information into the question answering model to obtain answer information that matches the question request.

[0057] Question-answering models refer to language models that can generate answers based on natural language questions and contextual clues. Retrieval enhancement information refers to the suggestion or reference information generated by combining highly relevant context retrieved through a tree-structured search when a large-scale language model generates answers.

[0058] When a computer device receives a user's question request, it retrieves and queries a set of relevant nodes through the multi-document tree structure network, including summaries of relevant nodes, document metadata, etc., and inputs this retrieval enhancement information along with the user's question into the question-answering model to generate an answer that meets the user's needs.

[0059] The question-answering method based on structured document retrieval enhancement provided in this application, by performing structured information parsing and block processing on structured documents, can retain chapter numbers, clause numbers, and title information in the document preprocessing stage, ensuring the consistency of the document's hierarchical structure in subsequent processing, thereby avoiding the context loss problem caused by traditional block splitting. Specifically, by setting leaf nodes in the tree structure to carry original semantic units and setting non-leaf nodes and root nodes to generate summary information for subordinate child nodes, a hierarchical information representation with layer-by-layer semantic summarization can be formed within the document, allowing subsequent retrieval to obtain content combining details and summaries as needed, thereby improving the comprehensiveness of information. Furthermore, by combining the tree structures of multiple structured documents based on cross-document association relationships, it can support the tracing of logical connections between multiple documents, improve the coverage of cross-document queries and the completeness of document interpretation, and prevent one-sided answers caused by isolated interpretations. Moreover, when a user's question request is received, the retrieval enhancement information returned through the tree structure network can simultaneously provide refined detailed evidence and multi-level context summary information in the question-answering model, thereby significantly improving the accuracy and traceability of question-answer generation.

[0060] In particular, when the structured document is a policy document, due to the strong versioning relationships within policy documents (the same policy may have multiple versions), the computer device not only merges different documents into a tree structure based on explicit references but also detects multiple versions of the same policy document for joint processing. If differences between versions are detected, the device automatically marks the revisions and generates update description nodes, marking them as version differences in the tree structure, making it easy for users to view the document revision history and specific modifications.

[0061] In some embodiments, the above method further includes: when processing policy documents involving multiple versions, the computer device first detects and identifies the version information of each document, and performs version difference processing on different versions of the same policy document; when combining across documents, if multiple versions of the same policy document are found, the device automatically marks the version differences and generates update description nodes by comparing the revised content, ensuring that the document version information is traceable and accurately identified in the tree structure; wherein, the version difference information includes added clauses, deleted clauses, or modifications to the content of clauses, and the differences between all versions will be automatically marked in the generated tree structure network, and a corresponding update description node will be created for each version, so that the historical versions of the document can be clearly displayed.

[0062] In other words, this application embodiment also provides a version difference processing mechanism, which includes: the computer device establishing an association between the version difference node and the original clause node; if the revision content in the version affects the semantics or structure of certain clauses, the device will automatically update the revised content to ensure that the revision information of different versions can be displayed and compared in the multi-document tree structure. In the update description node, the device will provide a detailed revision summary of the version and mark the timestamp and version number of the changes made, ensuring that users can trace the specific revision history of the document.

[0063] Specifically, during cross-document collaboration, if different versions of the same policy are detected, the computer device can compare the tree structure corresponding to each version to determine one or more of the updated leaf nodes, non-leaf nodes, or root nodes. Furthermore, based on the existence of updated leaf nodes, non-leaf nodes, or root nodes, the computer device can automatically annotate the revised content and generate update description nodes based on the summary information corresponding to the nodes before and after the update, thereby supporting version association and traceability. Thus, by proposing a processing mechanism based on the unique hierarchical structure and version management of policy documents, efficient technical support is provided for the interpretation and tracking of policy documents, enabling different versions of policy provisions to be traced and accurately interpreted through a tree structure.

[0064] The following example illustrates the steps outlined above. Assume a policy document titled "XXX Protection Law," containing multiple chapters and clauses, each with specific policy content. A computer system performs structured information parsing on this policy document and constructs a tree structure based on the hierarchical relationships between chapters and clauses. This structured document includes version management information, and multiple versions of the policy document have been processed and annotated, allowing for the tracing of modifications made in each version.

[0065] The computer equipment first performs structured information parsing on the "xxx Protection Law," and the parsing results include the hierarchical numbering of each chapter and clause. For example:

[0066] Chapter 1: General Provisions

[0067] Article 1: General Provisions

[0068] Article 2: General Provisions II

[0069] Chapter 2: A

[0070] ...

[0071] Each clause is treated as a node, and the content is divided into blocks using hierarchical numbering to construct a tree structure. Leaf nodes represent the original text of the clause, while non-leaf nodes and the root node contain summary information for each chapter or section, forming a complete tree structure. This structure includes the content of multiple versions of the "Environmental Protection Law." That is, assuming there are multiple versions of the "Environmental Protection Law," the computer device automatically detects and marks version differences through a version management mechanism. For multiple versions of the same clause, the device uses the revised content and update description information as update nodes, marking the version update in the tree structure. When crossing documents, the computer device combines the tree structures of multiple versions through citation relationships and topic similarity to ensure that cross-version connections and differences can be clearly traced.

[0072] When a user asks a question, the computer uses a tree-structured network to generate enhanced retrieval information and inputs it into the question-answering model. For example, if the user's question is: "According to Article 2 of the 'xxx Protection Law,' how is xxx defined?"

[0073] The computer device then searches within the tree-structured network, first finding Article 2, "General Principles II," under Chapter 1, "General Principles." This node contains the original text and summary information of the relevant policy. The computer device then traces back through the tree structure hierarchy to obtain information about the parent nodes related to Article 2 (e.g., the summary of the "General Principles" chapter) and supplements the metadata of the relevant documents (e.g., issuing organization, revision date).

[0074] Finally, the computer device, combining the retrieved information and supplementary context, inputs the generated answer into the question-answering model. This model, incorporating information from all levels, provides the following answer: "According to Article 2 of the 'xxx Protection Law,' General Provisions II stipulate the xx standard, including provisions for xxx, covering relevant restrictions on xxx. According to the 2018 revision, the new standard requirements are more stringent than the previous provisions…". Because this policy document involves multiple versions, the computer device automatically marks the revision history of this clause from the 2015 version to the 2018 version and provides detailed update descriptions of the revisions.

[0075] In this embodiment, the computer device constructs a precise tree structure for policy documents through structured information parsing and chunking, ensuring that differences between document versions can be traced. When a user asks a question, the system can not only extract precise policy content through the tree structure, but also trace back historical information based on version differences, generating accurate and context-rich answers.

[0076] In this way, the unique features of policy documents, such as version management, hierarchical identification, and clause types, are effectively incorporated into the retrieval process, ensuring the accuracy and timeliness of policy and regulation interpretation. This approach is particularly suitable for policy and regulation documents and solves the problems of multi-version management and update tracking of policy documents.

[0077] Current methods for retrieval structured information between multiple documents often only perform simple keyword associations or explicit hyperlink-based citations between documents. However, they lack the ability to summarize and organize similar topic nodes within the documents. As a result, when multiple documents are combined, information can only be linked together at a shallow level and cannot be deeply integrated at the semantic level. This limits the explanatory power and traceability of multi-document question answering in complex business scenarios.

[0078] Therefore, in some embodiments, the step of combining the tree structures corresponding to multiple structured documents based on cross-document association relationships to form a multi-document tree structure network includes: extracting explicit reference relationships between each structured document; establishing connections between relevant nodes in the tree structures of each document according to the reference relationships to form cross-document node mapping relationships; determining nodes with similar topics in each structured document; performing topic clustering on the nodes with similar topics to form document association relationships between multiple documents under the same topic; and forming a multi-document tree structure network based on the node mapping relationships and document association relationships.

[0079] After constructing tree structures for multiple structured documents, the computer device first extracts the explicit reference relationships between documents for multi-document joint processing. These include reference numbers, clause hyperlinks, or other traceable identifiers. The device then uses regular expression matching and named entity recognition tools to extract these reference identifiers in a structured manner. Finally, it establishes cross-document node connections between the starting node of the reference and the target node of the reference at the tree structure level, thereby forming a directional reference mapping between multiple trees.

[0080] Secondly, the computer equipment performs topic-based clustering on the root nodes of each tree structure to categorize one or more root nodes according to topic similarity. Specifically, based on the vectorized representation of the summary, the computer equipment uses a Gaussian mixture model to cluster all root nodes. During the clustering process, the number of Gaussian distributions and the parameters of each distribution (e.g., mean, covariance, mixture coefficients) are first initialized. Then, within the iterative framework of the expectation-maximization algorithm, posterior probability calculations and parameter updates are performed until the model converges to determine the topic cluster to which each root node belongs in the embedding space. Finally, root nodes belonging to the same Gaussian distribution are grouped into a single topic cluster category.

[0081] The computer device then integrates all reference connections and topic clustering units into a unified index table of the tree-structured network, so that the reference logic can be traced across documents in subsequent retrieval requests, while the same topic can be consistently summarized across documents, forming a complete tree-structured network across multiple documents.

[0082] In policy document application scenarios, for example, consider a policy document titled "XX Law," which contains multiple versions (e.g., the 2010 and 2020 versions), each with different clauses and revision records. The computer device first performs structured information parsing on these two versions of the document and constructs a tree structure for each version based on chapter and clause numbers. Each node in the tree structure represents a clause; leaf nodes store the original text, and non-leaf nodes store summary information. When processing different versions of "XX Law," the computer device first extracts explicit citation relationships between the versions. For example, Article 5 of the 2010 version cites Article 8, and Article 6 of the 2020 version also cites Article 9. The device uses regular expression matching and named entity recognition tools to extract these citation identifiers from the documents and establishes cross-version citation mapping relationships for the corresponding clause nodes in the tree structure. The computer device then performs summary vectorization on the root nodes of each version and clusters the summary vectors using a Gaussian mixture model. For example, if a computer determines that several clauses in the 2010 and 2020 versions are highly similar in topic, it will categorize these similar clause nodes into a topic cluster. Based on the explicit citation relationships and topic similarity clusters extracted above, the computer will combine the tree structures of the two versions to form a multi-document tree structure network. In this network, each node can point to relevant nodes in other documents through citation relationships, and nodes under the same topic are aggregated into a topic cluster unit.

[0083] Therefore, by using explicit cross-document references and topic similarity clustering, the device can effectively compare and trace different versions of policy documents, ensuring that users can clearly view the modification history of the clauses, especially policy changes or additions.

[0084] By adopting the above method, not only can the reference tracing mechanism between documents be accurately established through node mapping relationships, but also the node information of the same topic in multiple documents can be automatically aggregated to build a unified topic view across documents. This enables more comprehensive and structured knowledge network support, provides better contextual basis for downstream question answering models to generate more consistent and accurate answers, and significantly improves the scalability and intelligence level of the multi-document question answering system.

[0085] In some embodiments, the step of extracting explicit reference relationships between structured documents and establishing connections between relevant nodes in the tree structure of each document according to the reference relationships to form a cross-document node mapping relationship includes: performing text analysis on the text content at preset positions in each structured document to detect the reference identifiers or link information contained therein; locating the corresponding node in the referenced document according to the reference identifiers or link information, and establishing a connection relationship between the reference node and the referenced node.

[0086] When extracting citation relationships from multiple structured documents, computer equipment first identifies predefined text regions (including but not limited to the beginning of clauses, the end of paragraphs, or footnotes) as key detection areas. It then analyzes the text within these regions using regular expression matching and a rule dictionary to identify possible citation numbers, document abbreviations, or hyperlink formats.

[0087] After identifying the reference identifier, the computer device uses a node number matching algorithm in the tree structure to locate the corresponding target node in the tree of the referenced document, establishes a cross-document node link between the two, and records the reference direction, source document information and location information, storing them in the reference index table.

[0088] In different policy document scenarios, taking the "xxxx Licensing Specification" as an example, Article 12 contains the statement "It shall be implemented in accordance with Article 15 of the 'xxxx Management Measures'". The computer device detects the reference identifier "refer to Article 15 of the 'xxxx Management Measures'" through text analysis of the note position of the article, and further retrieves the node corresponding to Article 15 in the referenced document. In the tree structure network, the reference node and the target node are connected, so as to ensure that when users query relevant penalty details during Q&A, they can trace back to the clause content of the "xxxx Management Measures" across documents.

[0089] In scenarios involving different versions of the same policy document, computer equipment detects and extracts explicit citations from both versions of the "xx Law". For example, Article 3 of the 2015 version cites Article 5, while Article 5 of the 2020 version also cites Article 6. The computer equipment automatically extracts these citation identifiers and link information from the document using regular expression matching and named entity recognition tools. For instance, if Article 3 of the 2015 version contains "see Article 5", and Article 5 of the 2020 version also contains "according to the provisions of Article 6", the computer equipment identifies and records these citations. Through analysis of citation identifiers and link information, the computer equipment locates the corresponding node in the cited document. For example, if Article 3 of the 2015 version cites Article 5, the computer equipment will locate the node in Article 5 of the 2020 version. The computer equipment automatically finds the corresponding node of the target clause based on the citation identifier and creates a connection relationship between the citing node and the cited node in a tree structure. The computer then applies the extracted reference relationships to the tree structure, ensuring that each referenced node is connected to its corresponding target node. This allows for cross-document node mappings between nodes across versions. For example, a mapping relationship is established between Article 3 of the 2015 version and Article 5 of the 2020 version, and the system can use these mapping relationships to trace the correspondence between clauses in different versions.

[0090] The above methods can systematically identify and utilize explicit reference identifiers or links in the body of structured documents, accurately establish mapping relationships between document nodes, thereby significantly enhancing the traceability and information coverage of the multi-document tree structure network, avoiding context loss due to unassociated reference information, and significantly improving the completeness and reliability of multi-document intelligent question answering.

[0091] In some embodiments, determining nodes with similar topics in each structured document and clustering these nodes to form document associations among multiple documents under the same topic includes: generating vector representations of the summary information corresponding to the root nodes of each structured document; clustering each vector representation using a Gaussian mixture model, and optimizing the clusters using an expectation-maximization algorithm under a set initial number of clusters to obtain the probability distribution of each node's topic clusters; and, based on the probability distribution, assigning the topic cluster with the highest probability as the primary affiliation of the root node, and merging root nodes in the same topic cluster into the same supernode to achieve grouping and aggregation of root nodes with similar topics.

[0092] After constructing a tree structure of multiple structured documents, the computer device generates a vector representation for the root node summary of each document using models such as BERT (Bidirectional Encoder Representations from Transformers) and its variants, forming a vector set of all root node summaries. The computer device further configures a Gaussian mixture model and sets an initial number of clusters, then executes the Expectation-Maximization (EM) algorithm to fit the probability density distribution of these vectors in a multidimensional space. After model convergence, each root node obtains a probability distribution across various topic clusters, and the computer device selects the topic cluster with the highest probability as its primary topic label.

[0093] Subsequently, the computer device creates a supernode for each topic cluster and maps all root nodes belonging to that topic cluster to the supernode through a pointer structure, ultimately forming a multi-document cross-file tree structure network organized by topic to support consistent access to multi-document topics during subsequent retrieval enhancements.

[0094] For example, the computer device generates summaries of the root nodes for each version of the "XX Law". For instance, Chapter 2, "Basic Laws of Taxation," of the 2015 version might contain a clause regarding "Value Added Tax," and the corresponding clause in the 2020 version has similar content. The computer device converts these root node summaries into vector representations using natural language processing techniques and encodes the summary of each root node using vectorization methods (such as Word2Vec and BERT). The computer device then clusters the generated vector representations using a Gaussian mixture model. First, the computer device initializes the number of Gaussian distributions (e.g., set to 3 clusters) and then optimizes it using an expectation-maximization algorithm. This algorithm calculates the probability distribution of the topic cluster to which each root node belongs, thus determining the most likely topic cluster to which each root node belongs. Based on the probability distribution of the topic cluster to which each root node belongs, the device selects the topic cluster with the highest probability as the primary affiliation of that root node. For multiple root nodes belonging to the same topic cluster, the device merges them into a single supernode, thereby achieving grouping and aggregation of root nodes with similar topics. For example, all root nodes related to "Article A" will be clustered into a single supernode, facilitating subsequent retrieval and question-and-answer processing. Thus, through topic clustering, computer devices can automatically identify and categorize topic-related content in policy documents, such as clauses related to "Basic Law A" or "Article A," regardless of whether they appear in different versions of the document; they can all be aggregated into the same topic. This method ensures consistency across multiple versions of policy documents and systematic processing of information.

[0095] The above methods not only establish connections between cross-document nodes with explicit reference relationships, but also effectively group multi-document nodes lacking direct references based on Gaussian mixture clustering of topic similarity, thereby forming an information aggregation view of topic clustering within a multi-document scope. This greatly enhances users' ability to fully grasp information on the same topic when searching for complex questions, and supports question-answering models to generate more aggregated and consistent answers.

[0096] Related technologies often only focus on splitting documents into several text blocks, but fail to hierarchically manage the contextual structure and summary information between the split blocks. This results in subsequent retrievals having to process each block individually, lacking both cross-level summarization capabilities and the ability to quickly locate document summaries. Such technologies are unsuitable for intelligent question-answering scenarios that require multi-layered interpretation, such as policy or regulatory documents.

[0097] Therefore, in some embodiments, performing structured information parsing and chunking on the structured document to construct a tree structure reflecting the hierarchical relationship of the document includes: identifying the hierarchical numbering information of chapters, clauses, or paragraphs in the document based on predefined hierarchical identification rules; chunking the document content according to the hierarchical numbering information, and generating leaf nodes and their parent-child relationships in the tree structure according to the chunking results; and generating summary information from bottom to top according to the content of each leaf node and its parent-child relationship, as the node content of non-leaf nodes and root nodes.

[0098] When computer equipment performs structured information parsing and block processing of structured documents, it first uses regular expressions and hierarchical numbering mapping tables to parse the document content based on predefined hierarchical identification rules, such as "Chapter X", "Article X" or "XXX" style, to extract the hierarchical number and title information of each block, thereby identifying the structural hierarchy of chapters, clauses and paragraphs.

[0099] Next, the computer device breaks down the document content into multiple semantic units based on these hierarchical numbering information, generating leaf nodes in a tree structure for each unit, and automatically establishing parent-child relationships for the leaf nodes according to the numbering structure. Subsequently, based on the content of each leaf node, the computer device uses language models or keyword extraction methods to recursively generate summary information from the bottom-level nodes upwards, forming summary information for non-leaf nodes and the root node, ensuring that each level of the document retains both the original text of the segments and the summary information of the upper levels.

[0100] For example, the computer device first identifies the hierarchical numbering information in the Company Law document according to predefined hierarchical identification rules. These hierarchical numbers include chapter numbers, clause numbers, and paragraph numbers, such as the clause number of Article 1 "Company Establishment" under Chapter 1 "General Provisions". The computer device treats each clause, chapter, or paragraph as an independent node, constructing a tree structure reflecting the hierarchical relationship of the Company Law. Based on the identified hierarchical numbering information, the computer device divides the content of the Company Law document into blocks and assigns each block to a leaf node according to its hierarchical information. Each leaf node represents the original text of a clause or paragraph, and a parent-child relationship is established between each node. For example, Chapter 1 "General Provisions" is a non-leaf node, and Article 1 "Company Establishment" is a leaf node, linked together through a tree structure. Based on each leaf node and its parent-child relationship, the computer device generates summary information for each non-leaf node and the root node in a bottom-up manner. For example, the summary generated for Article 1, "Company Incorporation," under Chapter 1, "General Provisions," can include the main provisions of that article. This summary is then passed to higher-level nodes, such as the root node of Chapter 1, forming a summary of the entire chapter. Ultimately, the root node contains a summary of the entire Company Law.

[0101] In the above embodiments, uniformity can be achieved in the splitting, structuring, and summary-level expression of structured documents, ensuring that each node has traceable original text information, while also possessing multi-level summarization capabilities. This provides extremely high flexibility for accessing documents at different granularities in subsequent question-and-answer scenarios, significantly improving the information acquisition efficiency and interpretability of complex structured documents in the retrieval enhancement and question-and-answer process.

[0102] Currently, when performing block segmentation, structured documents often rely directly on natural paragraph splitting or simple line break recognition, failing to fully utilize the original hierarchical numbering information of the document to determine content boundaries. This can easily disrupt the original logical organization of the document, especially when there are many chapters and deeply nested clauses. It can easily lead to overly detailed block segmentation or contextual confusion, thereby affecting the accuracy of the subsequent tree structure.

[0103] Therefore, in some embodiments, the step of dividing the document content into blocks according to the hierarchical numbering information and generating leaf nodes and their parent-child relationships in a tree structure according to the block division results includes: performing format matching on the hierarchical numbering information contained in each structured document to determine the boundary positions of chapters or clauses; dividing the document content into basic node units based on the boundary positions to form leaf nodes and establish their hierarchical correspondence with the parent nodes.

[0104] When a computer device segments a structured document based on its hierarchical numbering information, it first performs a format matching operation on the hierarchical numbering information contained in the document. For example, it matches hierarchical numbers such as "Chapter X", "Article X", and "XXX". This is done by comparing the results using predefined regular expressions or a hierarchical numbering dictionary to determine the boundary positions of each chapter, clause, or paragraph in the text. Subsequently, based on these boundary positions, the computer device precisely breaks down the document content into basic node units. Each basic node unit serves as a leaf node in a tree structure, and by utilizing the structural relationships implied by its hierarchical number, it establishes parent-child connections between the leaf node and its parent nodes within the tree structure, ensuring the integrity of the hierarchy and information transmission of the entire tree structure.

[0105] This approach avoids disrupting the logical boundaries of existing chapters and clauses during document segmentation, ensuring that leaf nodes have clear origins and boundaries, and preserving accurate hierarchical mapping relationships within the tree structure. This provides a reliable structural foundation for subsequent summary generation, retrieval enhancement, and question-answering reasoning, thereby improving the stability and interpretability of overall information processing.

[0106] In some embodiments, the above method further includes: setting a text block length threshold; traversing and checking the content length of each node's basic unit; when a text block length exceeding the text block length threshold is detected, splitting the node into multiple virtual nodes to ensure that the content length of each leaf node meets the preset requirements.

[0107] After the computer device completes the document segmentation and generates the leaf nodes of the tree structure, it continues to perform the text block length detection operation. A text block length threshold is preset (e.g., 1024 characters), and all leaf nodes are traversed in turn to check their actual content length. When the text length of a certain leaf node is found to exceed the threshold, the computer device splits the content of the node into multiple virtual nodes according to sentence boundaries or semantic segmentation points. The content of each virtual node does not exceed the threshold, and the hierarchical mapping relationship between these virtual nodes and the original parent node is maintained in the tree structure, thereby forming a controllable granularity leaf node set that can be processed later.

[0108] By setting a text block length threshold and performing virtual node splitting, computer devices can effectively prevent truncation problems caused by excessively long text blocks in subsequent retrieval enhancement or question answering model input, ensuring that each leaf node can be fully processed by the language model, thereby improving the accuracy of question answering, contextual integrity and reasoning consistency, and enhancing the engineering adaptability of large-scale structured document parsing systems.

[0109] In some embodiments, the step of generating summary information from bottom to top based on the content of each leaf node and its parent-child relationship, to serve as the node content of non-leaf nodes and the root node, includes: for each non-leaf node, generating corresponding summary information based on the text content of its direct subordinate child nodes using a preset summarization algorithm, and using the summary information as the node content of the non-leaf node; for the root node, generating an overview summary information based on the summary information of all subordinate nodes and leaf node information, and using it as the node content of the root node; and generating corresponding text embedding vectors and word segmentation results for the original text blocks contained in each leaf node and the summary information contained in the non-leaf nodes through a text embedding model and a word segmenter, respectively, and storing them on their respective nodes.

[0110] When generating a tree structure, the computer device preserves the original semantic information of the original text blocks of each leaf node. Based on this, it performs bottom-up summary generation according to the parent-child relationship of the tree structure: for each non-leaf node, it generates corresponding summary information based on the content of its directly subordinate child nodes using a preset summary algorithm, and writes the summary information into the non-leaf node; for the root node, it generates an overview summary as its node content by taking the summaries of all its subordinate non-leaf nodes and the core information of all leaf nodes as input.

[0111] In addition, the computer device selects a specified text embedding model (e.g., BERT, Sentence-BERT) and a word segmenter (e.g., Jieba or NLP general word segmenter) to generate corresponding text vectors and word segmentation results for the original text blocks of each leaf node and the summary information of non-leaf nodes and root nodes, and stores these features in their respective node structures for subsequent retrieval and question answering generation.

[0112] The above approach not only preserves the original text of the smallest information unit (leaf node) in the multi-layered structured document, but also realizes the hierarchical expression of intermediate summaries and overview summaries. Furthermore, by using node-level embedding vectors and word segmentation information, it provides unified and indexable structured features for retrieval enhancement and large-scale model reasoning, which greatly reduces the redundant context passing required during the answering process and improves the speed of question answering and the interpretability of information.

[0113] In some embodiments, generating corresponding retrieval enhancement information through the tree structure network includes: expanding the tree structure network into a single-layer structure, and performing candidate recall on nodes in the tree structure network based on a hybrid retrieval method combining semantic information and keyword information to obtain one or more recalled nodes; performing backtracking of node information according to the hierarchical relationship of one or more recalled nodes to supplement their context information and associated file meta-information; applying a re-ranking model to optimize the sorting of the recalled node results, and combining the optimized nodes with the corresponding file meta-information to generate retrieval enhancement information for the question-answering model.

[0114] The recall node can be a node in different policy documents or a node in different versions of the same policy document.

[0115] After receiving a user's query request, the computer device first expands all nodes in the multi-document tree structure network in a single-layer structure, and then calculates a comprehensive matching score based on a hybrid retrieval method, that is, combining the semantic vector of the node (e.g., generated by BERT or Sentence-BERT) with the keyword information of the node (e.g., inverted word segmentation index), performs a recall operation on the node, and obtains a set of candidate nodes that are highly relevant to the query.

[0116] Subsequently, the computer device traces back to the parent node information or expands to the child node information based on the hierarchical relationship of these candidate nodes in the tree structure, and at the same time supplements the document meta-information associated with the node (such as document source, version, and publication time) to enrich the context.

[0117] For example, when processing policy documents, computer devices can recall and query relevant nodes from different versions of the policy documents, and supplement the historical revisions based on the information from these nodes. For instance, the 2015 and 2020 versions of the Labor Law may contain the same clause (such as the "wage payment" clause), but because the versions are different, the device will retrieve the "wage payment" clause from both the 2015 and 2020 versions when recalling the node, allowing the system to provide a comparison of the two versions when responding.

[0118] After recalling a node, the computer device traces back the node's context information based on the hierarchical relationship in the tree structure and searches for related document metadata (such as revision time and revision summary). For example, when a user queries a clause, the device automatically provides a comparison of historical versions and revisions of that clause, along with explanations of the updates, helping the user understand the changes between current and historical policies. When multiple versions of policy document nodes are recalled, the device generates corresponding update explanation nodes based on the version difference handling mechanism, marking the version differences and incorporating them as part of the enhanced retrieval information. In this way, users can not only obtain information about the current version of the clause but also directly see the changes in the clause across different versions, ensuring that the differences between versions are accurately presented.

[0119] Finally, the computer device performs comprehensive sorting optimization on the recalled node set based on a reordering model (such as Cross-Encoder or a deep reordering model based on sentence matching), and combines the sorted node text, file metadata and query hints to form retrieval enhancement information that can be used by the question answering model to support high-quality answer generation.

[0120] For example, after receiving a user's query request, the computer device first expands all nodes in the multi-document tree structure network in a single-layer structure and constructs two types of retrieval indexes: one type is a text embedding vector generated by BERT or Sentence-BERT based on node text and summary information, which is stored in a vector database (such as FAISS or Milvus) for calculating semantic similarity; the other type is a keyword index based on inverted tables, which is constructed by performing word segmentation on node text and summary for quickly retrieving node identifiers containing specified keywords.

[0121] When a query arrives, the computer device simultaneously performs a vector similarity query (such as Euclidean distance or cosine similarity Top-K) in the vector database and queries the inverted index table for node identifiers containing the query keywords to obtain the keyword matching results.

[0122] Subsequently, the computer device combines the semantic retrieval score and the keyword matching score into a comprehensive matching score through a weighted fusion method (e.g., semantic score weight 0.7, keyword score weight 0.3), and selects the Top-K nodes as candidate nodes based on this comprehensive score.

[0123] For these candidate nodes, the computer device then traces back to their parent and sibling node summary information and expands down to their direct child node information according to their hierarchical relationship in the tree structure, and adds the metadata of their respective documents (document source, publication date, version information, etc.) to supplement the context.

[0124] Finally, the computer device uses a Cross-Encoder or a deep re-ranking model based on sentence pair matching to cross-score and re-rank these expanded candidate node sets with the user query, selects the top-ranked nodes, and combines them with query hints to form the final retrieval enhancement information, which is then input into the question answering model.

[0125] In the above embodiments, by quickly recalling relevant nodes using a hybrid retrieval method, and by using the contextual hierarchy of a tree structure to backtrack and supplement the recalled nodes, and by combining the reordering model to optimize the final information, the question-and-answer generation process is guaranteed to have contextual continuity and traceability, which effectively improves the accuracy, compliance and interpretability of the generated answers.

[0126] Currently, when extracting structured information from various document types (such as PDF, DOCX, and HTML), the method often relies solely on plain text OCR or simple HTML parsing, ignoring the original physical layout structure and semantic element distinctions of the document. This makes it difficult for the obtained structured units to map the true logical hierarchy and document structure in subsequent tree structure construction. This is especially true for government regulations and policy documents involving tables or formulas, where information loss and structural confusion are particularly prominent, limiting the adaptability to multi-document question answering.

[0127] Therefore, in some embodiments, the process of parsing structured information in structured documents includes: for PDF or DOCX type structured documents, detecting the document's layout structure and physical layout elements, and extracting the content of text areas, table areas, and formula areas respectively to obtain multiple structured text units; for HTML type structured documents, performing web page content collection according to a predetermined target page structure and data storage method, and extracting field information such as title, publishing organization, and body content to obtain multiple structured text units; and saving the parsed structured text units to a local storage medium to assist in the construction of a tree structure.

[0128] When computer equipment parses structured information from structured documents, for PDF or DOCX documents, it first uses a layout analysis engine (such as PaddleOCR based on visual layout or a deep learning table detection model) to detect its physical layout elements, including text segments, table areas, and formula areas, locates the boundaries and structure of these areas on the page, and extracts the content of each area to form a structured text unit. For HTML documents, based on a pre-configured target page structure template and data storage method, it uses an extensible web crawling framework (such as Scrapy) to collect web page nodes, extracting key fields such as titles, publishing organizations, main body text, and optional metadata to form corresponding structured text units.

[0129] Subsequently, the computer device writes the aforementioned structured text units into a programmable access format (such as JSON or Markdown files) to local storage media as input for subsequent tree structure construction, ensuring that subsequent processing can build a complete document tree based on accurate, hierarchical text blocks.

[0130] For example, when dealing with a PDF document titled "xxxx Management Standards" issued by a government department, the computer equipment first detects its layout elements and finds that it contains multiple flowcharts, approval forms, and main text clauses. The table recognition module is used to extract the table data containing the "List of Materials for Approval Processes," and the text block detection is used to obtain the paragraph text. At the same time, the table part is recognized by OCR, and the semantic information of the image is extracted. These text blocks, tables, and images are output as structured units and stored as JSON files.

[0131] For the HTML page "Instructions for Applying for xxxx" published on the official policy website, the computer device locates the article title, source organization, publication time, and main body of the text based on the preset DOM structure template, extracts them, and writes them into local JSON, which is then used as subsequent input for generating the tree structure.

[0132] By employing the above methods, not only can the original physical structure and layout information be preserved under different document formats, but special areas such as tables and formulas can also be distinguished. This avoids semantic information loss or structural disorder during the search enhancement and tree structure construction stages, thereby providing a clearer and more accurate structured data foundation for multi-document intelligent question answering, and greatly improving scalability and interpretability.

[0133] The following example illustrates the question-answering method based on structured document retrieval enhancement provided in this application. Figure 2 This is a schematic diagram of the retrieval information enhancement process provided in other embodiments of this application. For example... Figure 2 As shown, a computer device performing a question-answering method based on structured document retrieval enhancement may include the following steps:

[0134] Step 1: Parsing Phase

[0135] Computer equipment collects various structured policy documents issued by government agencies and uses appropriate text parsing technologies based on the document format, including chapter title recognition, clause number extraction, and hierarchical relationship annotation, to store this structured information as local structured document files for subsequent segmentation and modeling. This step ensures that subsequent analysis can accurately identify the logical boundaries and hierarchical structure of the documents.

[0136] Step 2: Segmentation Stage

[0137] Computer equipment uses predefined hierarchical numbering, heading formats, and semantic segmentation rules to accurately segment structured documents, identify the boundaries and hierarchical affiliations of each node, and break down long texts into the smallest independently processable semantic units. For extremely long text blocks, further segmentation is performed based on sentence boundaries or semantic features to form even smaller node units, thereby facilitating subsequent vectorized retrieval and summary generation.

[0138] Step 3: Tree Building Phase

[0139] After obtaining the segmentation results, the computer device constructs a tree structure based on the hierarchical characteristics of each segment unit, and attaches the document's chapter nodes, clause nodes, and paragraph nodes in the tree structure with parent-child relationships. In the case of multiple documents, the computer device combines the reference relationships between nodes and thematic similarity, and merges multiple tree structures through node mapping or supernode merging to form a cross-document tree structure network, and generates summary information for each non-leaf node and root node in it.

[0140] Step 4: Search Phase

[0141] The computer device utilizes tree-based hierarchical retrieval enhancement (RAG) technology. First, it expands and flattens the tree structure to obtain a set of searchable nodes. Then, using a hybrid strategy of semantic vector matching and keyword inverted index retrieval, it selects the Top-K nodes most relevant to the query. Subsequently, based on the hierarchical relationship of the tree structure, it traces back upwards or downwards to expand these nodes, supplementing the complete contextual information. This information is then combined into prompts and input into the question-answering model to generate a traceable answer to the user's query.

[0142] Step 1 specifically includes:

[0143] Step 1.1: For PDF and DOCX documents, the "layout analysis-OCR-formula / table detection" technical route is adopted. After the document preprocessing is completed, the physical layout elements in the document are detected first, and the structure and hierarchical relationship of the document are analyzed and understood. Then, the corresponding recognition models are used to extract the content of the text area, formula area and table area in the document respectively. Finally, the parsing results are saved to the local storage medium in Markdown format.

[0144] Step 1.2: For HTML and other web page content, first determine the target domain name based on the official website or data platform where the policy is released, clarify the content structure and data storage method of the target page, and then write and run the Scrapy crawler script to realize the automated collection of web page content, focusing on extracting key fields such as title, publishing organization, and body content, and saving the parsed results to local storage media in JSON format;

[0145] Step 1.3: After completing the above extraction work, use regular expressions and natural language processing techniques to clean up invalid information (redundant newlines and special symbols) in the text. At the same time, auxiliary content such as the table of contents is marked and separated. The corresponding content is only used to assist in the construction of the tree model and is not included in the node content.

[0146] Step 2 specifically includes:

[0147] Step 2.1: Scan the policy document line by line (if the document contains a table of contents, this can be simplified to scanning only the table of contents). Analyze and extract the hierarchical identifiers of headings and sections according to established rules, which will serve as the primary basis for document segmentation. Hierarchical identifiers include, but are not limited to, Arabic numerals (such as "1.", "1.1", "1.1.1"), Chinese serial numbers (such as "Chapter 1", "Article 2"), etc., and record the corresponding hierarchical structure.

[0148] Step 2.2: Based on the identified hierarchical identifiers, the sequence number format is parsed and matched using regular expressions to identify the dividing boundaries of each title or section, construct basic node units, and record the corresponding text block content;

[0149] Step 2.3: Set a text block length threshold T_max. Traverse all basic units of nodes to check if any text block length exceeds T_max. If so, split the node into multiple virtual nodes and further segment the text block content. Specifically, use "\n" as the first-level delimiter, ".", "?", "!", etc. as second-level delimiters, and "," "、", etc. as tertiary delimiters. Segment the long text block sequentially in descending order of priority until all sub-text blocks meet the length requirement.

[0150] Step 3 specifically includes:

[0151] Step 3.1: Construct a tree model based on the hierarchical relationships identified in Step 2.1 to achieve automatic nesting of the hierarchical structure. For example, "1.1" is nested under "1", and "1.1.1" is nested under "1.1". Record the original text block content on the leaf nodes, and generate a content summary for each non-leaf node from bottom to top based on the content of the lower-level nodes. Define the root node as the file layer, recording file metadata such as file title, issuing organization, and issuing time; define other nodes as the content layer.

[0152] Step 3.2: Above the file layer, a summary layer is constructed based on explicit citation relationships and topic similarity between files. Explicit citation relationships can be derived from semantic analysis of file content, while topic similarity is obtained through clustering using a Gaussian mixture model. After setting the initial number of clusters C in the Gaussian mixture model, the model parameters are optimized using the Expectation-Maximization (EM) algorithm to obtain the probability distribution of each text block belonging to each topic cluster. The cluster with the highest probability is taken as the primary affiliation of the text block, and similar text blocks are grouped and classified. Subsequently, a content summary is compiled for each node in the summary layer.

[0153] Step 3.3: Select an appropriate embedding model and word segmenter to generate text embedding and word segmentation results for the original text blocks in all leaf nodes and the summary information in non-leaf nodes, and store them on the corresponding nodes.

[0154] Figure 3 This is a schematic diagram of a tree structure provided in some embodiments of this application. For example... Figure 3 As shown, based on this application, a multi-level tree structure and cross-document clustering structure are constructed for structured documents, divided into three levels:

[0155] (1) Content layer

[0156] At the lowest level (content layer), computer devices break down the document content into multi-level nodes based on the chapter numbers, clause numbers, or paragraph numbers of the structured document:

[0157] First-level heading nodes (such as "Heading 1" and "Heading 2") represent major sections of a document;

[0158] Second-level heading nodes (such as "Second-level heading 1.1", "Second-level heading 1.2", etc.) represent clauses or sub-sections under first-level headings;

[0159] Level 3 content nodes (such as “Level 3 content 1.3.1”, “Level 3 content 1.3.2”, etc.) represent the smallest text unit, usually a paragraph or sub-clause.

[0160] When the length of the content of a third-level node exceeds the preset text block threshold (e.g., 1024 characters), the computer device can further split the node into multiple "virtual nodes" (as shown in "Virtual Node (I)" and "Virtual Node (II)") to ensure that the text size of each node can be efficiently processed by the question-answering model.

[0161] The ability to further segment excessively long child node content demonstrates that when the smallest level node (e.g., third-level content) is still too long, computer devices can further refine it based on sentence boundaries or semantic fragments, generating multiple virtual nodes and maintaining their hierarchical relationships within the tree structure to ensure that the depth and width of the overall tree structure can be flexibly adapted.

[0162] (2) File layer

[0163] At the file level, the computer device treats each structured document (files A1, A2, A3, B, and C in the example diagram) as the root node of a complete document tree, and attaches each node of the content layer to the root node of its respective document through a hierarchical structure to maintain the consistency of the parent-child structure.

[0164] Meanwhile, for reference relationships between multiple documents (such as a clause referencing a clause in another document), computer devices can establish cross-document node mapping connections by clustering related nodes in the file layer, forming a cross-document directional information tracing relationship.

[0165] (3) Summary layer

[0166] At the top level, the computer device performs topic analysis and vectorization on the root node or higher-level nodes of the document tree based on the similarity of topics among different files. It then uses topic similarity clustering algorithms to classify documents or nodes belonging to the same topic. As shown in the figure, firstly, based on the cross-document node mapping relationship, files A1, A2, and A3 are clustered using relation clustering, while files B and C each form their own cluster. Next, the computer device performs topic clustering, for example, merging files A1, A2, A3, and file B into one topic, while file C belongs to another topic. This facilitates access to all information on the same topic across multiple files at once when querying users, thus achieving integrated retrieval capabilities across documents and multiple perspectives.

[0167] Therefore, through the bidirectional organization of topic clustering and relation clustering, the tree structure not only supports fine-grained splitting and hierarchical summarization within structured documents, but also efficiently completes the integration, summarization and traceable citation of information across multiple documents, thereby meeting the multi-level retrieval enhancement needs of complex policy or normative texts.

[0168] Step 4 specifically includes:

[0169] Step 4.1: Expand the tree model constructed in Step 3 into a single-layer structure, and use a hybrid retrieval method that combines semantic retrieval (based on text embedding information) and keyword retrieval (based on word segmentation information) to obtain Top-K recall results. Based on the differences in the level of the recalled nodes, perform upward backtracking or downward extension to the file layer to obtain the corresponding file meta information.

[0170] Step 4.2: Select an appropriate reordering model, reorder the recall results, and then concatenate the prompt words, node content, and corresponding file meta information, and input them into the language model for content generation.

[0171] Therefore, by adopting a four-step framework of "parsing-blocking-tree building-retrieval", and combining rule-based blocking methods in the document preprocessing stage, we can ensure that the chapters, clauses and hierarchical relationships of policy documents are accurately parsed and preserved. This avoids the problem of destroying the document tree structure in the traditional RAG preprocessing process, and ensures that the model can accurately understand the hierarchical relationships and semantic dependencies between chapters, effectively improving the structured nature of document parsing.

[0172] Meanwhile, by constructing cross-document reference relationships and topic-level networks through supernodes, efficient modeling of policy document relationship networks is achieved, enabling the explicit expression of implicit semantic relationships between different policy documents and making up for the shortcomings of traditional RAG in integrating multi-document information.

[0173] Furthermore, by balancing the summary and detail of information through flattened recall during the retrieval stage, and combining it with a hierarchical backtracking strategy, the query results not only cover key content but also trace back to the original information of the document, ensuring the integrity and consistency of information recall and effectively improving the credibility of intelligent interpretation of policy documents.

[0174] The question-answering method based on structured document retrieval enhancement provided in this application can be executed by a question-answering device based on structured document retrieval enhancement. This application uses the execution of the question-answering method based on structured document retrieval enhancement by a question-answering device as an example to illustrate the question-answering device based on structured document retrieval enhancement provided in this application.

[0175] This application also provides a question-answering device based on structured document retrieval enhancement, which is applied to computer equipment.

[0176] Figure 4 This is a schematic diagram of the structure of a question-answering device based on structured document retrieval enhancement provided in some embodiments of this application. For example... Figure 4 As shown, the question-answering device based on structured document retrieval enhancement includes a construction module 401, a federation module 402, and a question-answering module 403. Wherein:

[0177] The construction module 401 is used to perform structured information parsing and chunking processing on structured documents to construct a tree structure that reflects the hierarchical relationship of the documents. In the tree structure, leaf nodes are used to represent semantic content based on chunking processing, and non-leaf nodes and root nodes are used to represent semantic content that summarizes the child nodes they include.

[0178] The union module 402 is used to unite the tree structures corresponding to multiple structured documents based on cross-document association relationships to form a multi-document tree structure network.

[0179] The question-answering module 403 is used to generate corresponding retrieval enhancement information through the tree structure network when a question request is received, and input the retrieval enhancement information into the question-answering model to obtain answer information that matches the question request.

[0180] According to the question-answering device based on structured document retrieval enhancement provided in this application embodiment, by performing structured information parsing and block processing on structured documents, chapter numbers, clause numbers, and title information can be retained in the document preprocessing stage, ensuring the consistency of the document's hierarchical structure in subsequent processing, thereby avoiding the context loss problem caused by traditional block splitting; in particular, by setting leaf nodes in the tree structure to carry original semantic units, and setting non-leaf nodes and root nodes to generate summary information for subordinate child nodes, a hierarchical information representation with layer-by-layer semantic summarization can be formed within the document, so that subsequent retrieval can obtain content combining details and summaries as needed, thereby improving the comprehensiveness of information; furthermore, by combining the tree structures of multiple structured documents based on cross-document association relationships, it can support the tracing of logical connections between multiple documents, improve the coverage of cross-document queries and the completeness of document interpretation, and prevent one-sided answers caused by isolated interpretation; furthermore, when a user's question request is received, the retrieval enhancement information returned through the tree structure network can simultaneously provide refined detailed evidence and multi-level context summary information in the question-answering model, thereby greatly improving the accuracy and traceability of question-answer generation.

[0181] In some embodiments, the union module is further configured to extract explicit reference relationships between structured documents, establish connections between relevant nodes in the tree structure of each document according to the reference relationships to form a cross-document node mapping relationship; determine nodes with similar topics in each structured document, perform topic clustering on the nodes with similar topics to form a document association relationship between multiple documents under the same topic; and form a tree structure network of multiple documents based on the node mapping relationship and the document association relationship.

[0182] In some embodiments, the joint module is further configured to perform text analysis on the main text content at preset positions in each structured document to detect the reference identifiers or link information contained therein; based on the reference identifiers or link information, locate the corresponding node in the referenced document, and establish a connection relationship between the reference node and the referenced node.

[0183] In some embodiments, the joint module is further configured to generate vector representations of the summary information corresponding to the root nodes of each structured document; cluster each vector representation using a Gaussian mixture model, and optimize it using an expectation-maximization algorithm under the condition of a set initial number of clusters to obtain the probability distribution of each topic cluster to which each node belongs; according to the probability distribution, the topic cluster with the highest probability is taken as the primary affiliation of the root node, and the root nodes in the same topic cluster are merged into the same supernode to achieve grouping and aggregation of root nodes with similar topics.

[0184] In some embodiments, the construction module is further configured to identify the hierarchical numbering information of chapters, clauses or paragraphs in a document based on predefined hierarchical identification rules; divide the document content into blocks according to the hierarchical numbering information, and generate leaf nodes and their parent-child relationships in a tree structure according to the block division results; and generate summary information from bottom to top according to the content of each leaf node and its parent-child relationship, so as to serve as the node content of non-leaf nodes and root nodes.

[0185] In some embodiments, the construction module is further configured to perform format matching on the hierarchical numbering information contained in each structured document to determine the boundary position of chapters or clauses; and to divide the document content into basic node units based on the boundary position, thereby forming leaf nodes and establishing their hierarchical correspondence with the parent nodes.

[0186] In some embodiments, the construction module is further configured to set a text block length threshold; traverse and check the content length of each node basic unit; when the text block length exceeds the text block length threshold, split the node into multiple virtual nodes to ensure that the content length of each leaf node meets the preset requirements.

[0187] In some embodiments, the construction module is further configured to generate corresponding summary information for each non-leaf node based on the text content of its direct subordinate child nodes using a preset summarization algorithm, and use the summary information as the node content of the non-leaf node; for the root node, generate overview summary information based on the summary information of all subordinate nodes and leaf node information, and use it as the node content of the root node; and generate corresponding text embedding vectors and word segmentation results for the original text blocks contained in each leaf node and the summary information contained in each non-leaf node through a text embedding model and a word segmenter, respectively, and store them on their respective nodes.

[0188] In some embodiments, the question-answering module is further configured to expand the tree structure network into a single-layer structure, and perform candidate recall of nodes in the tree structure network based on a hybrid retrieval method combining semantic information and keyword information to obtain one or more recalled nodes; perform backtracking of node information according to the hierarchical relationship of one or more recalled nodes to supplement their context information and associated file meta-information; apply a re-ranking model to optimize the sorting of the recalled node results, and combine the optimized nodes with the corresponding file meta-information to generate retrieval enhancement information for the question-answering model.

[0189] In some embodiments, the construction module is also used to detect the layout structure and physical layout elements of structured documents of the PDF or DOCX type, and extract the content of text areas, table areas and formula areas respectively to obtain multiple structured text units; for structured documents of the HTML type, according to the predetermined target page structure and data storage method, perform web page content collection, and extract field information such as title, publishing organization and body content to obtain multiple structured text units; save the parsed structured text units to local storage medium to assist in the construction of tree structure.

[0190] The question-answering device based on structured document retrieval enhancement in this application embodiment can be a computer device or a component within a computer device, such as an integrated circuit or a chip. The computer device can be a terminal device or a server. For example, the computer device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle computer device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.

[0191] The question-answering device based on structured document retrieval enhancement in this application embodiment can be a device with an operating system. This operating system can be a Microsoft (Windows) operating system, an Android operating system, an iOS operating system, or other possible operating systems; this application embodiment does not specifically limit it.

[0192] The question-answering device based on structured document retrieval enhancement provided in this application embodiment can realize the various processes implemented in the various method embodiments. To avoid repetition, it will not be described again here.

[0193] Figure 5 This is a schematic diagram of the structure of a computer device provided in some embodiments of this application. In some embodiments, such as Figure 5 As shown, this application embodiment also provides a computer device 500, including a processor 501, a memory 502, and a computer program stored in the memory 502 and executable on the processor 501. When the program is executed by the processor 501, it implements the various processes of the above-described method embodiments and can achieve the same technical effects. To avoid repetition, it will not be described again here.

[0194] It should be noted that the computer devices in this application embodiment include the mobile computer devices and non-mobile computer devices described above.

[0195] This application also provides a non-transitory computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described question-answering method embodiment based on structured document retrieval enhancement and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0196] The processor is the processor in the computer device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0197] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described question-answering method based on structured document retrieval enhancement.

[0198] The processor is the processor in the computer device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0199] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described question-answering method embodiment based on structured document retrieval enhancement, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0200] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0201] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0202] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0203] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0204] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0205] Unless otherwise specified, all embodiments and optional embodiments of this application can be combined to form new technical solutions.

[0206] Unless otherwise specified, all technical features and optional technical features of this application may be combined to form new technical solutions.

[0207] Unless otherwise specified, all steps of this application may be performed sequentially or randomly, preferably sequentially. For example, the method includes steps (a) and (b), indicating that the method may include steps (a) and (b) performed sequentially, or it may include steps (b) and (a) performed sequentially. For example, the mention that the method may also include step (c) indicates that step (c) may be added to the method in any order; for example, the method may include steps (a), (b), and (c), or it may include steps (a), (c), and (b), or it may include steps (c), (a), and (b), etc.

[0208] The above are merely preferred embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A question-answering method based on structured document retrieval enhancement, characterized in that, The method includes: Structured information parsing and chunking are performed on structured documents to construct a tree structure that reflects the hierarchical relationship of the documents. In the tree structure, leaf nodes are used to represent semantic content based on chunking, and non-leaf nodes and root nodes are used to represent semantic content that summarizes the child nodes they include. Based on cross-document relationships, the tree structures corresponding to multiple structured documents are combined to form a multi-document tree structure network. Upon receiving a question request, corresponding retrieval enhancement information is generated through the tree structure network, and the retrieval enhancement information is input into the question answering model to obtain answer information that matches the question request.

2. The method according to claim 1, characterized in that, The method of combining the tree structures corresponding to multiple structured documents based on cross-document association relationships to form a multi-document tree structure network includes: Extract the explicit reference relationships between the structured documents, and connect the relevant nodes in the tree structure of each document according to the reference relationships to form a cross-document node mapping relationship; Identify nodes with similar themes in each structured document, and cluster these nodes by theme to form document association relationships among multiple documents under the same theme; Based on the node mapping relationship and document association relationship, a tree structure network of multiple documents is formed.

3. The method according to claim 2, characterized in that, The step of extracting explicit reference relationships between structured documents and establishing connections between relevant nodes in the tree structure of each document according to the reference relationships to form a cross-document node mapping relationship includes: Text analysis is performed on the main text content at preset locations in each structured document to detect the reference identifiers or link information contained therein; Based on the reference identifier or link information, locate the corresponding node in the referenced document and establish a connection between the referenced node and the referenced node.

4. The method according to claim 2, characterized in that, The step of identifying nodes with similar topics in each structured document and clustering these nodes to form document associations among multiple documents under the same topic includes: Generate vector representations of the summary information corresponding to the root node of each structured document; A Gaussian mixture model is used to cluster the vector representations, and the expectation-maximization algorithm is used to optimize the clusters under the condition of a set initial number of clusters to obtain the probability distribution of each node to each topic cluster. Based on the probability distribution, the topic cluster with the highest probability is taken as the primary affiliation of the root node, and the root nodes in the same topic cluster are merged into the same super node to achieve grouping and aggregation of root nodes with similar topics.

5. The method according to claim 1, characterized in that, The process of performing structured information parsing and chunking on structured documents to construct a tree structure reflecting the hierarchical relationships within the documents includes: Based on predefined hierarchical identification rules, identify the hierarchical numbering information of chapters, clauses or paragraphs in a document; Based on the hierarchical numbering information, the document content is divided into blocks, and leaf nodes and their parent-child relationships in a tree structure are generated according to the block division results. Based on each leaf node and its parent-child relationship, summary information is generated from bottom to top according to the content of the leaf nodes, which serves as the node content for non-leaf nodes and the root node.

6. The method according to claim 5, characterized in that, The step of dividing the document content into blocks based on the hierarchical numbering information and generating leaf nodes and their parent-child relationships in a tree structure according to the block division results includes: The hierarchical numbering information contained in each structured document is formatted to determine the boundary positions of chapters or clauses; Based on the boundary positions, the document content is divided into basic node units, thereby forming leaf nodes and establishing their hierarchical correspondence with the parent nodes.

7. The method according to claim 6, characterized in that, The method further includes: Set a text block length threshold; The basic unit of each node is traversed and its content length is checked. When the length of a text block exceeds the text block length threshold, the node is split into multiple virtual nodes to ensure that the content length of each leaf node meets the preset requirements.

8. The method according to any one of claims 5 to 7, characterized in that, The process of generating summary information from bottom to top based on the content of each leaf node and its parent-child relationship, to serve as the node content for non-leaf nodes and the root node, includes: For each non-leaf node, based on the text content of its direct subordinate child nodes, a corresponding summary information is generated using a preset summary algorithm, and the summary information is used as the node content of the non-leaf node. For the root node, an overview summary is generated based on the summary information of all subordinate nodes and leaf node information, and this summary is used as the node content of the root node. By using a text embedding model and a word segmenter, corresponding text embedding vectors and word segmentation results are generated for the original text blocks contained in each leaf node and the summary information contained in the non-leaf nodes, and then stored on their respective nodes.

9. The method according to claim 1, characterized in that, The generation of corresponding retrieval enhancement information through the tree structure network includes: The tree structure network is expanded into a single-layer structure, and a hybrid retrieval method combining semantic information and keyword information is used to recall candidates for nodes in the tree structure network, resulting in one or more recalled nodes. Execute the backtracking of node information based on the hierarchical relationship of one or more recall nodes to supplement its context information and associated file meta information; The retrieved node results are sorted and optimized using a reordering model, and the optimized nodes are combined with the corresponding file metadata to generate retrieval enhancement information for the question-answering model.

10. The method according to claim 1, characterized in that, The process of parsing structured information from structured documents includes: For structured documents such as PDF or DOCX, the document's layout structure and physical layout elements are detected, and the content of text areas, table areas and formula areas are extracted to obtain multiple structured text units. For structured documents of the HTML type, based on the predetermined target page structure and data storage method, web page content is collected, and field information such as title, publishing organization and body content is extracted to obtain multiple structured text units; The parsed structured text units are saved to local storage media to assist in the construction of the tree structure.

Citation Information

Patent Citations

  • Hierarchical clustering method and system for mass document set

    CN106815310A

  • Multi-text information knowledge graph construction method based on tree structure

    CN115687650A

  • Cross-file question and answer knowledge extraction method and system and electronic equipment

    CN117851566A

  • Standard chapter splitting method

    CN119067097A

Cited By

  • Retrieval enhancement generation method and device based on content hierarchical weighting and storage medium

    CN121071060A

  • Content-based hierarchical weighted retrieval enhancement generation method, apparatus, and storage medium

    CN121071060B

  • Hierarchical retrieval enhancement generation method based on document citation graph, terminal and medium

    CN122509353A

  • Hierarchical retrieval enhancement generation method based on document citation graph, terminal and medium

    CN122509353B