A Multi-Layer Structure-Aware Retrieval Enhancement Method and System for Legal Documents

By constructing a hierarchical structure tree for legal documents and performing hybrid retrieval and noise filtering, and explicitly injecting regulatory timelines, the problem of insufficient structure-aware retrieval of legal documents is solved, and highly accurate and traceable retrieval enhancement generation is achieved.

CN122132427APending Publication Date: 2026-06-02INESA (GRP) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INESA (GRP) CO LTD
Filing Date
2026-03-16
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing retrieval enhancement generation systems cannot effectively perceive structural relationships when processing legal documents, resulting in semantic breaks and missing context. Noise content interferes with retrieval, and the inability to explicitly model the temporal relationships of regulations leads to unstable applicability judgments.

Method used

Construct a hierarchical structure tree of legal documents, perform hybrid retrieval and noise filtering, inject timeline information of regulations, and generate traceable answers with evidence citations.

Benefits of technology

It enables multi-level precise retrieval of legal documents, reduces document mismatch and chapter mismatch, improves the accuracy and applicability of retrieval, and meets the high requirements of rigor and traceability in legal application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122132427A_ABST
    Figure CN122132427A_ABST
Patent Text Reader

Abstract

This invention relates to a multi-layered structure-aware retrieval enhancement method and system for legal documents, comprising: filtering candidate legal documents based on legal identification information in the current requirement; calculating the relevance of each node to the current requirement by combining node information in the hierarchical structure tree with the candidate legal document set, and determining the chapter-level retrieval constraint range; performing a hybrid retrieval of keyword retrieval and semantic retrieval within the determined chapter-level retrieval constraint range, and performing score fusion and noise filtering to obtain highly relevant candidate clauses; performing context neighborhood completion, and maintaining timeline data for each legal document corresponding to the highly relevant candidate clauses; and assembling the model context as output according to a preset order. Compared with the prior art, this invention has the advantages of high accuracy in legal retrieval, strong rigor in generated content, high applicability and traceability, and strong universality and wide adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of large language model technology, and in particular to a multi-layer structure-aware retrieval enhancement generation method and system for legal documents. Background Technology

[0002] In intelligent question-answering and decision-making support scenarios for legal regulations, technical standards, and compliance documents, retrieval-enhanced generation techniques are often used to mitigate the shortcomings of large language models in terms of factual accuracy and traceability, such as generating illusions and missing evidence citations. Legal documents, with their strong normative, rigorous, and structured nature, are widely used in judicial decisions, corporate compliance, administrative supervision, and legal services, and their requirements for the accuracy, timeliness, and traceability of intelligent processing results are far higher than those for ordinary text. With the continuous growth in the number and increasing diversity of related documents, traditional retrieval-enhanced generation solutions are no longer sufficient to meet practical needs.

[0003] However, existing retrieval-enhanced generation systems still suffer from several prominent problems when processing legal documents. Legal documents have complex hierarchical structures, with clear distinctions between chapters, sections, articles, clauses, and items, and frequent cross-references and cross-paragraph connections. Existing systems often employ flattened text retrieval, failing to perceive structural relationships and easily leading to semantic breaks and missing context. Simultaneously, legal documents contain a significant amount of non-core noise content such as tables, appendices, page numbers, headers, and footers, which traditional retrieval mechanisms struggle to effectively distinguish, easily mistaking noise for valid information, crowding out model context space, and reducing retrieval and generation quality. Furthermore, most existing systems employ only single-layer retrieval strategies, either only locating at the document level with insufficient precision, or directly performing full-text clause searches, resulting in an overly broad scope and prone to document mismatches and chapter mismatches. More critically, the temporal relationships of regulations—such as promulgation, implementation, effectiveness, transition, replacement, and repeal—cannot be explicitly modeled and utilized, leading to unstable results when judging the applicability of regulations at specific points in time and the transition between old and new regulations, making applicability judgments unreliable.

[0004] Patent application CN120316120A discloses a table-based question-answering method based on table structure awareness and enhanced cell retrieval, including: table database construction, including constructing a pattern database and a cell database; query generation and expansion, including pattern query and cell query; information retrieval, including pattern retrieval and enhanced cell retrieval; and constructing prompts and performing language model inference, including integrating information, optimizing prompt format, and invoking language model inference. However, this method only designs cell and pattern retrieval for table scenarios and does not consider the hierarchical structure, noise interference, time applicability, and traceability requirements of complex structured legal documents, thus failing to address the pain points of intelligent legal question answering and compliance decision-making. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide a multi-layer structure-aware retrieval enhancement method and system for legal documents that has high accuracy in legal retrieval, high rigor in generated content, high applicability and traceability, and strong versatility and wide adaptability.

[0006] The objective of this invention can be achieved through the following technical solutions: A multi-layer structure-aware retrieval enhancement generation method for legal documents includes the following steps: Based on the regulatory identification information in the current requirements, a set of candidate regulatory documents is obtained, and the regulatory documents have a corresponding hierarchical structure tree. By combining the node information in the hierarchical structure tree with the candidate regulatory document set, the relevance of each node to the current requirement is calculated, and the scope of chapter-level retrieval constraints is determined. Within the defined chapter-level retrieval constraints, a hybrid retrieval combining keyword and semantic searches is performed, followed by score fusion and noise filtering to obtain highly relevant candidate terms. For the obtained highly relevant candidate clauses, context neighborhood completion is performed. During the completion process, deduplication and noise filtering are performed simultaneously. At the same time, timeline data is maintained for each regulatory document corresponding to the highly relevant candidate clauses, and timeline fact blocks are generated in a formatted manner based on the timeline data. The model context is assembled as output according to a preset order. The model context includes at least key structure segments, clause hit segments, neighborhood completion segments, relevant summary segments, and timeline fact blocks.

[0007] Furthermore, the construction of the hierarchical structure tree specifically includes: The original legal documents are processed for text normalization and noise removal. A hierarchical structure tree is constructed based on the title level and clause numbering pattern. The hierarchical structure tree includes node level, numbering information, text range and chapter path.

[0008] Furthermore, the text normalization and noise removal process includes standardizing blank lines, repairing hyphens, and cleaning up page number and header / footer noise.

[0009] Furthermore, the process of filtering candidate regulatory documents based on regulatory identification information in the current requirements specifically includes: Based on the constructed hierarchical structure tree, the regulatory identification information is subjected to synonym expansion and signal weighting, and a candidate regulatory document set is obtained by matching the pre-constructed number and alias index. In the signal weighting process, the regulation number, document number, and stable alias are set as strong positioning signals, while the region, institution, and generalized high-frequency words are set as weak positioning signals and are subject to weight reduction.

[0010] Furthermore, when generating the candidate regulatory document set, a supplementary recall is performed based on the summary information in the current requirements, and the matched regulatory documents are added to the candidate regulatory document set.

[0011] Furthermore, after the scope of the chapter-level retrieval constraints is determined, parent nodes or adjacent nodes are added according to preset rules, and a priority retention strategy is implemented for the introduction, preface, scope of application and terminology definition nodes.

[0012] Furthermore, in the hybrid retrieval, the text block vector is calculated offline and persistently stored during the database construction phase, while the query vector is calculated online during the question-and-answer phase. Noise filtering and term weighting constraints are uniformly performed before and after the retrieval.

[0013] Furthermore, the context neighborhood completion is only triggered for clauses whose scores exceed a preset threshold, and deduplication and noise filtering are performed during the completion process.

[0014] Furthermore, the timeline data of the regulations shall at least include the publication time, implementation time, effective time, transition rules, and replacement or repeal relationships.

[0015] This invention also provides a structure-aware retrieval and enhancement generation system for legal documents, comprising: The structure parsing module is used to standardize legal documents, remove noise, and build a hierarchical structure tree. The document filtering module is used to filter candidate regulatory documents based on regulatory identification information and signal weights to obtain a set of candidate regulatory documents. The chapter constraint module is used to determine the scope of chapter-level search constraints based on the hierarchical structure tree. The hybrid retrieval module is used to perform keyword and semantic hybrid retrieval within the scope of the chapter-level retrieval constraints and output highly relevant candidate terms; The neighborhood completion module is used to complete the neighborhood text of the highly relevant candidate clauses; The timeline injection module is used to maintain and inject regulatory timeline fact data; The answer generation module is used to assemble contextual information and generate constrained answers with evidence citations.

[0016] Furthermore, it also includes a noise filtering module, which is used to identify and filter noisy content such as directories, tables, and appendices, and to uniformly perform noise suppression during the document database creation, retrieval, and context completion stages.

[0017] Furthermore, the answer output by the answer generation module carries document identifiers and chapter path metadata, enabling traceable location of answer evidence.

[0018] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention performs structured parsing of legal documents to construct a hierarchical structure tree, and sequentially performs document-level positioning, chapter-level scope constraints, and clause-level hybrid retrieval to achieve multi-level, strongly constrained, and accurate retrieval. This significantly reduces problems such as cross-documentation of regulations, chapter mismatch, and interference from irrelevant text, thereby improving the accuracy and reliability of retrieval.

[0019] 2. By explicitly injecting timeline information such as the release, implementation, effectiveness, transition, replacement, and repeal of regulations during the question-and-answer phase, this invention provides clear temporal semantic constraints for the model, enabling it to stably support complex compliance scenarios such as applicability judgments at specific points in time and comparisons between new and old regulations, thereby significantly improving the timeliness and applicability of the answers.

[0020] 3. This invention employs a multi-stage unified noise suppression mechanism to identify, filter, or downgrade low-value content such as tables, appendices, headers, and footers during document database creation, retrieval, and context completion processes. This reduces noise interference with the retrieval and generation process, improves context utilization efficiency, and enhances the standardization of output content.

[0021] 4. This invention completes the contextual neighborhood of highly relevant candidate clauses and retains traceable meta-information such as chapter paths and document identifiers, so that the generated answer has complete contextual support and clear evidence sources, making the answer interpretable, verifiable and auditable, and meeting the high requirements of rigor and traceability in legal application scenarios.

[0022] 5. This invention adopts a framework that combines structure awareness, hierarchical constraints, and hybrid retrieval. It is not only applicable to laws and regulations, but can also be extended to various structured professional documents such as technical standards, industry norms, and policy documents. It has strong versatility and scalability, low adaptation cost, and wide applicability. Attached Figure Description

[0023] Figure 1 This is a flowchart of the present invention; Figure 2 This is a block diagram of the present invention. Detailed Implementation

[0024] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0025] Example 1 This embodiment provides a multi-layered structure-aware retrieval enhancement generation method for legal documents. First, the original legal documents are standardized and noise-removed. A structure tree is constructed by integrating title hierarchy and clause number features, outputting structured data. Then, regulatory identification information is extracted from user questions. A candidate set of regulatory documents is selected through signal hierarchical weighting, initially narrowing the search scope. Based on the structure tree, the relevance of each chapter to the user question is calculated, determining and outputting chapter search constraints to further narrow the search space. Subsequently, keyword retrieval and semantic retrieval are performed simultaneously. The search results are integrated to output highly relevant clauses. Neighborhood completion is performed on highly relevant clauses to improve their contextual relevance. A timeline is injected into the model context to provide timeline constraints. Finally, all information is integrated to generate the answer. This method, through explicit document structure modeling, hierarchical constraint retrieval, timeline semantic injection, and noise suppression mechanisms, can improve the retrieval accuracy and generation credibility of complex legal questions. This method is particularly suitable for high-precision scenarios such as regulatory interpretation, compliance review, standard comparison, and time applicability judgment.

[0026] like Figure 1 As shown, the method includes the following steps: S1. Structured parsing of legal documents: The original legal documents are processed for text standardization and noise removal. A hierarchical structure tree is constructed based on the title level and clause numbering pattern. The hierarchical structure tree includes node level, numbering information, text range and chapter path.

[0027] Specifically, in one implementation, the text representation of the original legal document is obtained (e.g., converting PDF, DOCX, etc., into a text format with hierarchical tags), and then normalized, including whitespace and line break normalization, line break and hyphen correction, and noise removal for page numbers and determinate headers and footers. Subsequently, based on the title hierarchy tags and clause numbering patterns, the structural boundaries of chapters, sections, and articles are identified, and a hierarchical structure tree containing node levels, titles / numbers, text ranges, and chapter paths is constructed. The structured result (e.g., a structure tree file and its flattened node table) is output for subsequent scope constraints and evidence location.

[0028] S2. Document-level candidate regulation location: Based on current needs, such as the regulation identification information in user questions, a set of candidate regulation documents is obtained.

[0029] Specifically, in one implementation, candidate regulatory documents are screened based on regulatory identifiers and summary information in the user's question: First, the question is expanded using abbreviation / alias mapping to extract strong positioning signals such as regulatory number, document number, stable aliases, and key phrases in the title, while weak signals such as region, institution, and generalized high-frequency words are weighted less. Then, a pre-constructed number / alias mapping index is used for matching and recall. When strong signals are insufficient or ambiguous, further recall and re-ranking can be performed based on document summary or chapter summary. Finally, a set of candidate regulatory documents and the criteria for their selection are obtained, providing core input for the precise limitation of the subsequent search scope.

[0030] S3. Chapter-level search scope constraint: Combining the node information in the hierarchical structure tree with the candidate regulatory document set, calculate the relevance of each node to the current requirement and determine the chapter-level search constraint scope.

[0031] Specifically, in one implementation, for candidate regulatory documents, relevance is calculated based on the titles and summaries of nodes in the tree structure, and the most relevant chapters or clauses are selected as search constraints; parent nodes or adjacent nodes can be supplemented according to preset rules to avoid overly narrow constraints. Meanwhile, a priority retention strategy can be set for key sections such as the introduction, preface, scope of application, and terminology definitions to improve the completeness of subsequent explanations.

[0032] S4. Clause-level hybrid retrieval and filtering: Within the defined chapter-level retrieval constraints, perform a hybrid retrieval of keyword retrieval and semantic retrieval, and perform score fusion and noise filtering to obtain highly relevant candidate clauses.

[0033] Specifically, in one implementation, a hybrid retrieval is performed on the clause text blocks within the aforementioned constraints: on the one hand, keyword retrieval (e.g., BM25) is performed, and on the other hand, semantic retrieval (e.g., based on a vector index library) is performed. The text block vectors can be calculated offline and persisted during the database construction phase, while the query vectors are calculated online during the question-answering phase. Noise filtering and necessary term constraints / weighting are applied uniformly before and after the retrieval. Finally, the scores from both approaches are normalized and fused to output a highly relevant candidate sequence of clauses, providing core material for subsequent contextual completion.

[0034] By using the document filtering, chapter constraints, and hybrid search methods described above, the search scope is gradually narrowed, thus solving the problem of low accuracy in existing single-layer searches.

[0035] S5. Context completion for highly relevant candidate clauses: For the obtained highly relevant candidate clauses, context neighborhood completion is performed. During the completion process, deduplication and noise filtering are performed simultaneously.

[0036] Specifically, in one implementation, for the aforementioned output sequence of highly relevant clause candidates, neighborhood completion is performed only on high-confidence clause text blocks with a hit score exceeding a preset threshold: when the hit score exceeds the preset threshold, adjacent text blocks before and after it are added in a configurable number, and mandatory retention blocks such as introductions and prefaces are also completed as necessary. During the completion process, deduplication and noise filtering are performed to ensure contextual coherence and prevent the introduction of low-value content, thereby improving the completeness of clause interpretation.

[0037] S6. Regulatory Timeline Information Injection: While performing the above context completion steps, maintain timeline data for each regulatory document corresponding to highly relevant candidate clauses, and generate timeline fact blocks based on the timeline data in a formatted manner.

[0038] Specifically, in one implementation, timeline data is maintained for each regulation, including at least fields such as publication, implementation, effectiveness, transition, and replacement or repeal; during the question-and-answer phase, the timeline is loaded and formatted as a timeline fact block, which is injected into the front area of ​​the model context to provide explicit time constraints for questions such as whether a certain point in time is applicable and the connection between old and new rules, thus solving the problems of existing systems being unable to model time relationships and unstable applicability judgments.

[0039] S7. Generation of answers to structural constraints: Assemble the model context as output according to a preset order. The model context includes at least key structural segments, clause hit segments, neighborhood completion segments, relevant summary segments, and timeline fact blocks.

[0040] Specifically, in one implementation, the outputs of all the aforementioned steps are integrated, combining key structural segments, clause-hitting segments, neighborhood completion segments, relevant summary segments, and timeline fact blocks. A model context is constructed by carrying traceable meta-information such as document identifiers and chapter paths. The output structure is then generated according to a preset prompt structure constraint language model, requiring the main rule to be given first, followed by exceptions, and key conclusions to be accompanied by corresponding evidence citations, thereby generating an interpretable and auditable answer. By binding evidence citations to meta-information such as chapter paths and document identifiers, prioritizing the main rule before exceptions, and attaching evidence to key conclusions, the problem of existing systems generating content without traceability and unverifiable results is solved.

[0041] Example 2 This embodiment provides a structure-aware retrieval and enhanced generation system for legal documents, such as... Figure 2As shown, the system includes a structure parsing module 1, a document filtering module 2, a chapter constraint module 3, a hybrid retrieval module 4, a neighborhood completion module 6, a timeline injection module 7, and an answer generation module 8. Specifically, the structure parsing module 1 is used to standardize legal documents, remove noise, and construct a hierarchical structure tree; the document filtering module 2 is used to filter candidate legal documents based on legal identifier information and signal weights to obtain a candidate legal document set; the chapter constraint module 3 is used to determine the chapter-level retrieval constraint range based on the hierarchical structure tree; the hybrid retrieval module 4 is used to perform keyword and semantic hybrid retrieval within the chapter-level retrieval constraint range and output highly relevant candidate clauses; the neighborhood completion module 6 is used to complete the neighboring text of the highly relevant candidate clauses; the timeline injection module 7 is used to maintain and inject legal timeline factual data; and the answer generation module 8 is used to assemble contextual information and generate constrained answers with evidence citations.

[0042] In a preferred embodiment, a noise filtering module 5 is also included, which is used to identify and filter noisy content such as directories, tables, and appendices, and to uniformly perform noise suppression during the document database creation, retrieval, and context completion stages.

[0043] The modules in this embodiment are described in detail below.

[0044] 1) Structure Parsing Module: This module performs structured processing on the original legal documents, achieving accurate extraction and standardized output of the document structure. It provides a unified and stable structural benchmark for the entire system, enabling page indexing and table of contents detection. This module provides a clear target for subsequent noise suppression by locating table of contents intervals; it accurately divides structural nodes such as chapters and clauses through joint parsing of titles and numbers, reducing structural mis-segmentation; it constructs a chapter tree that conforms to the hierarchical logic of legal documents through stack-based tree building, clearly presenting the internal structural relationships of the documents; it establishes a mapping relationship between text blocks and chapters through range calculation, providing support for evidence tracing; it avoids local parsing anomalies affecting the overall structural usability through consistency checks and fallback strategies; and it provides a unified data source for subsequent modules such as range constraints, summary injection, and evidence path display by outputting a structure tree and flattened results, ensuring the coordinated operation of all modules.

[0045] The specific functions of the structure parsing module include: Table of Contents Paragraph Location: Identify markers such as "Table of Contents" and patterns such as dotted lines and page number lists in the text representation to locate the table of contents range for subsequent filtering or demotion.

[0046] Joint parsing of titles and numbers: combining title level markers and numbering patterns to determine node boundaries reduces erroneous segmentation caused by relying on a single feature.

[0047] Stack-based tree construction: The stack structure is maintained according to the level depth. When a deeper level is encountered, it is pushed onto the stack, and when it is rolled back, it is popped off the stack, forming a stable chapter tree.

[0048] Range calculation: Calculate the text range (e.g., start and end line numbers or offset) covered by each structure node and generate a mapping that can be used to bind chapters to text blocks.

[0049] Consistency check and fallback: Detect situations such as excessive hierarchical jumps, nodes that are too short / too long, and abnormally repeated titles, and trigger fallback strategies (such as downgrading to a shallower level or retaining only chapters / sections).

[0050] Results Output and Reuse: Output structure tree files and flattened node results as a unified source for range constraints, summary injection, and evidence path display.

[0051] 2) Document Filtering Module: This module quickly and accurately filters candidate regulatory documents related to user questions, narrowing the subsequent search scope and improving search efficiency and accuracy. It achieves rapid matching of regulatory identifier information by building an offline number / alias mapping index; it improves matching accuracy by eliminating format differences and ambiguities in the query and document identifiers through normalization and disambiguation processing; it expands the search dimensions of user questions through query enhancement, supplementing synonym and alias information to avoid missed detections due to differences in expression; it highlights the role of strong positioning signals and reduces interference from weak signals through evidence hierarchical weighting, improving the stability of candidate filtering; it solves the problem of missing candidates in scenarios with insufficient strong signals or ambiguity through a summary backtracking recall strategy; and it integrates multiple candidate results and performs standardized sorting through candidate aggregation and gating, outputting the optimal candidate document set to provide reliable input for subsequent chapter-level searches.

[0052] The document filtering module's functions specifically include: Offline build number / alias mapping: Extract regulatory number, standard number and alias from document title, file name and metadata, generate number index and alias index, and invert the mapping to document identifier.

[0053] Standardization and deambiguity: Implement full-width and half-width character unification and symbol standardization; ensure continuous numbers are not truncated; remove short words and highly generalized words; mark regional / institutional words as weak evidence and participate in weight reduction.

[0054] Query Enhancement: Load the abbreviation / alias mapping table, expand the query with abbreviations and append aliases to form an enhanced query, and generate an optional set of required terms.

[0055] Evidence grading and weighting: Different weights are assigned to different types of evidence to improve the stability of candidate screening.

[0056] — Strong evidence: High weight is given when the regulation number, document number, or stable alias is matched.

[0057] — Medium evidence: When a title segment hits, it is given medium weight.

[0058] — Weak evidence: Keywords, regional / institutional terms, etc. are given low weight and can be further reduced in weight.

[0059] Abstract rollback recall: When the numbering mapping is insufficient or ambiguous, perform lightweight recall (e.g., keyword or semantic matching) based on document or chapter summary to supplement candidate documents.

[0060] Candidate aggregation and gating: Merge multiple candidates and normalize the scores to output TopK candidate documents; when the differences at the top are not obvious, the range is widened according to preset rules or a fallback strategy is triggered.

[0061] 3) Chapter Constraint Module: This module enables chapter summary matching, accurately locating the chapters relevant to the user's question within candidate regulatory documents. This further narrows the search space, avoids interference from irrelevant chapters, and ensures the accuracy of subsequent searches. This module constructs a chapter candidate pool, providing a unified matching framework for chapter relevance assessment; it quantifies the correlation between each chapter and the user's question through chapter relevance calculation, achieving precise chapter location; it supplements relevant nodes through scope expansion rules, avoiding evidence loss due to an overly narrow search scope; it improves the consistency of answers through a key paragraph priority strategy; it clarifies the text scope of subsequent searches through constraint output, providing explicit search constraints for the hybrid search module; and it employs a fallback strategy to avoid no results due to overly strict constraints, ensuring a smooth search process.

[0062] The chapter constraint module specifically includes the following functions: Construct a chapter candidate pool: Based on the structure node table, prepare matching text for each node, consisting of "node title + node summary (if any)".

[0063] Chapter relevance calculation: Node scores are calculated based on the hit rate of title keywords and the coverage rate of abstract terms to achieve controllable chapter positioning.

[0064] Range expansion rules: After a node is hit, the parent node, adjacent nodes, or child nodes are added according to the rules to avoid the lack of evidence due to an overly narrow range.

[0065] Prioritize key paragraphs: Set priorities or force the retention of nodes such as the introduction, preface, definitions and scope of application to improve the consistency of interpretation.

[0066] Constraint output: Generates a set of text blocks or chapter paths that are allowed to be searched, which will be used as filtering conditions for subsequent searches.

[0067] Narrow fallback: When the number of searchable text blocks after constraint is lower than the threshold, automatically expand to the next level node or broaden to the entire document.

[0068] 4) Hybrid Search Module: Within the constraints of chapters, this module accurately retrieves clause text blocks highly relevant to the user's question, providing core evidence to support answer generation. This module rewrites search queries, breaks down question keywords, and allocates weights appropriately to avoid irrelevant words interfering with clause ranking; it quickly matches clauses containing core keywords through keyword retrieval, ensuring search efficiency; it reuses offline vector semantic retrieval to reduce repetitive encoding operations, improving search speed while capturing semantic relationships in the text, avoiding the limitations of keyword matching; it integrates two search results through score normalization and fusion, balancing accuracy and comprehensiveness while removing duplicate content to improve the quality of search results; it strengthens the role of essential terms through hard constraints and soft weighting, increasing the probability of hitting key clauses; and it avoids excessive concentration of search results through result quotas and diversity control, ensuring the comprehensiveness of evidence.

[0069] The functions of the hybrid search module specifically include: Search query rewriting: The question is broken down into document keywords, time keywords, and clause subject keywords. During the clause retrieval stage, keywords and time keywords are downgraded or removed to avoid them interfering with the clause ranking.

[0070] Keyword search: Perform keyword search within a limited set of document / chapter text blocks, and apply noise filtering rules consistent with those used in database creation to maintain consistency.

[0071] Semantic retrieval (reusing offline vectors): Text block vectors are pre-written into the vector index during the database construction phase; during the question answering phase, only the query vector is calculated and retrieved, avoiding repeated encoding of text blocks.

[0072] Score normalization and fusion: Keyword scores and semantic similarity are normalized separately and then weighted and fused, and duplicate text blocks are deduplicated and merged.

[0073] Hard constraints and soft weighting: Filtering or weighting the set of required terms to improve the stability of key terms hitting.

[0074] Results quota and diversity control: Limit the proportion of a single document or chapter when necessary to avoid the results being over-consumed by a single paragraph.

[0075] 5) Noise Filtering Module: This module filters directories and tables, identifying and suppressing noisy content in legal documents throughout the entire process. This reduces interference from invalid information in retrieval and answer generation, improving the accuracy and stability of system output. The module specifically identifies noise in directories, tables, appendices, etc., accurately marking various non-core semantic content. Through a false positive protection strategy, it avoids misjudging text blocks carrying core legal semantics as noise, ensuring no loss of core information. By uniformly applying noise rules in the three key stages of database construction, retrieval, and context completion, it forms a closed-loop noise suppression system, comprehensively cleaning up invalid information. During database construction, it removes or downweights noise to ensure database quality; during retrieval, it filters or downweights noise to improve the purity of retrieval results; and during context completion, it prohibits noise completion to avoid introducing invalid content. Through audit output, it statistically analyzes the noise processing effect and provides optimization directions, facilitating iterative improvement of module functions and ensuring the reliability of noise suppression.

[0076] The noise filtering module specifically includes the following functions: Noise identification in the table of contents: Identify patterns such as dotted line guides, consecutive page number lists, densely arranged chapter titles, and "Table of Contents" and mark them as noise content.

[0077] Table noise identification: Identify morphological features such as Markdown table separators, HTML table remnants, and dense column alignment symbols, and mark them as noise content.

[0078] Appendix / Guide Page Recognition: Set separate rules for Annex / Appendices, Terminology Index, Cover Information Page, etc., to reduce the probability of accidentally entering the main text search.

[0079] False positive protection: When a text block contains normative expressions such as obligations, prohibitions, and conditions, its retention priority is increased to reduce the probability of it being falsely judged by noise rules.

[0080] Multi-stage application: Noise rules can be executed uniformly in the database construction, retrieval and neighborhood completion stages to form a closed-loop suppression.

[0081] — Database building phase: Remove or mark noisy text blocks as having lower weight.

[0082] — Retrieval phase: Filter or reduce the weight of noisy text blocks.

[0083] — Completion stage: Noise text blocks are prohibited from being used as neighborhood completion content.

[0084] Audit output: Statistical analysis of noise percentage and sampling output of suspected misjudgments to facilitate iterative rule adjustments.

[0085] 6) Neighborhood Completion Module: This module completes the context of high-confidence hit clauses, improving the contextual relationships between clauses and enhancing the completeness and accuracy of the answer explanation. This module uses trigger conditions to complete only high-confidence clauses, avoiding contextual redundancy caused by indiscriminate completion; it defines adjacency relationships to ensure the relevance between the completed content and the hit clauses, guaranteeing contextual coherence; it controls the number of completed clauses to avoid information redundancy caused by excessive completion, balancing completeness and conciseness; it eliminates invalid and repetitive information in the completed content through noise and duplication suppression, improving completion quality; it links key paragraphs to complete key paragraphs such as introductions and prefaces, ensuring that the necessary preconditions for the answer explanation are not missing; and it records relevant information during the completion process through debugging and visualization, providing support for module parameter optimization, ensuring the stability and reliability of the completion function, and helping the answer generation module output more complete and rigorous explanations.

[0086] The specific functions of the neighborhood completion module include: Triggering conditions: Expansion is only triggered for high-confidence text blocks in the search results whose scores exceed a preset threshold, to avoid indiscriminate completion that leads to context bloat.

[0087] Adjacency definition: Selecting adjacent text blocks based on their natural order within the same document and chapter path, and restricting expansion to the same clause if necessary.

[0088] Expanded quantity control: Set a configurable number of neighbors for each hit point and set a global limit to prevent multiple hits from accumulating and getting out of control.

[0089] Noise and Repetition Suppression: Noise filtering is performed again before adding neighbors; duplicate text blocks are deduplicated; low-information-density neighbors can be downgraded or skipped.

[0090] Key paragraph linkage: Necessary completion can also be made for mandatory reserved blocks such as introductions and prefaces to ensure that definitions, scope and exception preconditions are not missing.

[0091] Debugging and visualization: Records the completion range, completion reason and context usage of the hit point, and supports parameter tuning and regression verification.

[0092] 7) Timeline Injection Module: Maintains timeline data for each regulation, including at least information related to the regulation's publication, implementation, effectiveness, transition, replacement, or repeal. Before generating the answer, the timeline data is formatted into standardized fact blocks and injected into the model context.

[0093] 8) Answer generation module: Integrates key information from the hierarchical structure tree, highly relevant clauses, context completion content, and timeline fact blocks to assemble a model context that meets the requirements. At the same time, it controls the output format through constraint rules to generate standardized answers with accompanying evidence citations and traceability.

[0094] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A multi-layer structure-aware retrieval enhancement generation method for legal documents, characterized in that, Includes the following steps: Based on the regulatory identification information in the current requirements, a set of candidate regulatory documents is obtained, and the regulatory documents have a corresponding hierarchical structure tree. By combining the node information in the hierarchical structure tree with the candidate regulatory document set, the relevance of each node to the current requirement is calculated, and the scope of chapter-level retrieval constraints is determined. Within the defined chapter-level retrieval constraints, a hybrid retrieval combining keyword and semantic searches is performed, followed by score fusion and noise filtering to obtain highly relevant candidate terms. For the obtained highly relevant candidate clauses, context neighborhood completion is performed. During the completion process, deduplication and noise filtering are performed simultaneously. At the same time, timeline data is maintained for each regulatory document corresponding to the highly relevant candidate clauses, and timeline fact blocks are generated in a formatted manner based on the timeline data. The model context is assembled as output according to a preset order. The model context includes at least key structure segments, clause hit segments, neighborhood completion segments, relevant summary segments, and timeline fact blocks.

2. The multi-layer structure-aware retrieval enhancement generation method for legal documents according to claim 1, characterized in that, The construction of the hierarchical structure tree specifically includes: The original legal documents are processed for text normalization and noise removal. A hierarchical structure tree is constructed based on the title level and clause numbering pattern. The hierarchical structure tree includes node level, numbering information, text range and chapter path.

3. The multi-layer structure-aware retrieval enhancement generation method for legal documents according to claim 2, characterized in that, The text normalization and noise removal process includes standardizing blank lines, fixing hyphens, and cleaning up page number and header / footer noise.

4. The multi-layer structure-aware retrieval enhancement generation method for legal documents according to claim 1, characterized in that, The candidate regulatory document set obtained by filtering based on the regulatory identification information in the current requirements specifically includes: Based on the constructed hierarchical structure tree, the regulatory identification information is subjected to synonym expansion and signal weighting, and a candidate regulatory document set is obtained by matching the pre-constructed number and alias index. In the signal weighting process, the regulation number, document number, and stable alias are set as strong positioning signals, while the region, institution, and generalized high-frequency words are set as weak positioning signals and are subject to weight reduction.

5. The multi-layer structure-aware retrieval enhancement generation method for legal documents according to claim 1, characterized in that, When generating the candidate regulatory document set, a supplementary recall is performed based on the summary information in the current requirements, and the matched regulatory documents are added to the candidate regulatory document set.

6. The multi-layer structure-aware retrieval enhancement generation method for legal documents according to claim 1, characterized in that, Once the scope of the chapter-level search constraints is determined, parent nodes or adjacent nodes are added according to preset rules, and a priority retention strategy is implemented for the introduction, preface, scope of application and terminology definition nodes.

7. The multi-layer structure-aware retrieval enhancement generation method for legal documents according to claim 1, characterized in that, In the hybrid retrieval, text block vectors are calculated offline and persistently stored during the database construction phase, while query vectors are calculated online during the question-and-answer phase. Noise filtering and term weighting constraints are uniformly performed before and after the retrieval.

8. The multi-layer structure-aware retrieval enhancement generation method for legal documents according to claim 1, characterized in that, The timeline data for the regulations should include at least the publication date, implementation date, effective date, transitional rules, and replacement or repeal relationships.

9. A structure-aware retrieval and enhanced generation system for legal documents, characterized in that, include: The structure parsing module is used to standardize legal documents, remove noise, and build a hierarchical structure tree. The document filtering module is used to filter candidate regulatory documents based on regulatory identification information and signal weights to obtain a set of candidate regulatory documents. The chapter constraint module is used to determine the scope of chapter-level search constraints based on the hierarchical structure tree. The hybrid retrieval module is used to perform keyword and semantic hybrid retrieval within the scope of the chapter-level retrieval constraints and output highly relevant candidate terms; The neighborhood completion module is used to complete the neighborhood text of the highly relevant candidate clauses; The timeline injection module is used to maintain and inject regulatory timeline fact data; The answer generation module is used to assemble contextual information and generate constrained answers with evidence citations.

10. A structure-aware retrieval and enhancement generation system for legal documents according to claim 9, characterized in that, It also includes a noise filtering module, which is used to identify and filter noisy content such as directories, tables, and appendices, and to uniformly perform noise suppression during the document database creation, retrieval, and context completion stages.

Citation Information

Patent Citations

  • Table question and answer method and system based on table structure perception and cell retrieval enhancement

    CN120316120A