Clause generation and auditing reporting method and device, equipment and medium
By receiving demand information, analyzing key elements, generating draft terms, establishing relationships, and performing compliance checks and revisions, this technology solves the problem of relying on manual operation for term generation and review in existing technologies. It achieves automated and closed-loop management of term generation and review in the fintech and healthcare fields, improving efficiency and accuracy.
Patent Information
- Application Number
- CN202511800759.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-02
- Publication Date
- 2026-03-17
AI Technical Summary
In existing technologies, the generation and review of terms in the fintech and healthcare sectors rely on manual operation, resulting in low efficiency, high error rates, and difficulty in version tracing, as well as a lack of a unified automated and closed-loop management mechanism.
This paper provides a method for generating, reviewing, and filing clauses, which includes receiving requirement input information and a set of supporting materials, parsing key elements, generating draft clauses, establishing relationships, performing compliance checks, generating a review report, performing revision operations, and finally generating a filing material package.
It has achieved fully automated processing of the entire process from the generation of clause documents to their filing, reducing manual intervention and repetitive operations, improving the efficiency and accuracy of clause drafting and review, and ensuring the traceability of versions and the standardization of filing materials.
Smart Images

Figure CN121683718A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, equipment, and medium for generating, reviewing, and filing clauses. Background Technology
[0002] In the fintech sector, the drafting and compliance management of insurance terms or financial contracts often rely on manual processes, which presents significant shortcomings. First, the drafting process is inefficient. Product developers need to manually review policy documents, historical contracts, and actuarial data, which is not only time-consuming but also prone to inconsistencies in format and high repetition rates, lacking intelligent processing support. Second, the review process relies excessively on manual comparison of regulatory rules and business specifications, resulting in a massive workload. Furthermore, in scenarios with frequent clause revisions, manual review is prone to overlooking key differences, leading to compliance risks. Third, existing tools suffer from inadequate interactive design. Users typically need to switch between multiple interfaces or pop-ups, hindering efficient multitasking. The management and correlation analysis capabilities for attachments are limited, failing to meet the complex compliance review requirements of financial terms. In addition, the generation and revision of terms generally lack robust version control mechanisms, making it difficult to trace modification history, which is detrimental to compliance audits and liability determination, thus impacting the reporting and launch processes of financial products.
[0003] Similar shortcomings exist in the healthcare sector. Documents such as medical service terms, clinical informed consent forms, and health insurance contracts typically require manual drafting, incorporating policy regulations, medical guidelines, and historical cases. This process is inefficient, and ensuring consistency in formatting and logical integrity is difficult. Document review largely relies on manual comparison of each item against medical standards and compliance requirements. With frequent revisions or multiple versions running concurrently, key changes are easily overlooked, or work is often duplicated. Existing systems generally offer poor user experience, often using window-based interfaces that fail to effectively support collaboration among doctors, compliance personnel, and administrators. Furthermore, the lack of a systematic version management and modification log mechanism makes it difficult to accurately track and verify document changes, hindering medical compliance checks and accountability. Finally, from the generation of terms or documents to final compliance reporting, the process is usually fragmented across multiple stages and systems, lacking end-to-end closed-loop support. Poor data transfer increases uncertainty and compliance risks in the process. Summary of the Invention
[0004] The main objective of this invention is to provide a method, apparatus, device, and storage medium for generating, reviewing, and filing clause documents, aiming to solve the technical problems in the prior art where the generation, review, and filing of clause documents rely on manual operation, lack a unified automation and closed-loop management mechanism, resulting in low efficiency, error susceptibility, and difficulty in version tracing.
[0005] To achieve the above objectives, the present invention provides a method for generating, reviewing, and filing clauses, comprising: Receive demand input information and a set of supporting materials, parse the demand input information to extract key elements, and form an element set; Based on the set of elements, matching information is retrieved from the knowledge base to form a matching result, and a draft clause is generated according to the matching result in a rich text structure; The draft terms are loaded into the editing interface to form a terms document, and the association between the terms document and the set of supporting materials is established during the review process; Based on the aforementioned relationships, a compliance check is performed on the terms and conditions document, and an audit report and issue location information are generated based on the compliance check results. Based on the problem location information, perform revision operations in the editing interface to generate a revised clause document; A second review operation is performed on the revised clause document to generate a review result. When the review result meets the approval conditions, a filing material package is generated. Extract the changes from the revised clause document to generate a clause change description, embed the clause change description into the filing material package, and output the filing material package containing the clause change description.
[0006] Furthermore, to achieve the above objectives, the present invention provides a clause generation, review, and filing device, comprising: The element analysis module is used to receive demand input information and a set of auxiliary materials, analyze the demand input information to extract key elements, and form an element set. The clause generation module is used to retrieve matching information from the knowledge base based on the element set to form a matching result, and generate a draft clause according to the rich text structure based on the matching result; The editing association module is used to load the draft clauses into the editing interface to form a clause document, and to establish the association between the clause document and the set of supporting materials during the review process; The compliance detection module is used to perform compliance detection on the clause document based on the aforementioned relationship, and generate an audit report and issue location information based on the compliance detection results; The revision processing module is used to perform revision operations on the editing interface based on the problem location information and generate a revised clause document; The re-examination and filing module is used to perform a re-examination operation on the revised clause document and generate a re-examination result. When the re-examination result meets the pass conditions, a filing material package is generated. The change description module is used to extract the change content of the revised clause document, generate a clause change description, embed the clause change description into the filing material package, and output a filing material package containing the clause change description.
[0007] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a clause generation and review reporting program stored in the memory and executable on the processor, wherein when the clause generation and review reporting program is executed by the processor, it implements the steps of the clause generation and review reporting method as described above.
[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a clause generation and review reporting program, wherein when the clause generation and review reporting program is executed by a processor, it implements the steps of the clause generation and review reporting method described above.
[0009] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a method, apparatus, device, and medium for generating, reviewing, and filing clauses, comprising: receiving demand input information and a set of supporting materials and extracting key elements to generate an element set; retrieving matching information from a knowledge base based on the element set to generate a draft clause; loading the draft clause onto an editing interface to form a clause document and establishing an association with the set of supporting materials; performing compliance checks based on the association to generate an audit report and problem location information; performing a revision operation on the editing interface based on the problem location information to generate a revised clause document; performing a second audit operation on the revised clause document to generate a re-audit result and generating a filing material package when the pass conditions are met; extracting the changed content of the revised clause document to generate a clause change description and embedding it into the filing material package; and finally outputting a filing material package containing the clause change description. This invention establishes an integrated closed-loop process for clause generation, review, revision, and filing, organically combining semantic parsing, compliance detection, version snapshot management, and filing material generation. This achieves fully automated processing of clause documents from generation to filing, reducing manual intervention and repetitive operations, improving the efficiency and accuracy of clause drafting and review, and ensuring the traceability of clause versions and the standardization of filing materials. Attached Figure Description
[0010] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for the clause generation, review and filing method in one embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the method for generating, reviewing, and filing the terms of this invention. Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the device for generating, reviewing and filing the terms of this invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0011] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0012] The clause generation, review, and filing method provided in this embodiment of the invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can receive requirement input information and a set of supporting materials from the client, extract key elements to generate an element set, retrieve matching information from the knowledge base based on the element set to generate a draft clause, load the draft clause into the editing interface to form a clause document and establish an association with the set of supporting materials, perform compliance checks based on the association to generate an audit report and problem location information, perform revision operations in the editing interface based on the problem location information to generate a revised clause document, perform a second audit operation on the revised clause document to generate a re-audit result, and generate a filing material package when the pass conditions are met, extract the changes in the revised clause document to generate a clause change description and embed it into the filing material package, and finally output a filing material package containing the clause change description. This invention establishes an integrated closed-loop process of clause generation, audit, revision and filing, organically combining semantic parsing, compliance checks, version snapshot management and filing material generation, realizing the fully automated processing of clause documents from generation to filing, reducing manual intervention and repetitive operations, improving the efficiency and accuracy of clause writing and auditing, and ensuring the traceability of clause versions and the standardization of filing materials. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will now be described in detail through specific embodiments.
[0013] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the clause generation, review, and filing method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0014] like Figure 2 As shown, the method for generating, reviewing, and filing clauses proposed in this invention includes the following steps: S10, Receive demand input information and a set of auxiliary materials, parse the demand input information to extract key elements, and form an element set; In this embodiment, the first step is to process the received requirement input information and supporting materials. The requirement input information is typically provided by business personnel, product planning teams, or external organizations, and can take the form of natural language descriptions, policy notices, marketing promotional materials, or clause drafting request forms. Due to the diverse sources of information, the receiving process should support multiple input channels, including typing, file uploads, and API calls. The supporting materials mainly cover policy and regulatory documents related to clause generation, historical clause templates, actuarial analysis documents, market research data, and other materials from different sources. The file formats need to be parsed upon receipt to ensure the content can be processed by the system. The supporting materials set has a high degree of structural complexity, including both plain text and composite documents containing structured content such as tables, charts, and cross-references. Therefore, a unified storage and management method needs to be established after receipt.
[0015] After receiving the input information, a parsing operation is required to extract key elements. The core of parsing lies in converting unstructured natural language expressions and implicit information from multi-source materials into computable and indexable structured expressions. For natural language input, elements can be identified through processing steps such as word segmentation, part-of-speech tagging, dependency parsing, and semantic role labeling. For example, text like "a corn planting period insurance policy for agricultural insurance" needs to be parsed to extract elements such as the type of insurance, the type of insured object, and the applicable period. For document-type auxiliary materials, format parsing and content extraction are required to convert complexly formatted documents into plain text streams. Then, pattern matching and entity recognition algorithms are combined to identify key elements, such as the payout ratio and deductible in actuarial documents, or the regulatory clause number in policy documents.
[0016] During the generation of key elements after parsing, it is necessary to construct an element set. An element set refers to the organization of various elements extracted from different inputs according to a unified data structure, facilitating subsequent knowledge retrieval and clause generation. An element set typically includes element types (such as product category, scope of liability, exclusions, and compensation conditions), element values (such as specific proportions, time ranges, and geographically applicable areas), and element attribute tags. To ensure the completeness and accuracy of the element set, consistency checks and conflict detection need to be performed on the parsing results. If an element differs in materials from different sources, it should be selected based on preset priority rules, or the user should be guided to confirm the final value through a manual confirmation interface.
[0017] The final set of elements is not merely a single-dimensional record, but a hierarchical and interconnected semantic structure. Dependencies need to be established between the elements; for example, compensation conditions must be linked to the scope of liability, and the applicable region must be linked to the geocoding in the policy document. This semantically structured set of elements not only forms the input basis for generating subsequent draft clauses but is also a crucial prerequisite for ensuring consistency between the clause text and business rules.
[0018] In the implementation process, various technical means can be used to optimize element extraction in different scenarios. For text input, deep learning semantic parsing models can be used to improve the accurate recognition of complex expressions; rule-based entity recognition methods can also be used to enhance stability in highly formatted texts such as legal provisions. For file input, a document parsing engine can support the parsing of multiple formats such as PDF, Word, and Excel, and extract structured information such as paragraphs, tables, and formulas during the parsing process to ensure the comprehensiveness of element extraction.
[0019] When fusing multi-source information, inconsistencies between different sources can be addressed by constructing a unified data intermediary layer. For example, in the financial sector, product brochures and actuarial model data may differ in their descriptions of payout ratios; these can be merged using weighted averages or confidence scoring mechanisms. In the healthcare sector, guidelines and historical case documents may differ in their descriptions of applicable conditions; semantic alignment techniques can be used to establish mapping relationships and unify these into a single set of elements.
[0020] When adapting to different business scenarios, the structure of the element set can be flexibly adjusted. For example, in the agricultural insurance scenario, specific elements such as weather conditions and crop growth cycles need to be added to the element set; in the health insurance scenario, medical elements such as treatment items and the scope of application of clinical guidelines can be added. These expansions are all carried out while maintaining a unified data structure to ensure compatibility in subsequent processing.
[0021] This embodiment receives the required input information and a set of supporting materials, and extracts key elements to form an element set. This allows for the establishment of a complete and accurate semantic input foundation before clause generation. This approach reduces the workload of manual searching and organization, improves processing efficiency, and ensures the diversity of information sources and the consistency of results. It provides traceable and calculable input conditions for the automation of clause generation, review, and filing.
[0022] S20, based on the element set, retrieve matching information from the knowledge base to form a matching result, and generate a draft clause according to the matching result in a rich text structure; In this embodiment, the element set is a structured dataset formed through prior processing, containing content across different dimensions such as product category, scope of liability, compensation conditions, exclusions, and applicable regions. To correlate these elements with existing data, matching information needs to be retrieved from a knowledge base. The knowledge base typically includes policy and regulatory databases, historical clause template databases, actuarial parameter databases, etc., and its internal storage format can be a relational database, a vectorized index database, or a graph-based knowledge graph. During retrieval, a query statement or semantic vector needs to be constructed based on the feature values of the element set. Through keyword matching, semantic similarity calculation, and regularized pattern recognition, the clause fragments, parameter records, or case templates corresponding to the element set are located from the knowledge base.
[0023] The formation of matching information is not a simple single-source retrieval, but a comprehensive result obtained by merging and comparing matching results from multiple databases. For example, policy and regulation matching information may include legal provision numbers and their text content, historical clause template matching information may include clause paragraphs that highly overlap with the element set, and actuarial parameter matching information may provide numerical parameters. The system needs to structurally integrate this matching information, remove redundancy and conflicts, and form a unified matching result set internally.
[0024] After obtaining the matching results, a draft clause needs to be generated according to a rich text structure. A rich text structure refers to a document structure framework that simultaneously supports paragraph hierarchy, heading numbering, tables, lists, cross-references, and highlighting. When generating the draft clause, the document's chapter structure must first be determined, generally including sections such as general provisions, insurance liability, exclusions, compensation handling, and supplementary provisions. Then, the matching results are populated into different chapters according to the categories corresponding to the element sets; for example, compensation conditions are entered into the insurance liability section, and exclusions are entered into the exclusions section. During the population process, cross-reference markers need to be automatically inserted for quick navigation and consistency checks in subsequent documents. Finally, the populated content undergoes a logical consistency check to ensure there are no conflicts between contexts, and the output is a formatted rich text draft clause.
[0025] At the implementation level, different knowledge retrieval technologies can be adopted. Keyword indexing can be used for fast retrieval, suitable for scenarios involving rule-based provisions in legal databases; semantic retrieval models can be used to encode the element set into semantic vectors and perform similarity calculations with vectorized clause fragments, thereby improving the matching ability for complex semantic expressions; graph reasoning methods can also be used to infer clause nodes that match the element set using nodes and relationships in the knowledge base.
[0026] When integrating matching results, a confidence threshold can be designed to remove matches below the threshold, ensuring the accuracy of the draft clauses. When generating the draft clauses, a template-based rendering method can be chosen to embed the matching information into a predefined clause framework, or a text generation engine can be used to perform natural language concatenation on the matching information, enhancing the readability of the draft. For rich text structure construction, it can be generated using an editor that supports HTML, Markdown, or RTF formats, or rendered using a dedicated clause editor, ensuring the complete functionality of cross-references, tables, and lists.
[0027] Adaptability can be tailored to different scenarios. In the generation of product terms in the financial sector, emphasis can be placed on matching actuarial parameters with regulatory compliance provisions to ensure that the draft covers fund flows, risk control, and regulatory requirements. In the generation of terms in the healthcare sector, more emphasis can be placed on matching policy documents and clinical guidelines to ensure that the draft covers the scope of treatment, exclusions, and reimbursement conditions.
[0028] This embodiment retrieves matching information from a knowledge base based on a set of elements and generates matching results. Then, it generates draft clauses according to a rich text structure based on the matching results, thus automating the transition from requirement input to initial clauses. This reduces the workload of manual word-by-word drafting, improves the efficiency and accuracy of clause writing, and ensures that the clause content is consistent with policies, historical cases, and actuarial parameters in the knowledge base, providing a higher-quality foundation for subsequent review and revision.
[0029] S30, Load the draft clauses into the editing interface to form a clause document, and establish the association between the clause document and the set of supporting materials during the review process; In this embodiment, the process of loading the draft clauses into the editing interface involves the transformation of document data from the generation module to the interaction module. The draft clauses are a rich text document generated through matching and filling, containing formatting information such as chapter structure, cross-references, and tables. To achieve loading, the editor interface needs to be called to parse the rich text content into renderable objects. The editing interface can be a web-based rich text editor or a desktop-specific editing tool, whose basic functions include text rendering, formatting control, paragraph style adjustment, and interactive modification support. During the loading process, the system needs to import the original content and formatting information of the draft clauses to ensure that the structure and style of the draft can be fully displayed in the editing interface.
[0030] Once the clause document is created, the review process requires establishing a link between the clause document and the supporting materials. The supporting materials include policy documents, historical clauses, actuarial reports, and other multi-source documents, which may be in PDF, Word, Excel, or structured database formats. To establish this link, the supporting materials first need to be parsed to extract text, tables, or parameter information. The extracted information then undergoes standardized processing, such as cleaning redundant symbols, segmenting, and deduplication. Subsequently, it is transformed into a high-dimensional vector using an embedding model or keyword indexing and stored in a vector database.
[0031] The connection between the clause document and the supporting materials is established based on a unified semantic index space. Paragraphs or key sentences in the clause document also undergo the same vectorization process, enabling similarity calculations with the supporting materials within the same space. This semantic matching mechanism allows for the identification of corresponding policy provisions or historical cases for each part of the clause document. To support subsequent tracking and auditing, each established relationship is assigned a unique session identifier and bound to a specific clause document paragraph and its corresponding supporting material record. These semantic relationships can be invoked in real-time during the review process, supplementing the clause document with contextual evidence.
[0032] When implementing the loading of draft terms, an HTML rendering engine can be used to perform rich text conversion, or a Markdown parser or RTF parser can be used to ensure compatibility with document formats from different sources. In practical applications, the DOM node structure or document flow of the draft file can be directly written into the editing interface by calling the rich text editor API. To ensure rendering consistency, cross-page tables, numbered lists, and multi-level headings need to be processed again to avoid misalignment in the interface.
[0033] Different semantic matching methods can be used to establish the association between the clause document and the set of supporting materials. One method is to use a pre-trained semantic embedding model to calculate the cosine similarity between paragraphs in the clause document and paragraph vectors in the supporting materials to obtain the most relevant candidate matches. Another method is to use a knowledge graph-based retrieval mechanism to map entities and relationships in the clause document to knowledge nodes in the supporting materials, forming a structured association. Furthermore, a rule engine can be combined to overlay regularization rules and dictionary mappings on top of semantic matching to improve accuracy.
[0034] The implementation of the system will differ depending on the scenario. In financial business scenarios, the focus can be more on establishing connections between the terms and conditions documents and regulatory rule bases and actuarial data to quickly identify compliance risks. In healthcare business scenarios, the emphasis can be more on the correspondence between the terms and conditions and clinical policy guidelines and payment standards to ensure that the document content covers the scope of treatment and reimbursement conditions.
[0035] This embodiment loads the draft clauses into the editing interface to form a clause document, and then establishes a link between the clause document and the set of supporting materials during the review process, enabling semantic linkage between the clause document and external materials. This not only improves the operability of the clause document during the editing process, but also provides contextual evidence support for the clauses during the review process, thereby reducing the workload of manual comparison and improving the accuracy and efficiency of the review.
[0036] S40, Perform compliance checks on the terms document based on the aforementioned relationship, and generate an audit report and issue location information based on the compliance check results; In this embodiment, to perform compliance checks on clause documents based on relationships, it is necessary to first clarify the role of these relationships. Relationships are semantic mappings between clause documents and sets of supporting materials, providing corresponding paths between clause paragraphs and external evidence. Through this mapping mechanism, the system can call upon supporting materials in real time when checking clauses, ensuring that the results not only rely on the text of the clauses themselves but also compare them with policy documents, historical clauses, and actuarial parameters, thereby forming a compliance judgment supported by evidence.
[0037] Compliance checks are typically divided into three categories. Format constraint checks examine whether the structure and layout of the clause document conform to regulations, specifically including elements such as chapter completeness, numbering order, heading levels, table boundaries, and list nesting. In implementation, the system parses the rich text structure of the clause document, extracts heading nodes and paragraph levels, and compares them with the regulatory-specified structure template to identify missing or incorrectly ordered parts.
[0038] Semantic constraint detection is used to determine whether the semantic expression of the clause text contains content that violates policy rules. This can be achieved by calculating the similarity between the sentence vectors of the clause document and the vectors of prohibited or restrictive expressions in the regulatory rule base; if the similarity exceeds a threshold, it is marked as a potential violation. Alternatively, a natural language processing model can be introduced to perform entity recognition and relation extraction on the clause text to determine whether the expression contains elements that violate regulatory provisions.
[0039] Logical consistency checks are used to analyze whether there are conflicts in the logical relationships between different parts of a clause, such as whether the statements regarding the scope of liability and exclusions contradict each other. This can be achieved by constructing a rule graph, using liability-related statements and exclusion clauses as nodes, and using an inference engine to check for overlapping or conflicting relationships. Alternatively, a Boolean logic model can be used to model the clause conditions and detect inconsistencies.
[0040] After the inspection is completed, the results need to be integrated to generate an audit report and issue location information. The audit report is a structured output, containing the overall conclusions of the compliance inspection results and the status of each inspection item. The issue location information is based on the mapping relationship between the clause documents and supporting materials, binding the discovered issues to specific paragraphs of the clauses, supporting subsequent revisions and source tracing. During the generation of the audit report, each issue needs to be labeled with the inspection type, violation content, and related evidence, and machine-generated modification suggestions need to be provided.
[0041] In practical implementation, format constraint detection can be accomplished using a regular expression parser and a tree structure comparator. The clause document is first parsed into a tree structure, and then the system compares it layer by layer with a standard structure template. Semantic constraint detection can employ a Transformer-based language model, where the clause text is segmented and input into the model, and the output vector is matched against existing expressions in the rule base. Logical consistency detection can be achieved using knowledge graph methods, abstracting liability clauses and exclusion clauses into entities and relationships, and using graph matching algorithms to detect contradictory relationships.
[0042] The implementation methods differ depending on the scenario. In the fintech sector, a key focus can be on strengthening the matching between regulatory documents and actuarial parameters. For example, this can be achieved by establishing an actuarial parameter database to numerically verify the reimbursement ratios or rates expressed in the terms and conditions. In the healthcare sector, emphasis can be placed on semantic comparison between clinical guidelines and payment catalogs. For instance, this involves mapping the scope of treatment clauses against the medical insurance payment item database to identify any statements of liability that exceed the scope of coverage.
[0043] For the implementation of audit reports, JSON or XML structured storage can be used for subsequent visualization or interaction with external systems. Issue location information can be implemented by adding unique identifiers to paragraphs in the terms document and binding the issue item to this identifier in a database. This allows users to jump to the corresponding paragraph by clicking on the issue report in the editing interface.
[0044] This embodiment performs compliance checks on clause documents based on relationships and generates audit reports and issue location information based on the results. This transforms clause inspection from a process relying solely on manual comparison into an automated, traceable, and intelligent process. This reduces the workload of manual review and lowers the risk of omissions and inconsistencies. Simultaneously, the audit report provides structured results, and the issue location information ensures the traceability of problems, thereby improving overall audit efficiency and accuracy.
[0045] S50, based on the problem location information, perform a revision operation in the editing interface to generate a revised clause document; In this embodiment, revision operations are performed on the editing interface based on the issue location information. The core of this approach is to link the issue location generated during compliance testing with the editing process of the clause document. Issue location information is typically stored in a structured, marked format, including paragraph identifiers, violation types, violation content, and revision suggestions. In practice, the editing interface parses the issue location information, maps the issue items to specific locations in the clause document, and marks them with highlights, underlines, or annotations, allowing users to view and process them intuitively within the interface.
[0046] Revision operations can be performed in two ways: manual revision and intelligent revision. Manual revision involves reviewers manually modifying the clause text in the editing interface based on location information, such as deleting non-compliant expressions, adding missing content, or adjusting the logical structure. Intelligent revision, on the other hand, involves the system automatically generating new text that meets compliance requirements based on a revision suggestion library or language generation model, and providing replacement or merging options. Regardless of the method used, the revision results must be updated in the interface in real time to ensure consistency between the editing status and the clause document.
[0047] A version management mechanism needs to be established during the generation of revised clause documents. The system will store the original clause documents and the revised clause documents separately, and record the revision time, reviser, and summary of the changes for later review and comparison. The revised clause document is a rich text object generated based on the text finally confirmed in the editing interface, retaining the original format and structure while replacing or adjusting the content related to the issues, forming a compliant version.
[0048] In practical applications, the editing interface can adopt a dual-view mode, displaying the terms document on one side and a list of issue location information on the other. When the user clicks on an issue item in the list, the interface automatically jumps to the corresponding paragraph and provides an editing entry. Another approach is to embed issue markers directly into the terms document, allowing users to view the reasons for violations and revision suggestions when the cursor hovers over them, and to perform modifications in a floating window.
[0049] The revision methods differ across different business scenarios. In financial business scenarios, revision operations can be linked to a regulatory parameter database. For example, when a compensation ratio is detected to exceed the prescribed range, the system directly provides a reference value for the legal range on the revision interface for the user to quickly select. In healthcare business scenarios, revision operations can be integrated with a medical knowledge base. For example, when a liability description is found to contain items not covered by the payment catalog, the system prompts the user to delete or replace them, and provides legal items within the catalog as candidate alternatives.
[0050] Automatic revisions can be achieved through a configurable rules engine. For example, when an issue is marked as "missing format," the system can automatically add missing section titles or numbers; when an issue is marked as "prohibited expression," the system can replace it with a compliant expression template according to preset rules. These automatic revision results undergo manual review before generating the revised clause document to ensure that the modifications do not introduce new issues.
[0051] This embodiment significantly improves the efficiency and accuracy of clause revision by performing revision operations based on issue location information within the editing interface and generating revised clause documents. The precise mapping of issue location information avoids the inefficient process of manually searching for problems word by word, while the version management mechanism of the revised clause documents ensures the traceability and verifiability of the revisions. By combining manual and intelligent revision, the review process can be accelerated while maintaining compliance, thereby shortening the clause finalization cycle.
[0052] S60, perform a second review operation on the revised clause document to generate a review result, and generate a filing material package when the review result meets the pass conditions; In this embodiment, performing a second review of the revised clause document essentially involves re-inputting the results of the previous revision into the review engine for secondary verification using the same or extended set of compliance detection rules. The revised clause document, as the input, needs to be parsed into structured text units, including chapter numbers, item numbers, and text paragraphs. The second review process is not merely a repetition of the initial review; rather, it focuses on the revised content areas based on the initial review results, while simultaneously performing a rapid consistency comparison on the unchanged parts to reduce redundant checks.
[0053] When generating the review results, the system distinguishes the status of the output, such as fully passed, partially problematic, or with new anomalies. The review results are generally stored as a data report, including compliance status, a list of issues, difference comparison results, and a stability evaluation of the revised version. The difference comparison can be achieved through text hashing, syntactic analysis, or semantic embedding comparison to confirm whether the scope of differences between the revised and initial clause documents meets expectations.
[0054] When the review result meets the approval criteria, the process of generating the filing materials package is triggered. The filing materials package includes formatting the revised clause documents, attaching necessary compliance testing records, review signature information, and supplementary tables containing regulatory requirements. This materials package is typically packaged in compressed or containerized document format to ensure that different types of files and metadata can be packaged and output in a uniform format for easy uploading or archiving.
[0055] In one implementation, the re-audit operation is performed by invoking a pre-defined regulatory rules engine. The rules engine verifies each item in the revised terms document to confirm whether the revision has completely eliminated the problems identified in the initial audit. To improve efficiency, a deep audit can be triggered only for the areas of difference, while a quick consistency comparison is performed on the unchanged parts.
[0056] In another implementation, the re-examination process introduces a multi-level review strategy, which adds a manual confirmation step on top of automated review. After the system generates the re-examination result, it automatically determines whether it passes or fails, but reviewers are still allowed to manually confirm key clauses, such as those involving the scope of compensation, exemptions from liability, or medical catalog limitations, to avoid missing risk points due to incomplete coverage by automated rules.
[0057] The generation methods for the reporting materials package also differ. In the financial sector, the package can automatically incorporate standard report formats required by regulatory authorities and attach a clause revision history table to the clause documents for audit tracking. In the healthcare sector, the package can include medical references, applicable catalogs, and reports on differences before and after revisions to meet the traceability requirements of healthcare regulatory authorities. The integrity and immutability of the reporting materials package can be enhanced through encrypted storage and digital signatures, and it can also be configured for multi-format output, such as both PDF and XML, to adapt to different system interface requirements.
[0058] This embodiment ensures the validity and completeness of revisions by re-reviewing the revised clause documents and generating a re-review result, preventing duplicate issues or newly added violations from being overlooked. When the re-review result meets the approval criteria, a filing material package is generated, ensuring that the clauses have a compliant and complete set of documents before formal submission. This achieves closed-loop management of clauses from revision to filing, improving compliance reliability and reducing the time cost of repeated manual verification.
[0059] S70, extract the changes from the revised clause document to generate a clause change description, embed the clause change description into the filing material package, and output the filing material package containing the clause change description.
[0060] In this embodiment, extracting the changes to the revised terms document typically requires comparing the document structure and text content of different versions. The revised terms document, as input, is first broken down into units such as clause paragraphs, heading numbers, and additional explanations, and then compared with the previous version of the terms document. The comparison methods can include line-by-line text comparison, structured comparison based on syntax trees, or vector matching based on semantic embedding. These methods can identify newly added, deleted, or modified portions and represent the changed areas in the form of highlights, tags, or indexes.
[0061] The process of generating a clause change notice involves identifying the differences and extracting the changed areas into a written explanation. The clause change notice not only lists the locations of the changes but also summarizes the type of change (such as adjusting reimbursement ratios, adding exclusion clauses, or updating the medical insurance catalog), and cites relevant policies or regulatory requirements. The explanation document can use a structured rich text format, including a title, changed items, original content, revised content, and corresponding regulatory basis, ensuring intuitive readability and audit traceability for subsequent use.
[0062] The process of embedding the clause change description into the filing material package requires reserving space in a standardized filing template to automatically populate the description document into the designated area. The embedding method typically relies on a unified document generation engine capable of integrating the text description with other filing materials (such as compliance reports, full text of clauses, and supplementary material reference tables) into a complete package. The final output filing material package will use a unified format, such as a multi-file compressed package or a multi-level document container, ensuring logical and version consistency between the clause change description and the clause text.
[0063] In one implementation, the extraction of changes is accomplished through hash-based difference detection, that is, calculating a hash value for each paragraph of the terms document, comparing it with the previous version to quickly pinpoint the change points, and then generating a terms change description.
[0064] In another implementation, the change description document is generated by an intelligent parsing module. This module not only extracts textual differences but also automatically identifies paragraphs that involve regulatory priorities, such as rate adjustments in financial terms or catalog updates in medical terms, and adds compliance references to the description document.
[0065] The embedding method for the clause change instructions can also be flexibly chosen. In financial business scenarios, the change instructions can be directly attached to the end of the filing materials package, and a table of contents can be generated for regulators to quickly locate them. In healthcare business scenarios, the change instructions can be displayed side by side with the revised clause documents, along with a comparison table to visually show the differences between the clauses before and after the modification. The filing materials package can ultimately be exported as a digitally signed PDF document or a structured XML file to meet the system interface requirements of different regulatory agencies.
[0066] This embodiment ensures the transparency and traceability of revisions by extracting the changes from the revised clause documents and generating clause change descriptions, reducing the risk of omissions or misjudgments during the review process. Embedding the change descriptions into the filing materials package allows regulatory reviewers to directly grasp the scope of revisions and compliance basis in a single review, thereby shortening the filing and review cycle and improving the overall efficiency of clause management and filing.
[0067] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a method, apparatus, device, and medium for generating, reviewing, and filing clauses, comprising: receiving demand input information and a set of supporting materials and extracting key elements to generate an element set; retrieving matching information from a knowledge base based on the element set to generate a draft clause; loading the draft clause onto an editing interface to form a clause document and establishing an association with the set of supporting materials; performing compliance checks based on the association to generate an audit report and problem location information; performing a revision operation on the editing interface based on the problem location information to generate a revised clause document; performing a second audit operation on the revised clause document to generate a re-audit result and generating a filing material package when the approval conditions are met; extracting the changed content of the revised clause document to generate a clause change description and embedding it into the filing material package; and finally outputting a filing material package containing the clause change description. This invention establishes an integrated closed-loop process for clause generation, review, revision, and filing, organically combining semantic parsing, compliance detection, version snapshot management, and filing material generation. This achieves fully automated processing of clause documents from generation to filing, reducing manual intervention and repetitive operations, improving the efficiency and accuracy of clause drafting and review, and ensuring the traceability of clause versions and the standardization of filing materials.
[0068] In one embodiment, step S10 includes: S101, receive natural language description input containing product feature descriptions, and receive file upload input containing policy documents and historical terms; S102, perform semantic intent recognition and entity extraction on the natural language description input to generate structured semantic parsing results; S103, perform format parsing and content extraction on the uploaded file to generate multi-source file parsing results; S104, The structured semantic parsing result and the multi-source file parsing result are fused together to generate a preliminary set of key elements; S105, Perform integrity verification on the preliminary set of key elements, and perform compliance verification on the verified preliminary set of key elements based on the regulatory strategy library; S106 will classify the key elements that pass the compliance verification in multiple dimensions according to product type, agreement subject matter and regional characteristics, and generate an element set containing classification tags and attributes.
[0069] In this embodiment, the input stage targets two types of original information sources: one is requirement input information, which comes from product planning documents, text descriptions submitted by business teams, task orders issued by external organizations, or interface pushes; the other is a collection of auxiliary materials, which comes from policy and regulation repositories, historical clause archives, actuarial analysis archives, market research databases, etc. The input channel simultaneously supports three paths: text input, batch file upload, and system integration. After entering the unified collection queue, file fingerprint registration, character encoding normalization, format metadata extraction, and security verification are completed to ensure the availability and traceability of data in subsequent parsing stages. Requirement input information and the collection of auxiliary materials establish source records and timestamps during the receiving stage. Any subsequent derived information retains a source traceability pointer, ensuring that each item can be linked back to the location section or page number of the original material.
[0070] Natural language descriptions are input into a semantic processing pipeline, aiming to generate structured semantic parsing results. The text is first segmented and denoised, then aligned with a terminology dictionary, named entity recognition, and intent determination are performed. Key slots for clause generation are extracted, including product type, coverage scope, exclusions, payout trigger conditions, applicable region, index or parameter name, time interval, subject, and object. To control ambiguity, entities are stored in a standardized dictionary and bound to a thesaurus; cross-sentence references are resolved through reference parsing; numerical values and units of measurement are normalized using a unit library; and date expressions are converted into standard calendar objects using time regularization. The structured semantic parsing results output by this pipeline are represented as key-value pairs, hierarchical paths, and source evidence triples, which can be consumed by the search engine and facilitate subsequent consistency verification.
[0071] File uploads cover PDFs, Word documents, Excel spreadsheets, scanned images, and semi-structured tables. Format parsing prioritizes layout reconstruction and hierarchical syntax tree extraction. Tables are reconstructed using cell coordinates and cell merging relationships; cross-page tables retain cross-page links; and footnotes and attachments are linked using anchor points. Images and scanned documents undergo text recognition and layout reconstruction; equal-width tables and multi-column layouts are incorporated into the regular text flow after reconstruction. Content extraction operates in parallel across paragraphs, tables, and lists, extracting article numbers, chapter titles, clause fragments, parameter tables, caliber descriptions, and citation links to generate multi-source file parsing results. Each parsed entry retains its page position, row and column coordinates, and page block identifier to ensure accurate positioning for subsequent comparisons and evidence presentation.
[0072] The structured semantic parsing results and multi-source document parsing results are aligned into a unified set of semantic objects during the fusion process, producing a preliminary set of key elements. The alignment order follows two principles: semantic matching priority and explicit citation priority. Entries with article numbers or explicit title anchors are directly bound to semantic objects; entries lacking anchors are bound through a collaborative retrieval process using vector similarity and keyword matching. Conflict resolution employs a joint adjudication of source priority and confidence level. Priority is derived from material type and authority, while confidence level is derived from matching score and evidence density. If the adjudication is still not unique, parallel candidates are retained and marked as pending confirmation as intermediate states. The fusion process completes unit conversion, regional name standardization, and industry terminology mapping. All elements are accompanied by source pointers, matching scores, and cleaned records to ensure interpretability and revocability.
[0073] The initial set of key elements is then processed for completeness and compliance verification. Completeness verification performs missing and mutual exclusion checks based on scenario templates and regulatory checklists: for example, coverage must be accompanied by liability definition and triggering conditions; mutually exclusive exclusions cannot appear simultaneously with their corresponding liability items. Rules are stored as declarative constraints, and the verifier outputs missing fields, conflict pairs, and suggested completion paths. Compliance verification interfaces with a regulatory strategy library, covering policy provisions, guidelines, filing format requirements, and a list of sensitive expressions. The verification process first performs semantic matching to identify suspected non-compliant expressions, then compares templates to confirm structural constraints, and, if necessary, invokes parameter-based checks to compare rate caps, compensation ranges, or geographical restrictions. Verification results are bound to elements, indicating the triggered clause nodes, evidence fragments, and rectification suggestion categories, providing evidence-based constraint boundaries for downstream draft generation and subsequent review.
[0074] Key elements validated are entered into a multi-dimensional classification system, outputting a set of elements containing classification tags and attributes. The classification dimensions cover three main axes: product type, agreement target, and regional characteristics. Product type is mapped to a unified product tree and its hierarchical path is recorded. Agreement targets are represented by standard target library entries with accompanying attribute tables (such as risk category, insurable scope, and pricing caliber). Regional characteristics are bound to administrative division codes, climate zones, or geographical zoning tables. During classification, index keys and inverted items are generated simultaneously for easy subsequent retrieval and chapter generation for specific dimensions. The attribute area records caliber, unit, version number, validity period, and source tags to ensure cross-version comparisons and expiration reminders are triggered. Results are stored in a dual-track system using a structured object graph and renderable fragments. One track serves knowledge retrieval and consistency checks, while the other serves rich text generation and cross-reference insertion.
[0075] This embodiment unifies the access requirements and supporting material sets, completes semantic parsing, format parsing, and evidence alignment, and merges them into an element set after integrity verification and comparison with the regulatory strategy library. It also establishes multi-dimensional classification and attribute tags according to product type, agreement subject, and regional characteristics, which can significantly reduce the cost of manual sorting and standard alignment. It forms a calculable, traceable, and constrained input base in advance at the generation front end, reducing repeated rework and omissions in subsequent generation, review, and revision. At the same time, it provides structured input with evidence for compliance modules and version management, so that subsequent document generation and compliance detection operate on consistent data semantics and maintain cross-version comparability and auditability.
[0076] In one embodiment, step S20 above includes: S201, based on the element set, retrieve matching content from the policy and regulation database of the knowledge base to form a policy and regulation matching subset; S202, based on the element set, retrieve matching content from the historical clause template database of the knowledge base to form a historical clause matching subset; S203, based on the element set, retrieve matching content from the actuarial parameter database of the knowledge base to form an actuarial parameter matching subset; S204, perform multi-source data fusion on the policy and regulation matching subset, the historical clause matching subset, and the actuarial parameter matching subset to generate a comprehensive matching result; S205, Determine the clause structure framework based on the comprehensive matching results; S206, Fill in the specific content of each chapter of the initial draft clauses according to the clause structure framework, and insert cross-reference marks during the content filling process; S207, Perform logical consistency verification on the completed draft clauses, and format the draft clauses that pass the logical consistency verification according to the rich text structure to generate the final draft clauses.
[0077] In this embodiment, the element set provides semantic anchors and filtering conditions for retrieval, including items such as product type, scope of liability, exclusions, compensation triggering conditions, subject matter of the agreement, geographical application, scope of application, and time limit. Each item includes a source pointer and normalized expression, facilitating the generation of an executable retrieval intent. The retrieval intent consists of two parts: one part is for precise field matching, using field names and normalized values to generate a structure for filtering; the other part is for semantic similarity, using vector expressions to cover synonymous descriptions and cross-text rewriting. The two parts are jointly executed in the same query to balance interpretability and recall coverage.
[0078] The policy and regulation database contains regulatory documents, industry guidelines, and filing format requirements. Articles are organized using fields such as hierarchical numbering, title, body text, and footnotes. A list of prohibited and restricted expressions and citation relationships between articles are added to facilitate downstream constraint verification. The historical clause template database collects filed or archived clause fragments and complete texts, labeled with chapter type, applicable scenarios, and applicable regions. Inverted indexes and paragraph vectors are built at the fragment granularity for rapid assembly and reuse. The actuarial parameter database stores pricing standards and parameter descriptions related to liability definitions. It is organized using fields such as parameter name, applicable objects, time range, regional scope, and standard description, and supports cross-expression retrieval using parameter semantic vectors. All three databases provide dual-channel interfaces for keyword indexing and semantic indexing, supporting conditional filtering, field weighting, and source priority control.
[0079] The query pipeline is generated based on a set of elements. The first pipeline targets a policy and regulation database, prioritizing precise location using article numbers, chapter titles, and prohibited / restricted word lists. Semantic vectors are then used to supplement similar articles, forming a policy and regulation matching subset. Each record in this subset includes the article path, evidence fragment, and appropriate score. The second pipeline targets a historical clause template database, filtering by chapter type and applicable scenario. Semantic vectors for scope of liability, exclusions, and triggering conditions are then used to retrieve similar fragments, forming a historical clause matching subset while preserving cross-references between fragments and the original templates. The third pipeline targets an actuarial parameter database, limiting the search domain by the agreement's subject matter and regional scope. Semantic matching is then performed using parameter names and definitions, forming an actuarial parameter matching subset, and binding textual descriptions of parameter definitions and applicable boundaries.
[0080] Multi-source data fusion focuses on alignment and adjudication. The alignment phase maps three subsets to a unified fragment model, standardizing field names and semantic labels, and establishing a mapping table for each entry in the feature set. This ensures that each entry is associated with one or more of the three types of evidence: policy basis, historical expression, and parameter caliber. The conflict adjudication phase jointly sorts the data based on source authority, time relevance, and matching score. It first selects the primary evidence, then merges similar evidence on the same topic into parallel evidence, avoiding content fragmentation. The redundancy removal phase merges semantically repetitive fragments and fragments with similar formats, retaining representative expressions and their source links. The fusion output is a comprehensive matching result, including the fragment text, source list, matching score, applicable scope label, and binding relationship with the feature set.
[0081] The clause structure framework is jointly determined by the subject mapping and the template library. The subject mapping selects the chapter skeleton and order based on the items in the element set, such as general provisions, liability, exclusions, accounting, claims, and supplementary provisions, and creates placeholders and constraints under each chapter. The template library provides chapter openings, common definitions, and table frameworks as content containers. After the structure framework is generated, content filling begins, writing the comprehensive matching results sequentially into the corresponding chapter placeholders according to the mapping table. Three types of adjustments are performed during the filling process: first, consistency in wording, merging similar expressions from multiple sources into a single paragraph while maintaining the evidence list; second, consistency in terminology, replacing synonyms according to a semantic dictionary to maintain consistency within and between clauses; and third, linkage of definitions, automatically attaching parameter definition fragments when parameter names appear in the liability definition to maintain consistent descriptions between definitions and pricing.
[0082] Cross-reference markers are generated synchronously during the population process. Each definition, table, parameter name, and clause node is assigned an anchor point and a reference key. The reference key is written at the reference point to establish a link between the referenced elements and the referenced elements, supporting navigation and consistency checks during editing and subsequent review. After cross-references are generated, a sanity check is performed to ensure that the referenced target exists and has not been pruned, and that the numbering, headings, and structural framework are consistent.
[0083] Logical consistency verification constructs a constraint graph based on the semantic relationships within and between chapters. The constraint graph uses nodes to represent responsibilities, exclusions, definitions, parameters, and triggering conditions, and edges to represent inclusion, exclusion, prerequisites, or dependencies. The verification engine traverses the constraint graph, checking for obvious conflicts, overlapping expressions, and unclosed references. It also performs value domain alignment for cross-chapter dependencies. When inconsistencies are found, it returns the conflict point, related fragments, and sources of evidence for correction. Text that passes verification enters the rich text formatting stage, applying an editor-compatible structured mode to standardize the rendering of heading levels, numbering systems, tables, lists, and references, ensuring clear structure and interactive usability in subsequent interfaces.
[0084] This embodiment uses a set of elements as the sole entry point for the search intent. It retrieves matching information from three sources: policies and regulations, historical clause templates, and actuarial parameters, and integrates them into a comprehensive matching result. Then, based on the structural framework, it generates a draft clause through chapter filling and cross-referencing, ultimately completing a draft clause that has undergone consistency verification and rich text formatting. This allows for evidence binding, terminology standardization, and structural standardization during the generation stage. This reduces omissions and inconsistencies caused by manual collection and splicing, shortens the drafting cycle, improves the consistency between clauses and regulatory interpretations, and between historical practices and parameter interpretations, and provides a traceable, verifiable, and maintainable content foundation for subsequent review and revision stages.
[0085] In one embodiment, step S30 above includes: S301, Load the draft terms into the editing interface of the rich text editor, and render the loaded draft terms in rich text format to form a structured terms document. S302, in the editing interface, enable the automatic numbering function for multi-level headings, the nesting function for lists, and the function for breaking rows across pages in tables; S303, During the review process, the format of the multi-format documents in the set of auxiliary materials is parsed to extract the original text content; S304, the original text content is cleaned, segmented and deduplicated, and the processed text content is transformed into a high-dimensional vector through an embedding model; S305, store the high-dimensional vector into a vector database to construct a unified semantic index space; S306, Based on the unified semantic index space, establish the semantic association between the structured clause document and the set of auxiliary materials; S307, In the process of establishing semantic association, a unique session identifier is assigned to the association; S308, Perform semantic association mapping based on the unified semantic index space to retrieve the most relevant policy provisions and historical cases from the vector database in real time as contextual evidence; S309, Generate an associated session context containing semantic association mapping information based on the context evidence; S310, the associated session context is bound and stored with the unique session identifier.
[0086] In this embodiment, before the draft terms enter the rich text editor's editing interface, the document object model is reconstructed and format metadata is extracted. During reconstruction, chapter hierarchy, paragraph styles, numbering system, table structure, and cross-reference placeholders are preserved. The rendering engine maps according to the rich text specifications supported by the editor, converting titles, body text, lists, tables, annotations, and anchors into a set of renderable nodes. During rendering, row break control and cell merging relationships are corrected for cross-page tables to ensure consistency in subsequent pagination display and export behavior. After rendering is complete, a structured terms document is formed. The structured meaning is that the document not only has a visual style but also has computable attributes such as node type, hierarchy, and reference key values, supporting subsequent location, retrieval, and constraint validation.
[0087] The editing interface enables three capabilities: automatic numbering of multi-level headings, nested lists, and cross-page row breaks in tables. These capabilities respectively ensure structural consistency, hierarchical expression, and stability in complex layouts. Automatic numbering of multi-level headings recalculates the numbering in real time and updates the cross-reference text synchronously when the user adjusts the hierarchy or inserts / deletes paragraphs. Nested lists provide validation of the legal nesting relationships of clauses, sections, and items, preventing hierarchical jumps and incomplete closures. Cross-page row breaks in tables maintain the strategy of repeating table headers and aligning row merge boundaries and paragraphs, avoiding misalignment during export or printing. All three capabilities are linked to the node attributes of the structured clause document; any editing operation will trigger consistency checks and necessary reflows.
[0088] The review process involves accessing a collection of supporting materials in various formats, prioritizing format parsing and text extraction. The parsing channel covers formatted documents, editable documents, and scanned images. Formatted documents utilize layout reconstruction and reading order prediction to output paragraph flows. Editable documents directly extract paragraphs, tables, and bookmarks based on an embedded style tree. Scanned images undergo text recognition and layout segmentation before entering the paragraph flow. The text extraction stage preserves page coordinates, row and column coordinates, bookmark paths, and source file summaries to form the original text content and add source pointers, facilitating evidence presentation and retrospection.
[0089] The original text content enters a cleaning, segmentation, and deduplication pipeline. Cleaning processes include standardized encoding, punctuation, and elimination of soft line breaks and redundant whitespace. Segmentation strategies combine punctuation, title features, and page breakpoints to generate stable semantic segments. Deduplication employs fingerprint hashing and near-duplicate detection, removing low-difference segments and merging highly similar segments to reduce subsequent indexing redundancy. After preprocessing, an embedding model is used to convert the segments into high-dimensional vectors. During vectorization, window sliding and overlapping strategies are introduced to ensure coverage of long segments. Special structures such as tables are generated as composite vectors using a joint encoding method of "structural summary + cell content," balancing structure and content.
[0090] High-dimensional vectors are stored in a vector database to construct a unified semantic index space. This index space incorporates an approximate nearest neighbor retrieval structure and metadata filtering capabilities, establishing filtering keys such as material type, source level, validity period, and regional tags, supporting conditional retrieval within a large-scale semantic vector set. To ensure consistency, paragraphs, definitions, tables, and anchor points in structured clause documents are also vectorized and written into the index space, ensuring that clause content and supporting materials fall within the same semantic coordinate system.
[0091] A semantic association between structured clause documents and supporting material sets is established within a unified semantic index space. The associated objects are bidirectional mappings between clause paragraphs and material fragments. The process initiates a semantic similarity search starting with the clause paragraph as the query endpoint, overlaid with metadata filtering (e.g., priority given to regulatory materials, materials from the same region, and the latest version), returning a set of candidate evidence. The candidate set undergoes relevance threshold pruning and topic redundancy removal, retaining representative evidence and recording similarity, source, page number, and evidence fragment range. Each mapping includes clause paragraph anchor points and material fragment location information upon saving, enabling bidirectional navigation from clauses to evidence and from evidence back to clauses.
[0092] Each semantic relationship is assigned a unique session identifier to manage all mappings, retrievals, and user interactions generated within the same round of review. The assignment is triggered when the first mapping is created. The identifier is bound to the document version number, user account, timestamp, and environment signature to ensure data isolation and audit traceability between different sessions. Semantic relationship mapping operations are performed under this identifier, iteratively searching around clause paragraphs and returning the most relevant policy provisions and historical cases from the vector database in real time as contextual evidence. The search results and clause paragraphs form mapping entries, each recording the semantic path, evidence text, source link, and similarity distribution curve.
[0093] Based on the aforementioned contextual evidence, a related session context is generated. The session context is a semantic session snapshot organized around a unique session identifier, containing a set of clause paragraphs, corresponding evidence sets, interaction history, filtering conditions, recalculation strategies, and consistency verification records. The generation process packages clause anchors, evidence anchors, similarity, material type, and source authority into a serializable object, which can be used for front-end rendering of the evidence sidebar and for evidence injection in subsequent compliance checks. Finally, the related session context is bound and stored with the unique session identifier, employing a pre-write verification and versioned write strategy to avoid concurrent overwriting and cross-user contamination, while retaining incremental changes for rollback and auditing.
[0094] This embodiment renders the draft clauses as a structured clause document and enables structural stability in the editing interface, ensuring consistency between the editing behavior of chapters, lists, and tables and cross-references, thus reducing positioning errors caused by layout disturbances. During the review process, format parsing, preprocessing, and embedding vectorization are introduced, incorporating the collection of supporting materials and clause content into a unified semantic index space. Semantic retrieval directly connects to evidence fragments and page number location, shortening the path for manual retrieval and comparison. Through semantic relationships and unique session identifiers, traceable evidence linkage is established, generating associated session contexts containing mapping information. This provides immediate, computable, and replayable contextual support for subsequent compliance checks and problem localization, reducing redundant comparisons and evidence loss, and improving the stability and processing throughput of the review process.
[0095] In one embodiment, step S40 above includes: S401, Based on the contextual evidence in the aforementioned relationship, perform format constraint detection on the clause document to verify the completeness of clause sections and the compliance of section order; S402, Based on the contextual evidence in the association relationship, perform semantic constraint detection on the clause document to analyze the semantic compliance of the clause text; S403, Based on the contextual evidence in the aforementioned relationship, perform a logical consistency check on the terms document to detect logical conflicts between the scope of liability and the exclusions of liability; S404, Based on the results of format constraint detection, semantic constraint detection, and logical consistency detection, compile the set of pre-audit issues; S405, mark the abnormality level and violation basis for each issue item in the pre-examination issue item set; S406, Generate modification suggestions for each issue item in the pre-examination issue item set; S407, Associate the problem items in the pre-examination problem item set with the specific locations in the clause document to generate problem location information; S408, Generate an audit report based on the pre-audit issue set, anomaly level label, violation basis, modification suggestions, and issue location information.
[0096] In this embodiment, when performing compliance checks on the clause documents based on the relationship, the contextual evidence, clause paragraph anchors, and material source pointers stored in the relationship are first loaded. The clause text, evidence fragments, and metadata are mapped to the same working set, establishing a one-to-one or one-to-many correspondence between clause fragments and evidence fragments. Format constraint checks run on the working set to perform structural parsing and layout consistency verification. The heading level, numbering sequence, list nesting, table skeleton, and cross-reference keys are read and matched one by one with the structural template. The integrity and order of chapters are verified. Structural alarms are generated for missing chapters, hierarchical jumps, disordered numbering, broken tables, and mismatched references. Each alarm is bound to the paragraph anchor point at the location where it is generated and the corresponding evidence fragment, forming an initial annotation set at the structural level. Semantic constraint detection integrates a regulatory expression list, a prohibited terminology database, and semantic entries in the clauses. It first segments the clause text into semantic units and aligns them with standard terminology. Then, it cross-references contextual evidence to identify sentences or segments that may contain prohibited, ambiguous, or unauthorized expressions. For each triggering entry, it records the matching phrases, contextual windows, supported guidance clauses, and similarity elements, outputting an initial semantic annotation set. Logical consistency detection constructs a relationship graph of liability scope and exclusions within the clause document. It abstracts liability entries, exclusions, triggering conditions, and applicable restrictions as nodes, and inclusion, exclusion, dependence, and premises as edges. Combining definitions and applicable boundaries from contextual evidence, it performs consistency reasoning to identify coverage conflicts, double exclusions, unclosed dependencies, and cross-chapter condition mismatches. For each conflict, it generates a causal chain and a set of involved nodes, while also referencing evidence fragments and clause locations, outputting an initial logical annotation set. When merging the three initial annotation sets, a deduplication and merging strategy is adopted. Multi-source hits in the same paragraph and the same semantic position are merged, and only entries with more complete information and more sufficient evidence are retained. Weak hit prompts are downgraded to annotations, and finally a computable detection result view is obtained.
[0097] When compiling the pre-audit issue set based on the results of format constraint checks, semantic constraint checks, and logical consistency checks, a fixed set of fields is used for each issue item, including issue number, trigger type, clause paragraph anchor, evidence source list, trigger condition description, scope of impact, location path, and the targeted revision point. A merging and grouping mechanism is introduced during the set generation process to combine multiple manifestations of the same cause in different locations into a single main issue, and to hierarchically represent composite issues triggered by multiple checks in the same location, ensuring the list is easy to browse and execute. During the issue grading stage, each issue item in the pre-audit issue set is marked with an anomalous level. The anomalous level is determined comprehensively based on factors such as trigger type, evidence strength, scope of impact, regulatory seriousness, and remediability. The grading process retains the judgment basis and adjudication path for easy review. The violation basis field is bound to clause citations and material page numbers from contextual evidence, and language fragments are added when necessary for audit tracing.
[0098] When generating modification suggestions for each issue item, a dual-track approach of template-based and evidence-based generation is employed. The template-based path connects to the standard revision sentence library and structured change patterns, providing directly replaceable or guideable suggested texts for scenarios such as responsibility definition, exclusions, terminology standardization, table standardization, and cross-reference correction. It also specifies the expected change locations, the reference keys requiring synchronous updates, and potentially affected adjacent paragraphs. The evidence-based path extracts the policy provisions or historical expressions associated with the issue item into acceptable constraints or reference expressions, providing the basis text and definitional boundaries for revision. The two paths are merged into a single suggestion item at the output stage, including the suggestion text, suggestion type, applicable conditions, and evidence links.
[0099] When associating and locating issues in the pre-audit issue set with specific locations in the clause documents, multi-granularity positioning is achieved using paragraph anchors, byte offsets, and structural node paths. Each issue is bound to one or more editable locations, and replayable positioning instructions are generated to ensure one-click navigation to the precise location within the editing interface, forming issue positioning information. This positioning information includes not only location and scope but also positioning reliability, peer dependencies, and suggested execution order to reduce secondary searches and conflicts during revisions. Finally, an audit report is generated based on the pre-audit issue set, anomaly level labeling, violation basis, modification suggestions, and issue positioning information. The report uses a dual-track output of structured description and renderable fragments. The structured description is used for system integration and statistical analysis, while the renderable fragments are used for interface display and interaction. The report body provides an overall inspection overview, a list grouped by type, evidence links for each issue, positioning navigation, and suggested execution entry points, while also retaining an inspection configuration summary and evidence snapshot index to support compliance auditing.
[0100] This embodiment, supported by correlation and contextual evidence, coordinates format constraint detection, semantic constraint detection, and logical consistency detection, merging the three verification links of structure, semantics, and relationships into a consistent result view. This significantly reduces the time consumption and probability of oversights associated with manual searching and comparison. By standardizing the expression of the pre-review issue set and anomaly levels, the complex issue space is compressed into an executable revision list. Combined with the basis for violations and modification suggestions, revision actions have directly actionable textual and evidentiary sources. Relying on issue location information, each issue is precisely bound to the editing location, forming a direct path from discovery to handling. This improves the interpretability, traceability, and one-time fix rate of the review output, and provides a stable input base for subsequent online revisions, version snapshots, and re-reviews.
[0101] In one embodiment, step S50 above includes: S501, Load the problem location information in the editing interface, and display the problem paragraph in the clause document with a preset prominent mark; S502, Receive a revision operation instruction for the problematic paragraph based on the problem location information; S503, according to the revision operation instruction, perform text content revision operation in the editing interface, generate revision clause document, and synchronize revision operation data to version control system; S504 creates an incremental version snapshot for each revision operation to record the revision content and revision time information; S505 adds an operator identifier and revision type marker to each version snapshot; S506, the version snapshot is associated with the revised terms document and stored to form a version history.
[0102] In this embodiment, after the issue location information enters the editing interface, structural parsing and anchor point alignment are performed first. The information unit includes the issue number, paragraph anchor point or character range, exception level, violation basis, suggested text, and dependency tags. The parsing module maps the anchor points to node paths and offsets in the document object model and inserts preset prominent markers without altering the semantics of the main text. These prominent markers operate both in the visual layer and are written into the node metadata, including marker type, start and end range, source issue number, and timestamp, ensuring stable positioning during pagination, export, and reflow.
[0103] The editing interface presents revision entry points in a linked manner with the issue list and the main text. The issue list is grouped based on exception level and chapter level, and supports filtering by evidence source and responsibility category. When a user selects a target issue in the list, the main text view scrolls to the corresponding anchor point and highlights a prominent mark, while the revision panel expands in the side area. The revision panel loads the suggested text and the basis for the violation, and provides operation buttons such as replace, insert, delete, rearrange, cross-reference repair, and format standardization. All buttons are bound to the current document node path and offset to avoid accidentally triggering adjacent paragraphs.
[0104] Before submission, revision instructions undergo consistency and dependency checks. Consistency checks verify that heading levels, list nesting, consecutive numbering, table structure, and cross-reference keys meet editing guidelines. Dependency checks identify related issues requiring simultaneous processing based on dependency tags in the issue location information, such as changes in scope of responsibility and corresponding exceptions. After passing these checks, the editor generates a serializable revision instruction object. Fields include the target node path, range, operation type, operation load, rollback fragment, and reference key update table. The instruction object is first applied locally to the document object model to create a temporary version, and then simultaneously streamed to the version control system via a network connection and written to the operation log queue.
[0105] The generation of revised clause documents adopts a transactional commit. The transaction boundary covers four stages: text replacement or structural adjustment, number recalculation, cross-reference rewriting, and format rendering. Failure in any stage triggers a complete rollback, and the reason for failure and the investigation path are marked in the issue list. After a successful commit, a new document version is generated, and a one-to-many mapping is established between the version number and the issue number in the commit record to facilitate subsequent difference analysis and accountability tracking.
[0106] Incremental version snapshots are created immediately after a transaction commit. The snapshot generation module compares the document object model before and after the commit, outputting the minimum changeset and the set of affected reference keys, while also calculating a content summary and a structure summary. Each snapshot records a revision content summary, revision time information, and a set of source issue numbers, and generates a content hash value and a parent version pointer. To reduce storage pressure, the text is stored in the database using fragment-level deduplication and differential storage. A reverse changeset is generated when necessary to support rapid rollback.
[0107] Operator identifiers and revision type tags are attached together when the snapshot is written. Operator identifiers, sourced from the unified identity management system, include account, organization, and device summaries, are written into audit fields, and used for access control. Revision type tags are automatically categorized based on operational load and triggered detection types into terms such as terminology standardization, scope adjustment, exclusion updates, format patching, reference patching, and parameter caliber alignment. Multiple tags can be applied simultaneously to cover complex revisions. The tagging system aligns with the issue list categories, ensuring semantic consistency across reports and versions.
[0108] Version history is maintained using a directed graph structure. Each node represents a document version, and each edge represents an incremental commit, carrying the operator's identifier, revision type flag, and changeset summary. When parallel revisions occur, branches appear on the graph, and merge operations are represented by explicit merge nodes. The merge record includes a conflict list and a summary of the resolution strategy. Version history also maintains an index table, supporting retrieval and replay by issue number, operator, time window, revision type, or chapter anchor. Any historical node can be restored to a working version with a single click, while retaining a new snapshot of the restoration action to maintain the integrity of the audit chain.
[0109] An idempotent write and replay mechanism enhances reliability between the editing interface and the version control system. Each revision command carries a globally unique sequence number and a session signature. The version control system automatically deduplicates duplicate sequence numbers and performs delayed rearrangement for out-of-order commands. In the event of a network interruption, local commands are placed in a temporary queue and periodically retried; synchronization proceeds in order upon connection restoration. To prevent overwriting caused by cross-user collaboration, the version control system verifies the base version number before committing. If a lagging base version is detected, the user is prompted to merge or fetch the latest version before committing, preventing historical data loss.
[0110] Issue location information and prominent markers remain updated after revision. If a revision causes the original interval to shift or the text to be split, the anchor migration module adjusts the node path and offset according to the changeset to ensure accurate location even in subsequent reviews or continuous revisions. Resolved issues are automatically downgraded in the list and moved to the processed partition, while unresolved or newly emerging issues remain in the pending partition, supporting continuous iteration.
[0111] This embodiment utilizes a linked revision and versioned storage driven by problem location information to create a single closed loop for locating, modifying, submitting, and retrospectively reviewing issues, reducing repetitive manual retrieval and secondary comparisons. Significant markers and anchor metadata ensure accurate correspondence between the visual and data layers; revision operation instructions and transactional commits guarantee consistent implementation of text adjustments and structural updates; incremental version snapshots and operator identifiers construct a complete audit and accountability chain; and revision type markers provide a direct entry point for subsequent statistical analysis and focused re-examination. Dependency verification and reference rewriting reduce the risk of cascading errors, while idempotent writes and sequential replay enhance robustness in collaborative environments.
[0112] In one embodiment, step S70 above includes: S701, Perform a difference detection operation on the revised terms document based on the version history to identify the changed content; S702, Generate a clause change description document based on the changes; S703, highlight the key changes and corresponding compliance basis in the document explaining the changes to the aforementioned terms; S704, embed the clause change description document into the template reserved position in the filing material package; S705, verify the completeness and format compliance of the reported materials package, and output the verified reported materials package containing the clause change explanation document.
[0113] In this embodiment, the difference detection constructs the comparison range based on the version history. Each version in the version history has an associated version number, commit time, operator identifier, and incremental change set. The revised clause document serves as the current baseline and enters the comparison pipeline. Before comparison, the alignment of paragraph anchors, chapter number trajectories, cross-reference keys, and table cell paths is completed, and invalid changes caused solely by layout re-runs are removed to ensure that the comparison focuses on semantically valid intervals. Text units and structural units are compared in parallel. Text units cover titles, body paragraphs, list items, and notes, while structural units cover chapter levels, clause numbers, tables, and references. Change types are divided into four categories: addition, deletion, modification, and position adjustment. Modifications record both the original and current content intervals, while position adjustments record both the original and current paths. For each change, a change content entry is generated, carrying a position index, change type, preceding and following segments, the set of affected reference keys, and a version snapshot identifier of the backlink, forming a traceable difference detail.
[0114] The clause change explanation document is generated based on the changed content. During generation, it is first aggregated by chapter and topic, folding multiple minor changes under the same clause into a single topic entry, while retaining a detailed index to avoid reading redundancy. Each topic entry includes the original fragment, the revised fragment, a description of the reason for the change, the scope of impact, dependencies, and the source of evidence. The description of the reason for the change is extracted from previous audit results and contextual evidence, such as semantic constraint trigger points, logical consistency verification conclusions, or parameter alignment explanations. Key modification points are highlighted with visual markers, and the highlighted areas maintain an anchor point relationship with the clause location, ensuring correct positioning across pages and during export. Compliance basis is presented as clause paths and evidence fragments, along with the location information and time stamp of the source materials, facilitating audit review. The document as a whole adopts a structured rich text organization, unifying heading levels, numbering systems, and cross-references to ensure consistent rendering on the editing and submission ends.
[0115] The embedding process writes the clause change description document into the template reserved space in the reporting materials package. The template reserved space consists of three parts: bookmark anchors, placeholders, and layout styles. During embedding, the template version and the validity of the reserved space are first verified, and then the document fragments are mapped and inserted according to the predefined styles. After embedding, the table of contents index and cross-references are rebuilt to ensure that new entries are correctly covered by the table of contents and headers and footers. If the template contains multilingual or multi-regional columns, the embedder generates parallel fragments according to regional characteristics and binds them to the reserved spaces of the corresponding columns, while maintaining consistent entry numbering.
[0116] Completeness and format compliance checks are performed after the materials are assembled. Completeness checks verify each item in the material list, checking the completeness of the full text of clauses, revised clause documents, clause change explanation documents, re-audit reports, and supporting evidence lists, and verifying the consistency of version number references. Format compliance checks focus on layout and structural key points, checking the continuity of heading levels, the completeness of numbering intervals, the correct row breaks across pages in tables, the absence of missing cross-references, and the conformity of margin information to the template style. All defects during the check process are output as location information and correction suggestions; blocked items are prevented from being exported, and warning items are recorded in the material check log for tracking.
[0117] In the output phase, a unified encapsulation is generated for the reported materials package. The encapsulation list includes the material name, version number, time stamp, and verification summary. Links between the clause change description document and the revised clause document are registered in the encapsulation table to ensure direct navigation at the receiving end. After encapsulation, a verification report summary and hash digest are generated for easy transmission verification and subsequent archiving comparison. The entire process maintains unidirectional traceability: any change can be linked back to a specific snapshot in the version history, and any compliance evidence can be linked back to the source and location of the evidence.
[0118] This embodiment transforms changes in revised clause documents into locatable, aggregateable, and auditable changes through version history-driven difference detection. The clause change description document then combines key modifications with compliance evidence, reducing repetitive manual verification and evidence collection. By embedding pre-reserved spaces in the template and conducting dual checks on completeness and format compliance, the assembled materials achieve structural consistency and evidence closure before output. When change items can be linked to version snapshots and evidence sources, the reporting material package has direct jump and one-click verification capabilities during submission and review, thereby shortening the cycle of repeated material returns, reducing the probability of omissions and contradictions, and improving the operability of cross-version comparison and compliance auditing.
[0119] In one embodiment, a clause generation, review, and filing device is provided, which corresponds one-to-one with the clause generation, review, and filing methods described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the clause generation, review, and filing device of the present invention. The modules include: element parsing module 10, clause generation module 20, editing and association module 30, compliance detection module 40, revision processing module 50, re-review and filing module 60, and change description module 70. Detailed descriptions of each functional module are as follows: Element analysis module 10 is used to receive demand input information and a set of auxiliary materials, analyze the demand input information to extract key elements, and form an element set; Clause generation module 20 is used to retrieve matching information from the knowledge base based on the element set to form a matching result, and generate a draft clause according to the rich text structure based on the matching result; The editing association module 30 is used to load the draft clauses into the editing interface to form a clause document, and to establish the association between the clause document and the set of supporting materials during the review process; The compliance detection module 40 is used to perform compliance detection on the clause document based on the aforementioned relationship, and generate an audit report and issue location information based on the compliance detection results. The revision processing module 50 is used to perform revision operations on the editing interface based on the problem location information and generate a revised clause document; The re-examination and filing module 60 is used to perform a re-examination operation on the revised clause document and generate a re-examination result. When the re-examination result meets the pass conditions, a filing material package is generated. The change description module 70 is used to extract the change content of the revised clause document to generate a clause change description, embed the clause change description into the filing material package, and output a filing material package containing the clause change description.
[0120] For a preferred embodiment of the clause generation and review filing device, please refer to the foregoing limitations on the clause generation and review filing method, which will not be repeated here. Each module in the aforementioned clause generation and review filing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0121] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When executed by the processor, the computer program implements the functions or steps of a terms generation and review reporting method on the server side.
[0122] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements the client-side functions or steps of a clause generation and review reporting method.
[0123] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Receive demand input information and a set of supporting materials, parse the demand input information to extract key elements, and form an element set; Based on the set of elements, matching information is retrieved from the knowledge base to form a matching result, and a draft clause is generated according to the matching result in a rich text structure; The draft terms are loaded into the editing interface to form a terms document, and the association between the terms document and the set of supporting materials is established during the review process; Based on the aforementioned relationships, a compliance check is performed on the terms and conditions document, and an audit report and issue location information are generated based on the compliance check results. Based on the problem location information, perform revision operations in the editing interface to generate a revised clause document; A second review operation is performed on the revised clause document to generate a review result. When the review result meets the approval conditions, a filing material package is generated. Extract the changes from the revised clause document to generate a clause change description, embed the clause change description into the filing material package, and output the filing material package containing the clause change description.
[0124] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Receive demand input information and a set of supporting materials, parse the demand input information to extract key elements, and form an element set; Based on the set of elements, matching information is retrieved from the knowledge base to form a matching result, and a draft clause is generated according to the matching result in a rich text structure; The draft terms are loaded into the editing interface to form a terms document, and the association between the terms document and the set of supporting materials is established during the review process; Based on the aforementioned relationships, a compliance check is performed on the terms and conditions document, and an audit report and issue location information are generated based on the compliance check results. Based on the problem location information, perform revision operations in the editing interface to generate a revised clause document; A second review operation is performed on the revised clause document to generate a review result. When the review result meets the approval conditions, a filing material package is generated. Extract the changes from the revised clause document to generate a clause change description, embed the clause change description into the filing material package, and output the filing material package containing the clause change description.
[0125] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0126] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0127] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0128] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A clause generation and audit reporting method, characterized in that, The method comprises the following steps: receiving demand input information and auxiliary material set, analyzing the demand input information to extract key elements, forming an element set; based on the element set, retrieving matching information from the knowledge base to form a matching result, and generating a draft clause according to the matching result in a rich text structure; loading the draft clause into an editing interface to form a clause document, and establishing an association between the clause document and the auxiliary material set during the review process; based on the association, performing compliance detection on the clause document, and generating a review report and problem positioning information according to the compliance detection result; performing revision operations on the editing interface according to the problem positioning information, and generating a revised clause document; performing a re-review operation on the revised clause document to generate a re-review result, and generating a submission material package when the re-review result meets the pass condition; extracting the change content of the revised clause document to generate a clause change description, embedding the clause change description in the submission material package, and outputting the submission material package containing the clause change description.
2. The clause generation and audit placement method of claim 1, wherein, receiving demand input information and auxiliary material set, analyzing the demand input information to extract key elements, forming an element set, comprising: receiving natural language description input containing product feature description, and receiving file upload input containing policy file and historical clause; performing semantic intent recognition and entity extraction on the natural language description input to generate a structured semantic analysis result; performing format analysis and content extraction on the file upload input to generate a multi-source file analysis result; fuse the structured semantic analysis result and the multi-source file analysis result to generate a preliminary key element set; performing integrity verification on the preliminary key element set, and performing compliance verification on the verified preliminary key element set based on the regulatory policy library; classify the key elements that pass the compliance verification according to product type, agreement subject and regional characteristics to generate an element set containing classification tags and attributes.
3. The clause generation and audit placement method of claim 1, wherein, based on the element set, retrieving matching information from the knowledge base to form a matching result, and generating a draft clause according to the matching result in a rich text structure, comprising: based on the element set, retrieving matching content from the policy and regulation database of the knowledge base to form a policy and regulation matching subset; based on the element set, retrieving matching content from the historical clause template database of the knowledge base to form a historical clause matching subset; based on the element set, retrieving matching content from the actuarial parameter database of the knowledge base to form an actuarial parameter matching subset; fuse the policy and regulation matching subset, the historical clause matching subset and the actuarial parameter matching subset to generate a comprehensive matching result; determine the clause structure framework based on the comprehensive matching result; fill in the specific content of each chapter of the initial draft clause according to the clause structure framework, and insert cross-reference marks during the content filling process; perform logical consistency verification on the filled draft clause, format the draft clause that passes the logical consistency verification according to the rich text structure to generate the final draft clause.
4. The clause generation and audit placement method of claim 1, wherein, loading the clause draft into an editing interface to form a clause document, establishing an association between the clause document and the auxiliary material set during the review process, including: loading the clause draft into an editing interface of a rich text editor and performing rich text format rendering processing on the loaded clause draft to form a structured clause document; enabling multi-level heading automatic numbering, list nesting, and table cross-page line breaking functions in the editing interface; during the review process, performing format analysis on the multi-format documents in the auxiliary material set to extract original text content; performing cleaning, segmentation, and deduplication processing on the original text content, and converting the processed text content into high-dimensional vectors through an embedding model; storing the high-dimensional vectors into a vector database and constructing a unified semantic index space; based on the unified semantic index space, establishing a semantic association between the structured clause document and the auxiliary material set; during the establishment of the semantic association, assigning a unique session identifier to the association; based on the unified semantic index space, performing semantic association mapping to retrieve the most relevant policy provisions and historical cases in the vector database as contextual evidence in real time; based on the contextual evidence, generating an associated session context containing semantic association mapping information; binding and storing the associated session context with the unique session identifier.
5. The clause generation and audit placement method of claim 1, wherein, based on the association, performing compliance detection on the clause document, and generating a review report and problem positioning information based on the compliance detection results, including: based on the contextual evidence in the association, performing format constraint detection on the clause document to verify clause chapter integrity and chapter order compliance; based on the contextual evidence in the association, performing semantic constraint detection on the clause document to analyze clause text semantic compliance; based on the contextual evidence in the association, performing logical consistency detection on the clause document to detect logical conflicts between responsibility scope and excluded responsibilities; based on the results of format constraint detection, semantic constraint detection, and logical consistency detection, compiling a pre-review issue item set; labeling each issue item in the pre-review issue item set with an abnormality level and a violation basis; generating a modification suggestion for each issue item in the pre-review issue item set; associating and positioning each issue item in the pre-review issue item set with a specific location in the clause document to generate problem positioning information; generating a review report based on the pre-review issue item set, abnormality level labeling, violation basis, modification suggestion, and problem positioning information.
6. The clause generation and audit placement method of claim 1, wherein, performing revision operations in the editing interface according to the problem positioning information to generate a revised clause document, including: loading the problem positioning information in the editing interface and displaying the problem paragraph with a pre-set prominent marker in the clause document; receiving revision operation instructions based on the problem positioning information performed on the problem paragraph; performing text content revision operations in the editing interface according to the revision operation instructions to generate a revised clause document, and synchronizing revision operation data to a version control system; creating an incremental version snapshot for each revision operation to record revision content and revision time information; adding an operator identification and a revision type label to each version snapshot; storing the version snapshots in association with the revision clause document to form a version history record.
7. The clause generation and audit placement method of claim 1, wherein, extracting change content of the revision clause document to generate a clause change description, embedding the clause change description into the filing package, and outputting the filing package containing the clause change description, including: performing a difference detection operation on the revision clause document based on the version history record to identify change content; generating a clause change description document according to the change content; highlighting key modification points and corresponding compliance bases in the clause change description document; embedding the clause change description document into a template reserved position in the filing package; verifying integrity and format compliance of the filing package, and outputting the verified filing package containing the clause change description document.
8. A clause generation and audit reporting device, characterized by, The clause generation and review filing device includes: an element analysis module configured to receive demand input information and a set of auxiliary materials, analyze the demand input information to extract key elements, and form an element set; a clause generation module configured to retrieve matching information from a knowledge base based on the element set to form a matching result, and generate a clause draft in a rich text structure according to the matching result; an editing association module configured to load the clause draft into an editing interface to form a clause document, and establish an association relationship between the clause document and the set of auxiliary materials during the review process; a compliance detection module configured to perform compliance detection on the clause document based on the association relationship, and generate a review report and problem positioning information according to the compliance detection result; a revision processing module configured to perform a revision operation in the editing interface according to the problem positioning information, and generate a revised clause document; a re-review filing module configured to perform a re-review operation on the revised clause document to generate a re-review result, and generate a filing package when the re-review result meets a passing condition; a change description module configured to extract change content of the revised clause document to generate a clause change description, and embed the clause change description into the filing package, and output the filing package containing the clause change description.
9. A computer device, comprising: The computer device includes a memory, a processor, and a clause generation and review filing program stored on the memory and executable on the processor, and the clause generation and review filing program, when executed by the processor, implements the steps of the clause generation and review filing method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a clause generation and review filing program, and the clause generation and review filing program, when executed by the processor, implements the steps of the clause generation and review filing method according to any one of claims 1-7.
Citation Information
Cited By
Audit report automatic generation method based on natural language processing
CN121997896A
Audit method, device and apparatus for vulnerability mechanism
CN122333490A