A method and system for intelligent identification of contract risks
By building a review knowledge graph and a standardized template library, combined with semantic vector matching and stance recognition models, contract terms are automatically analyzed, solving the problems of low efficiency and poor accuracy in existing technologies, and achieving efficient and comprehensive contract risk identification and compliance checks.
Patent Information
- Application Number
- CN202511045556.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-29
AI Technical Summary
In existing technologies, the identification of contract clause risks mainly relies on the experience of professionals, resulting in low efficiency and difficulty in ensuring accuracy, and making it impossible to effectively identify and avoid contract risks.
By building a review knowledge graph and a standardized contract template library, using semantic vectors to match the knowledge graph, automatically analyzing contract terms, identifying potential risks and conducting multi-dimensional compliance checks, and combining position recognition models and disagreement resolution analysis, efficient and comprehensive risk identification and assessment can be achieved.
It improves the efficiency and accuracy of risk identification of contract terms, realizes automated, objective and comprehensive risk identification, reduces reliance on manual experience, ensures that contracts comply with regulations and corporate rules, and enhances the structural integrity and logical fluency of contracts.
Smart Images

Figure CN120542439B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of text recognition technology, and in particular to a method and system for intelligently identifying contract risks. Background Art
[0002] Contracts are widely used in various fields of production and life because of their strong binding force and flexible adaptability. A complete contract consists of several contract clauses that clearly define the rights and obligations of both parties. The rigor of the contract clauses directly affects the effectiveness and feasibility of a contract.
[0003] However, due to the particularity of the contract, the long-term nature of the contract, and the diversity and complexity of the contract performance, the risk responsibilities contained in the contract terms are unavoidable for both parties to the contract. Therefore, it is necessary to identify the risks of the contract terms before signing the contract in order to avoid risks.
[0004] Currently, the risk assessment of contract clauses relies primarily on professionals' expertise, practical experience, the needs of the contracting parties, and current regulations to determine whether a contract clause poses a risk. This is a time-consuming and labor-intensive process, not only creating a huge workload for the relevant personnel but also reducing the efficiency of the entire process. Furthermore, different professionals have different levels of execution, making it difficult to ensure the accuracy of risk identification. Summary of the Invention
[0005] In order to solve the problem that existing contract clause risk identification methods mainly rely on the professional experience of professionals, resulting in low risk identification efficiency and low accuracy, the present invention provides a contract risk intelligent identification method, which includes:
[0006] Obtain a file to be identified, pre-process the file to be identified to obtain a first file, obtain the contract terms of the first file, and extract the semantic vectors of the contract terms; construct a review knowledge graph and a review list; obtain a first risk identification result based on the semantic vector and the review list; obtain the contract category of the first file based on the semantic vector; semantically match the semantic vector and the review knowledge graph to obtain potential risks; obtain a first risk category based on the potential risks and the review knowledge graph, obtain a preset risk judgment rule based on the first risk category, and obtain a second risk category based on the preset risk judgment rule; obtain a second risk identification result based on the contract category and the second risk category; construct a clause rule library, and obtain a compliance result based on the semantic vector and the clause rule library; construct different standardized contract template libraries according to different contract types, and obtain a complete result based on the semantic vector and the standardized contract template library, wherein the standardized contract template library includes several lists of essential contract terms and the clause statements corresponding to the lists of essential contract terms; obtain a total risk identification result based on the first risk identification result, the second risk identification result, the compliance result, and the complete result.
[0007] This method performs unified preprocessing on contract documents to be identified, improving post-identification efficiency and maintainability through a unified data processing process. Semantic vectors, such as keywords, semantic vectors, and contextual features, are extracted from the clause text. A review knowledge graph and review checklist are constructed. A customized review checklist can quickly integrate external regulations, internal corporate rules, and regulatory requirements to identify potential risks and non-compliant clauses, improving review speed and accuracy and providing support for contract review and decision-making. Semantic vectors are semantically matched with rule nodes in the knowledge graph and mapped with a predefined risk category feature library in the knowledge graph. Based on the matching results, potential risks are identified by reasoning through relationships in the graph (such as violation, non-conformity, and possible consequences). Identified risks are classified according to dimensions such as commercial risk, performance risk, and reputation risk to obtain a first risk category. Within this risk category, corresponding professional domain knowledge subgraphs are automatically invoked for different contract types to further refine the risk category. A hierarchical classification model is used to first determine the main risk category and then subdivide it into subcategories, narrowing the risk range and improving the accuracy and professionalism of risk identification. Contract clause risks are further analyzed from dimensions such as integrity and compliance, providing more comprehensive risk identification.
[0008] This method utilizes the review knowledge graph, review checklist, clause rule library and standardized contract template library to realize multi-dimensional automated analysis of contract clauses. It does not need to rely on manual experience review, and risk identification is more efficient, objective and comprehensive.
[0009] Furthermore, the specific steps of obtaining compliance results based on the semantic vector and the clause rule library include: obtaining several semantically identical clauses based on the semantic vector, extracting key elements of each semantically identical clause, judging whether the key elements are consistent, and if inconsistent, obtaining conflict results based on the key elements and the conflict table; matching the contract clauses with the clause rule library to obtain a first compliance result; matching the contract clauses with the preset industry rule library to obtain a second compliance result; matching the contract clauses with the custom rule library to obtain a third compliance result; matching the contract clauses with the historical dispute rule library to obtain a fourth compliance result; obtaining the compliance result based on the first compliance result, the second compliance result, the third compliance result and the fourth compliance result.
[0010] Clause conflict detection determines whether contract clauses conflict through semantic relevance and key elements. It conducts multi-dimensional compliance checks, including regulatory compliance, industry standard compliance, internal corporate compliance, and historical case comparisons, to ensure that the contract complies with regulations and corporate rules. The inspection results are more comprehensive and accurate.
[0011] Furthermore, the specific steps of obtaining a complete result based on the semantic vector and the standardized contract template library include:
[0012] Based on the contract category, the contract terms are matched with the standardized contract template library to obtain a first missing result; similarity is calculated between the contract terms and the standardized contract template library to obtain semantic similarity, and a second missing result is obtained based on the semantic similarity; based on the first missing result and the second missing result, the missing terms are obtained; the historical contract position of the missing terms is obtained, and the optimal revision position of the missing terms is obtained based on the historical contract position; the complete result is obtained based on the missing terms and the optimal revision position.
[0013] Build a standardized contract template library, and establish a standardized template library based on different contract types, which includes a list of necessary terms for various contracts and their standard expressions. A two-way detection mechanism is adopted. On the one hand, structured analysis is used to detect whether the contract contains the necessary terms in the template library. On the other hand, semantic similarity calculation is used to identify whether the terms exist in different expressions. The detection is more comprehensive and accurate. For missing terms detected, their optimal insertion position is determined based on their historical contract position, thereby improving the structural integrity and logical fluency of the contract.
[0014] Furthermore, if the first risk category is commercial risk, the preset risk judgment rules include: obtaining commercial elements based on the semantic vector, the commercial elements including price terms, delivery terms, and restrictive terms, the restrictive terms including exclusivity clauses, non-competition clauses, confidentiality clauses, and unilateral termination clauses; obtaining benefits and costs of different contract parties based on the commercial elements, and determining whether the benefits and costs are balanced based on a first preset range; obtaining a contract price based on the commercial elements, and determining whether the contract price is abnormal based on a preset average price level and a preset floating range; obtaining a payment node, delivery node, down payment ratio, and payment period based on the commercial elements, determining whether the down payment ratio exceeds a preset ratio, determining whether the payment node and delivery node match based on a preset time range, and determining whether the payment period is abnormal based on a preset period; obtaining a form of liability for breach of contract based on the commercial elements, and determining whether the form of liability for breach of contract is abnormal based on a preset breach of contract amount; determining whether the restrictive terms are abnormal based on preset restriction data; obtaining potential restriction data based on the restrictive terms, obtaining a restriction score based on a preset restriction table and the potential restriction data, and determining whether it is abnormal based on the restriction score and a preset threshold.
[0015] By identifying commercial elements such as price terms, payment terms, and delivery terms, and breaking them down into subcategories, we can assess commercial unfairness and determine whether it imposes unreasonable economic burdens or commercial restrictions on our party.
[0016] Furthermore, the method also includes: analyzing the contract terms based on a position recognition model to obtain unbalanced terms; obtaining responsibility elements based on the unbalanced terms, constructing a responsibility-obligation balance matrix based on the responsibility elements, and obtaining the distribution ratio of responsibilities and powers of the different contract parties based on the responsibility-obligation balance matrix, wherein the responsibility elements include responsible parties, responsibilities and obligations, exercise of powers, and power restrictions; constructing an unfavorable terms indicator vocabulary, matching the semantic vector with the unfavorable terms indicator vocabulary to obtain target vocabulary, obtaining target text based on the target vocabulary and a preset text range, obtaining conditional relationships based on the target text, and obtaining semantic roles of the target text based on the conditional relationships; matching the target vocabulary with preset language elements to obtain semantic modifiers, wherein the preset language elements include qualifiers, modifiers, and exceptions; constructing a professional domain knowledge base and an industry standard terms library, simulating the perspectives of different contract parties based on the professional domain knowledge base, the industry standard terms library, the distribution ratio, the semantic roles, and the semantic modifiers, analyzing the unbalanced terms to obtain evaluation results, and obtaining biased terms based on the evaluation results.
[0017] The position identification model is used to analyze the position of the clause text and identify the clause's tendency towards the parties to the contract; a responsibility-obligation balance matrix is constructed to quantitatively analyze the distribution ratio of responsibilities and rights of each party; risk assessment is triggered by keywords using the adverse clause indicator vocabulary, which not only analyzes the target vocabulary itself but also considers the preceding and following texts to capture the complete semantic environment and make the analysis more accurate; the clause content is interpreted from the perspective of different contract parties, and the different impacts of the same statement on different parties are evaluated to achieve a more comprehensive and neutral risk assessment.
[0018] Furthermore, the method also includes: matching the semantic vector with preset disagreement words to obtain a dispute resolution clause; matching the dispute resolution clause with preset resolution words to obtain an integrity result; matching the dispute resolution clause with preset fuzzy words to obtain an enforceability result; matching the dispute resolution clause with preset special words to obtain a compliance result; obtaining a rationality result based on the integrity result, the enforceability result and the compliance result; extracting the geographic location entity and dispute resolution method of the dispute resolution clause, obtaining the association between the geographic location entity and the dispute resolution method, and obtaining the jurisdiction based on the association and the geographic location entity; obtaining the address information of the contract subject of each party based on the semantic vector, and obtaining the disadvantage degree result of the contract subject of each party based on the address information and the jurisdiction; obtaining the jurisdiction compliance result of the jurisdiction based on the address information, the geographic location entity and the jurisdiction; obtaining the dispute resolution result based on the rationality result, the disadvantage degree result and the jurisdiction compliance result.
[0019] The contract text is matched with preset disagreement words to locate the relevant clauses for dispute resolution, and the identified dispute resolution clauses are evaluated in multiple dimensions. Completeness: Check whether the necessary elements such as dispute resolution method, initiation procedure, and time limit are included; Enforceability: Evaluate whether the clause is clear, specific, and unambiguous, and reduce vague expressions; Compliance: Check whether the dispute resolution method complies with the regulations. Extract geographic location entities from the clauses, identify the relationship between the geographic entity and the dispute resolution method, and determine whether it is a place for negotiation or a place for dispute resolution; combine information such as the company's registration place and the place of contract performance to evaluate the degree to which the designated dispute resolution place is favorable or unfavorable to the company; and verify whether the designated jurisdiction complies with the regulations. The review of dispute resolution clauses not only determines whether the dispute resolution clause exists or is reasonably described, but also determines whether the dispute resolution address or dispute resolution location exists and is compliant, thereby improving the comprehensiveness and accuracy of the review.
[0020] Furthermore, the semantic vector of the contract terms is extracted based on the semantic model. In accordance with the calling order, the semantic model includes an encoding module, a multi-granularity self-attention module, a hierarchical progressive encoding module and a semantic integration module. The encoding module is used to encode the contract document according to the hierarchical structure, the multi-granularity self-attention module is used to extract global semantic features, the hierarchical progressive encoding module is used to extract local semantic features, and the semantic integration module is used to fuse features and generate a semantic vector. The multi-granularity self-attention module includes a nested self-attention layer, and the nested self-attention layer includes a hierarchical attention mechanism and a relative position-aware attention mechanism.
[0021] Furthermore, the hierarchical progressive encoding module includes a plurality of cross-layer connection layers, the cross-layer connection layers include a cross-layer connection mechanism, and based on a preset number of interval layers, the cross-layer connection layers are connected through residual connections; the semantic integration module includes a gating mechanism.
[0022] The introduction of nested self-attention layers and cross-layer connection mechanisms enables the model to process local and global semantic features simultaneously. A hierarchical attention mechanism is adopted to divide long texts into multiple semantic blocks, and attention calculations with different parameters are applied within and between blocks, which greatly reduces the computational complexity. In terms of attention mechanism, a relative position-aware attention mechanism is introduced to better capture long-distance dependencies, which is especially suitable for remote reference parsing between clauses in contract texts. In terms of position encoding, a segmented recursive position encoding is adopted to encode documents according to a hierarchical structure, maintaining the relative position information between different levels, and effectively solving the gradient vanishing problem of traditional position encoding when processing ultra-long texts. By optimizing the attention mechanism and position encoding method, the long text processing capability can be significantly improved.
[0023] Furthermore, the semantic model also includes a multi-level semantic feature extraction module, which is connected to the hierarchical progressive encoding module and the semantic integration module. The multi-level semantic feature extraction module includes a surface semantic extraction layer, a deep semantic fusion layer and a latent semantic reasoning layer. The surface semantic extraction layer is used to extract text expressions and keyword features, the deep semantic fusion layer is used to extract contextual semantics, and the latent semantic reasoning layer is used to construct a sentence relationship graph.
[0024] A multi-layered semantic feature extraction mechanism was introduced, comprising three layers: surface semantic extraction, deep semantic fusion, and latent semantic inference. The surface semantic extraction layer uses the bag-of-words model and word embedding techniques to capture direct textual representations; the deep semantic fusion layer utilizes a bidirectional LSTM combined with a self-attention mechanism to capture contextual semantics; and the latent semantic inference layer uses a graph neural network to construct a relationship graph between sentences and infer implicit regulatory concepts and relationships through a message passing mechanism. Features from each layer are selectively fused through a gating mechanism to form a comprehensive semantic representation.
[0025] The present invention also provides a contract risk intelligent identification system, the system comprising:
[0026] Preprocessing module: used for obtaining a file to be identified, and preprocessing the file to be identified to obtain a first file;
[0027] Extraction module: used for obtaining the contract terms of the first document and extracting the semantic vectors of the contract terms;
[0028] Construction module: used to build review knowledge graph and review checklist;
[0029] A first identification module: configured to obtain a first risk identification result based on the semantic vector and the review list;
[0030] A second identification module is configured to obtain the contract category of the first document based on the semantic vector; and to perform semantic matching between the semantic vector and the review knowledge graph to obtain potential risks; obtain a first risk category based on the potential risks and the review knowledge graph, obtain a preset risk judgment rule based on the first risk category, and obtain a second risk category based on the preset risk judgment rule; and obtain a second risk identification result based on the contract category and the second risk category.
[0031] A third recognition module is configured to construct a clause rule library and obtain compliance results based on the semantic vector and the clause rule library; and to construct different standardized contract template libraries according to different contract types and obtain complete results based on the semantic vector and the standardized contract template library, wherein the standardized contract template library includes a list of essential contract terms and the clause expressions corresponding to the list of essential contract terms;
[0032] Result module: used to obtain a total risk identification result based on the first risk identification result, the second risk identification result, the compliance result and the complete result.
[0033] The principles and effects of this system are similar to those of this method, so no further description will be given of this system.
[0034] The one or more technical solutions provided by the present invention have at least the following technical effects or advantages:
[0035] 1. This method performs unified preprocessing on contract documents to be identified, improving post-identification efficiency and maintainability through a unified data processing process. Semantic vectors, such as keywords, semantic vectors, and contextual features, are extracted from the clause text. A review knowledge graph and review checklist are constructed. The customized review checklist can quickly integrate external regulations, internal corporate rules, and regulatory requirements to identify potential risks and non-compliant clauses, improving review speed and accuracy and providing support for contract review and decision-making. Semantic vectors are semantically matched with rule nodes in the knowledge graph and mapped with a predefined risk category feature library in the knowledge graph. Based on the matching results, potential risks are identified by reasoning based on relationships in the graph (such as violation, non-conformity, and possible consequences). Identified risks are classified according to dimensions such as commercial risk, performance risk, and reputational risk to obtain a first risk category. Within this risk category, corresponding professional domain knowledge subgraphs are automatically invoked for different contract types to further refine the risk category. A hierarchical classification model is used to first identify the main risk category and then break it down into subcategories, narrowing the risk range and improving the accuracy and professionalism of risk identification. Contract clause risks are further analyzed from the perspectives of integrity and compliance, providing more comprehensive risk identification. It also uses review knowledge graphs, review checklists, clause rule libraries, and standardized contract template libraries to achieve multi-dimensional automated analysis of contract terms, without relying on manual experience review. Risk identification is more efficient, objective, and comprehensive.
[0036] 2. Clause conflict detection: This function uses semantic relevance and key elements to determine whether contract clauses conflict. This system performs multi-dimensional compliance checks, including regulatory compliance, industry norms compliance, internal corporate compliance, and historical case comparisons, to ensure that contracts comply with regulations and corporate rules. The results are more comprehensive and accurate.
[0037] 3. Build a standardized contract template library. A standardized template library is established based on different contract types, including a list of required terms for various contracts and their standard expressions. A two-way detection mechanism is adopted. On the one hand, structured analysis is used to detect whether the contract contains the necessary terms in the template library. On the other hand, semantic similarity calculation is used to identify whether the terms exist in different expressions. The detection is more comprehensive and accurate. For missing terms detected, their optimal insertion position is determined based on their historical contract position, thereby improving the structural integrity and logical fluency of the contract.
[0038] 4. Conduct position analysis on the clause text using a position identification model to identify the clause's bias towards the contracting parties; construct a responsibility-obligation balance matrix to quantitatively analyze the distribution ratio of responsibilities and rights among the parties; conduct risk assessment by triggering keywords using a vocabulary of adverse clause indicators, analyzing not only the target vocabulary itself but also the surrounding text to capture the complete semantic context for more accurate analysis; interpret the clause content from the perspective of different contracting parties, assessing the different impacts of the same statement on different parties, and achieving a more comprehensive and neutral risk assessment.
[0039] 5. Use preset disagreement words to match in the contract text, locate the clauses related to dispute resolution, and conduct a multi-dimensional assessment of the identified dispute resolution clauses. Completeness: Check whether it contains necessary elements such as dispute resolution methods, initiation procedures, and time limits; Enforceability: Evaluate whether the terms are clear, specific, and unambiguous, and reduce vague expressions; Compliance: Check whether the dispute resolution method complies with regulations. Extract geographic location entities from the clauses, identify the relationship between the geographic entity and the dispute resolution method, and determine whether it is a place for negotiation or a place for dispute resolution; combine information such as the company's registration place and the place of contract performance to evaluate the degree to which the designated dispute resolution place is favorable or unfavorable to the company; and verify whether the designated jurisdiction complies with regulations. The review of dispute resolution clauses not only determines whether the dispute resolution clauses exist or are reasonably described, but also determines whether the dispute resolution address or dispute resolution location exists and is compliant, thereby improving the comprehensiveness and accuracy of the review.
[0040] 6. The introduction of nested self-attention layers and cross-layer connection mechanisms enables the model to process local and global semantic features simultaneously. A hierarchical attention mechanism is adopted to divide long texts into multiple semantic blocks, and attention calculations with different parameters are applied within and between blocks, which greatly reduces the computational complexity. In terms of attention mechanism, a relative position-aware attention mechanism is introduced to better capture long-distance dependencies, which is particularly suitable for remote reference parsing between clauses in contract texts. In terms of position encoding, a segmented recursive position encoding is adopted to encode documents according to a hierarchical structure, maintaining the relative position information between different levels, and effectively solving the gradient vanishing problem of traditional position encoding when processing ultra-long texts. By optimizing the attention mechanism and position encoding method, the long text processing capability can be significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of the present invention, and do not constitute a limitation of the embodiments of the present invention;
[0042] Figure 1 This is a flow chart of a method for intelligently identifying contract risks in the present invention;
[0043] Figure 2 It is a structural diagram of the semantic model;
[0044] Figure 3 It is a schematic diagram of the structure of the review model;
[0045] Figure 4 This is a schematic diagram of the Transformer encoder structure;
[0046] Figure 5 It is a structural diagram of the MOE module. DETAILED DESCRIPTION
[0047] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present invention and the features therein can be combined with each other without conflict.
[0048] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.
[0049] Example 1
[0050] refer to Figure 1 This embodiment provides a method for intelligently identifying contract risks, the method comprising:
[0051] Obtain a file to be identified, and pre-process the file to be identified to obtain a first file. In this embodiment, the pre-processing method may include:
[0052] Text format standardization: Convert contract documents in various formats, such as doc, docx, txt, and pdf, into a unified format. For example, use an existing dedicated document parsing engine to convert different formats into UTF-8 encoded plain text strings, and use OCR technology to recognize and convert doc, docx, and pdf. For tabular data, use existing conversion technology to retain its structured features and convert it into a standard JSON format. For special symbols and professional terms, establish a mapping dictionary for unified processing.
[0053] Document structure parsing: Contract structure parsing is performed using a combination of rule-based and deep learning methods. For example, the hierarchical structure of the document (title, chapter, clause, item, paragraph, etc.) is first identified, and then a tree-like data structure is constructed to represent the hierarchical relationship of the document. For different types of contracts (such as procurement contracts, labor contracts, and lease contracts), specialized structural templates are constructed to enable adaptive identification of specific structural elements of the contract (such as the Party A information area, the Party B information area, the contract body, and attachments). The parsed data is stored in XML format, preserving the structural information and semantic associations of the original document.
[0054] Paragraph Segmentation and Numbering: Based on the contract text's natural delimiters (such as line breaks, punctuation, and paragraph markers) and semantic integrity (e.g., pre-trained semantic models), the text is segmented into reasonable paragraph units. Each paragraph is assigned a unique number to facilitate subsequent processing and reference. Specifically, potential paragraph boundaries are preliminarily marked using natural delimiters (such as line breaks, periods, and other punctuation). For each potential paragraph formed by a candidate segmentation point, a pre-trained semantic model is used to calculate its internal semantic cohesion score (e.g., using the average cosine similarity of semantic vectors between sentences within the paragraph). The semantic coupling between adjacent potential paragraphs is also assessed. Segmentation points that result in highly cohesive and less coupled paragraphs are prioritized. If a segmentation point results in a paragraph's semantic incompleteness (e.g., a complete logical argument is split) or if a paragraph contains multiple relatively independent semantic topics, the segmentation point is adjusted based on the semantic evaluation results, for example, by merging short, semantically cohesive segments or splitting long segments containing multiple topics, until each paragraph unit meets the pre-set semantic integrity standard.
[0055] Key Information Extraction: Identify and extract core information from contracts, including basic contract information (contract title, number, signing date, effective date, and expiration date); party information (company name, legal representative, registered address, contact information, bank account, and account number); and key contract elements (contract amount, payment method, payment time, delivery conditions, performance period, liability for breach of contract, and dispute resolution method). Key information can be extracted using named entity recognition (NER) combined with a deep learning model for contract-specific entity types. For key information within tables, table structure recognition algorithms can be used to locate and extract the information.
[0056] Obtaining the contract terms of the first document and extracting semantic vectors of the contract terms;
[0057] Construct a review knowledge graph and a review checklist; in this embodiment, the review knowledge graph may include a multi-level knowledge graph of different risk category characteristics, different professional field knowledge, regulations, industry norms, historical cases and internal corporate regulations; in this embodiment, the review checklist is a contract review point pre-defined according to business needs, including internal corporate regulations, external regulations and regulatory requirements, and the contract is reviewed in focus according to the review checklist. The custom review rules in the review checklist allow users to customize the review rules according to specific needs and conduct contract reviews according to user-defined rules.
[0058] Obtaining a first risk identification result based on the semantic vector and the review checklist; obtaining a contract category of the first document based on the semantic vector;
[0059] Perform semantic matching between the semantic vector and the review knowledge graph to obtain potential risks; for example, based on the matching results, perform reasoning through relationships in the graph (such as "violation," "non-compliance," "may result in," etc.) to identify potential risks;
[0060] Based on the potential risk and the review knowledge graph, a first risk category is obtained; for example, a feature of the potential risk is mapped and matched with a risk category feature library predefined in the review knowledge graph;
[0061] A preset risk judgment rule is obtained based on the first risk category, and a second risk category is obtained based on the preset risk judgment rule. In this embodiment, the second risk category is a subcategory of the first risk category. If the first risk category is commercial risk, the second risk category may include: imbalance between income and expenditure, price anomalies, payment anomalies and unreasonable economic burden, etc. All identified risk categories can be output (both the identified first risk category and the second risk category can be output), or a corresponding classifier can be set for each category. Combined with the confidence of each classifier, a weighted voting mechanism is used to determine the final risk category, and the risk classification result and confidence score are output.
[0062] Obtaining a second risk identification result based on the contract category and the second risk category;
[0063] A clause rule base is constructed, and compliance results are obtained based on the semantic vector and the clause rule base. In this embodiment, the clause rule base includes correct, incorrect, high-risk, medium-risk, and low-risk clauses.
[0064] Building different standardized contract template libraries according to different contract types, and obtaining a complete result based on the semantic vector and the standardized contract template library, wherein the standardized contract template library includes a list of essential contract terms and the corresponding clause expressions of the list of essential contract terms;
[0065] A total risk identification result is obtained based on the first risk identification result, the second risk identification result, the compliance result, and the complete result.
[0066] The specific steps of obtaining a compliance result based on the semantic vector and the clause rule base include:
[0067] Based on the semantic vector, several semantically identical clauses are obtained, and the key elements of each semantically identical clause are extracted to determine whether the key elements are consistent. If they are inconsistent, the conflict result is obtained based on the key elements and the conflict table; for example, a large language model is used to identify semantically related clauses in the contract, such as payment terms and breach of contract liability, delivery period and acceptance criteria, etc.; key elements such as time, amount, subject and behavior are extracted from the related clauses; the key elements in the related clauses are compared to determine whether there are contradictions or inconsistencies, such as inconsistent descriptions of the same liability in different clauses, inconsistent amount calculation methods, etc.; and rating is performed according to the severity of the regulatory consequences that may be caused by the conflict, thereby obtaining the conflict result.
[0068] For example, conflicts are categorized based on identified key elements (e.g., matching keywords), potentially into data description conflicts (e.g., inconsistent dates, quantities, or amounts), conflicts in the definition of rights and obligations (e.g., unclear or conflicting definitions of responsible parties), and conflicts in preconditions or subsequent conditions for performance. Based on the review knowledge graph (which integrates regulations, judicial precedents, industry practices, and expert experience), the regulatory consequences that such conflicts may typically trigger in similar contractual contexts are analyzed, such as partial or complete invalidity of the contract, rescission of the contract, or the incurrence of liability for breach of contract.
[0069] Similar contract backgrounds can be obtained by searching for contract cases of the same type as similar backgrounds based on the contract type, and then performing similarity searches on contract cases of the same type based on industry, transaction size, subject, region, time, and core clause structure (such as rights and obligations structure, payment method), thereby obtaining similar contracts.
[0070] Combine the risk assessment rules preset in the review knowledge graph with historical case data to establish a rating model for the severity of regulatory consequences. Risk assessment rules may include:
[0071] Economic impact assessment rules: If the amount involved in the conflict accounts for more than 30% of the total contract amount, the rating is severe; if it accounts for 10%-30%, the rating is general; if it accounts for less than 10%, the rating is minor; Contract performance obstacle assessment rules: If the conflict involves a major inconsistency in the performance time (such as the difference in delivery time exceeds 50% of the original period), and this time is critical to the realization of the purpose of the contract, the rating is severe; Regulatory compliance risk assessment rules: If the clause wording caused by the conflict violates the provisions of mandatory regulations, it will be rated as severe regardless of the amount; If the conflict involves a contradiction in the dispute resolution clause (such as agreeing on two resolution methods at the same time, but the two methods are mutually exclusive), the rating is general; Industry specificity assessment rules: If the conflict in the construction project contract involves a major inconsistency in the performance time (such as the difference in delivery time exceeds 50% of the original period), and the time is critical to the realization of the purpose of the contract, the rating is severe; Regulatory compliance risk assessment rules: If the conflict causes the clause wording to violate the mandatory regulations, the rating is severe regardless of the amount; If the conflict involves a contradiction in the dispute resolution clause (such as agreeing on two resolution methods at the same time, but the two methods are mutually exclusive), the rating is general; Industry specificity assessment rules: If the conflict in the construction project contract involves a major inconsistency in the performance time (such as the difference in delivery time exceeds 50% of the original period), and the time is critical to the realization of the purpose of the contract, the rating is severe; If the conflict involves inconsistencies in interest rates and fee calculation methods in financial services contracts, the rating will be severe; if the conflict involves inconsistencies in interest rates and fee calculation methods, the rating will be general to severe based on the amount; time sensitivity assessment rule: if the matters involved in the conflict are irreversible (such as a production process that has already started or a service that has already been initiated), the rating will be increased by one level; remediability assessment rule: if the conflict can be resolved within 7 working days through simple text modifications or supplementary agreements, the rating will be decreased by one level; if the resolution of the conflict requires complex commercial re-negotiation or involves the consent of a third party, the rating will be increased by one level; cumulative effect assessment rule: if there are three or more general-level conflicts in the same contract and these conflicts are interrelated, the overall rating will be severe.
[0072] The above risk assessment rules can be applied individually or in combination. Based on the specific conflict situation, the appropriate combination of rules can be selected, and the final risk level is determined through a weighted average or highest severity approach. Furthermore, these rules can be dynamically adjusted and updated based on feedback from actual applications and new regulatory changes.
[0073] The model uses one or more of the following factors to assign a quantitative score or qualitative rating (e.g., “minor,” “moderate,” “severe,” etc.) to the severity of a conflict:
[0074] Potential economic impact: Assess whether the conflict will directly lead to economic losses for one or more parties, and the estimated range of potential losses; Obstacles to the performance of core obligations: Determine whether the conflict will substantially affect the performance of core obligations of the contract, such as key terms such as payment, delivery, and provision of services; Regulatory risk level: Analyze whether the conflict may violate mandatory regulations or is likely to trigger complex regulatory disputes; Degree of obstruction to the realization of the contract's intended fundamental purpose: Assess the extent to which the conflict hinders the realization of the contract's intended fundamental purpose; Remediability and remediation costs of the conflict: Consider whether the conflict can be remedied simply and effectively through negotiation, supplementary agreement, etc., or whether complex procedures are required to resolve it, as well as the associated costs.
[0075] Initial severity ratings are dynamically adjusted based on the specific contract type (e.g., sales contract, lease contract, service contract, etc.), the transaction context, the industry characteristics, and the specific demands and risk tolerance of both parties to the contract, ensuring the pertinence and accuracy of the assessment results. For example, in a critical project with a large contract value, a minor inconsistency regarding a delivery date may be assessed as a higher severity level due to its potential significant cascading impact.
[0076] Matching the contract terms with the terms rule library to obtain a first compliance result; matching the contract terms with the preset industry rule library to obtain a second compliance result; matching the contract terms with the custom rule library to obtain a third compliance result; matching the contract terms with the historical dispute rule library to obtain a fourth compliance result;
[0077] Ensure that contracts comply with regulations and corporate rules through multi-dimensional compliance checks: (1) Regulatory compliance: Match the contract terms with the latest regulatory terms and conditions to check whether there are any illegal or invalid terms; (2) Industry norms compliance: Build an industry rules library based on the characteristics of different industries to check whether the contract complies with industry standards and regulatory requirements; (3) Internal corporate compliance: Check whether the contract complies with internal management requirements based on the company's customized compliance rules library (such as approval authority, risk control standards, commercial terms bottom line, etc.); (4) Historical case comparison: Compare the contract terms with historical disputes to identify potential risk points.
[0078] The compliance result is obtained based on the first compliance result, the second compliance result, the third compliance result, and the fourth compliance result.
[0079] The specific steps of obtaining a complete result based on the semantic vector and the standardized contract template library include:
[0080] Based on the contract category, the contract terms are matched with the standardized contract template library to obtain a first missing result; and structured analysis is used to detect whether the contract contains necessary terms in the template library.
[0081] Similarity calculation is performed on the contract terms and the standardized contract template library to obtain semantic similarity, and a second missing result is obtained based on the semantic similarity; and whether the terms exist in different expressions is identified through semantic similarity calculation.
[0082] Obtaining missing terms based on the first missing result and the second missing result;
[0083] Obtaining a historical contract location of the missing clause, and obtaining an optimal revision location of the missing clause based on the historical contract location;
[0084] The complete result is obtained based on the missing clause and the optimal revision location. In this embodiment, contextual adaptation can also be performed based on the subject information and core parameters of the current contract to generate supplementary clauses that conform to the overall style and logic of the contract. In this embodiment, the subject information may include the exact names, abbreviations, and roles (e.g., Party A, Party B, guarantor, etc.) of the two or more parties to the contract; core parameters may include the specific description, quantity, specifications of the subject matter of the contract, the amount involved, payment method, performance period, location, as well as specific data directly related to the missing clause (e.g., confidentiality period, liquidated damages ratio, etc.).
[0085] By capturing the logical and reference relationships between the missing clauses and other existing clauses in the contract, for example, the supplementary payment clauses need to be coordinated with the agreed delivery clauses at the time node.
[0086] Analyzing contract styles:
[0087] Lexical features: Natural language processing technology is used to analyze the vocabulary selection of the clauses used in the contract, such as obtaining common regulatory terms, industry-specific terms, and the tendency of formal or informal vocabulary; Syntactic features: Analyze sentence structure, such as average sentence length, common sentence patterns (for example, obtaining mutual consent, should, must not, etc.), and the frequency of use of active / passive voice; Format specifications: Identify format features such as the contract's layout, numbering system (such as the hierarchical structure of clause numbers and item numbers), and title usage; Learning and representation: Quantify or encode the style features obtained from the above analysis to form descriptive parameters or models of the current contract style.
[0088] Analyze the contract logic:
[0089] Identification of relationships between clauses: Based on semantic understanding and review of knowledge graphs, identify dependencies (such as prerequisites, subsequent obligations), definition relationships, constraint relationships, and exception relationships between clauses; Understanding of text structure: Analyze the overall organizational structure of the contract, such as the distribution pattern of rights and obligations, and the logical main line of risk allocation; Maintain consistency: Ensure that newly generated clauses do not logically conflict with existing clauses and conform to the overall regulatory logic and business logic of the contract.
[0090] Generate stylistically and logically correct supplementary clauses:
[0091] Combining templates and generation: For common missing clauses (such as standard confidentiality clauses and notice clauses), the system can match preliminary templates from a library of standardized contract templates. Context adaptation and style injection: Extracted contextual information (subject matter, core parameters) is populated into the selected template or generation framework. The system uses captured contract style features (lexical and syntactic features) to adjust and polish the initially generated clause text, for example, by replacing synonyms and transforming sentence structures to align it with the original text. Logical verification and integration: This ensures that the generated supplementary clauses are accurate in content, logically align with the context, and conform to the overall logical structure of the contract. For example, a newly generated breach of contract clause must correctly reference the relevant obligation clauses. Iterative optimization: Generated supplementary clauses can undergo one or more rounds of evaluation and revision to ensure they meet requirements for regulatory rigor, commercial plausibility, textual fluency, and consistency with the overall contract style. For example, language models can be used to score the readability and naturalness of generated clauses and make adjustments accordingly.
[0092] If the first risk category is commercial risk, the preset risk judgment rules include:
[0093] Obtaining business elements based on the semantic vector, the business elements including price terms, delivery terms, and restrictive terms, the restrictive terms including exclusivity terms, non-competition terms, confidentiality terms, and unilateral termination terms;
[0094] Based on the commercial elements, the benefits and costs of different contract parties are obtained, and whether the benefits and costs are balanced is judged based on a first preset range; for example, a rights-consideration evaluation model is constructed to quantitatively analyze whether the rights and interests obtained by each party are balanced with the costs borne. The specific implementation may include decomposing the terms into multiple rights and obligations units, assigning different value scores to each unit, and calculating the total value difference rate of each party. When the difference rate exceeds a threshold, it is determined to be significantly unfair.
[0095] The contract price is obtained based on the commercial factors, and whether the contract price is abnormal is determined based on the preset average price level and the preset floating range; for example, an industry price index database is constructed, which includes the average price levels and floating ranges of different industries and regions. When the contract price deviates from the industry average and reaches a certain threshold, it is determined to be a price abnormality.
[0096] Based on the commercial elements, the payment node, delivery node, down payment ratio and payment period are obtained, and it is determined whether the down payment ratio exceeds the preset ratio. Based on the preset time range, it is determined whether the payment node and the delivery node match. Based on the preset period, it is determined whether the payment period is abnormal; such as the down payment ratio is too high, the payment node and the delivery node are seriously mismatched, and the payment period is too short.
[0097] The form of liability for breach of contract is obtained based on the commercial elements, and whether the form of liability for breach of contract is abnormal is judged based on the preset breach of contract amount; if the advance payment time is too long; the liquidated damages ratio is too high; the deposit ratio is too high; the additional fee clauses (such as management fees, service fees) account for more than 15% of the main transaction amount, etc., it is judged to be an unreasonable economic burden.
[0098] Based on the preset restriction data, it is determined whether the restriction clause is abnormal. In this embodiment, the preset restriction data may include keywords related to restrictions, such as prohibition of cooperation with third parties, non-competition restriction, confidentiality period, and unilateral termination.
[0099] Based on the restrictive clauses, potential restriction data is obtained. Based on a preset restriction table and the potential restriction data, a restriction score is obtained. A determination is made as to whether an abnormality exists based on the restriction score and a preset threshold. In this embodiment, the potential restriction data may include the scope of the restriction, the duration of the restriction, and liability for breach of contract. For example, the analysis of the potential restrictions a clause may impose on the company's future business development may include restrictions on technology use, market access, and customer resources. The preset restriction table includes scores corresponding to different restrictions. A combined score for all restrictions is calculated. If the score exceeds a preset threshold, a high-risk business restriction is identified.
[0100] Example 2
[0101] On the basis of the first embodiment, in this embodiment, the method further includes:
[0102] The contract terms are analyzed based on the position recognition model to obtain unbalanced terms; for example, a deep learning model is trained with a large amount of labeled data to perform position analysis on the clause text and identify the tendency of the clauses towards the parties to the contract.
[0103] Obtaining liability elements based on the imbalance clause, constructing a liability-obligation balance matrix based on the liability elements, and obtaining the distribution ratio of responsibilities and powers of different contract parties based on the liability-obligation balance matrix, wherein the liability elements include the responsible party, responsibilities and obligations, exercise of power, and power restrictions;
[0104] Constructing a library of unfavorable clause indicative words, matching the semantic vector with the library to obtain a target vocabulary, obtaining a target text based on the target vocabulary and a preset text range, such as the preceding and following text of the target vocabulary (e.g., the five preceding and following sentences), obtaining a conditional relationship based on the target text, and obtaining the semantic role of the target text based on the conditional relationship, such as identifying conditional clauses and hypothetical statements, such as the conditional clause: If...then..., distinguishing different semantic roles in the conditional clause to reduce misjudgment of hypothetical content;
[0105] For example, in a conditional clause: If Party B fails to pay the amount before the agreed date, Party A has the right to terminate the contract and require Party B to pay liquidated damages, the following semantic roles can be identified:
[0106] Condition: Party B fails to pay the amount by the agreed date - a condition on the occurrence or non-occurrence of an event;
[0107] Actor / responsible party (in the condition part): Party B - the subject of the behavior or responsibility; behavior / event (in the condition part): failure to pay the money - the specific behavior or event; time limit (in the condition part): before the agreed date - the time when the conditional behavior occurs; object / subject matter (in the condition part): money - the object of the behavior; result / right (in the main clause part): Party A has the right to terminate the contract and require Party B to pay liquidated damages - the result or right granted after the condition is met or not met; right holder (in the main clause part): Party A - the beneficiary of the result or right; specific actions / measures (in the main clause part): termination of the contract, requirement for payment of liquidated damages - the specific behavior or measures included in the result; related object / subject matter (in the main clause part): contract, liquidated damages - the object of the main clause behavior.
[0108] Through the above semantic roles, we can obtain:
[0109] Hypothetical: The entire if...then... structure expresses a hypothetical scenario and its possible consequences, rather than an established fact; Triggering Relationship: It clearly states which condition (e.g., Party B's failure to pay) will trigger which consequence (e.g., Party A's termination of the contract); Ownership of Rights and Responsibilities: It clearly identifies the rights and obligations of each party under the hypothetical conditions.
[0110] In addition to the above roles, other semantic roles may be identified depending on the specific sentence structure and semantic complexity, such as:
[0111] Place: the place where the incident occurs; Method: the way the behavior occurs; Tool: the tool used to complete the behavior; Cause: the reason that led to the incident; Purpose: the purpose to be achieved by the behavior; Beneficiary: the beneficiary of the behavior (when different from the right holder); Loss: the person who suffers damage from the behavior.
[0112] Accurately identifying semantic roles helps to gain a deeper understanding of the true meaning of the clauses, especially when dealing with complex conditions, assumptions, exceptions, etc., avoiding misjudging hypothetical content as deterministic descriptions, thereby improving the accuracy of risk identification.
[0113] In this embodiment, the unfavorable clause indicator word library may include unfavorable clause indicator words such as unilateral decision, no notification required, and no liability.
[0114] The target vocabulary is matched with pre-set language elements to obtain semantic modifiers. The pre-set language elements include qualifiers, modifiers, and exceptions. For example, qualifiers, modifiers, and exceptions that may change the basic meaning of the keyword are detected. For example, the distinction between "shall not be unilaterally determined" and "shall be unilaterally determined" is not allowed except in special circumstances.
[0115] In this embodiment, exceptions may specifically refer to statements in contract terms that exclude specific circumstances, conditions, or objects from a provision, right, obligation, or responsibility. These statements typically serve to limit or modify the scope of the primary statement. Specific content may include, but is not limited to, the following types:
[0116] Statements that explicitly exclude specific objects or situations:
[0117] For example: Exclusions, Exclusions, This does not apply to. Specific phrases include: Except for, but not limited to, Does not apply to the following: This warranty does not cover...
[0118] Statements that set conditional exceptions: for example, clauses introduced by "unless...", "if not...", "only if...", etc., indicate that the provisions of the main clause are applicable or not applicable when specific conditions are met or not met.
[0119] Statement of force majeure or exemption from liability:
[0120] For example: a dedicated force majeure clause, or a disclaimer immediately following the description of specific obligations.
[0121] References to other clauses as exceptions:
[0122] For example, some clauses will specify that their provisions are subject to limitations or adjustments made by other specific terms in the contract.
[0123] A statement that qualifies or narrows a previous general statement:
[0124] For example: first give a general rule, and then introduce exceptions through words such as but, however, and notwithstanding the foregoing.
[0125] By identifying the specific language patterns and structures of the exceptions mentioned above, the system can more accurately understand the scope of application and conditions of the clauses, thereby accurately assessing potential risks and the true rights and obligations of the parties to the contract.
[0126] Build a professional domain knowledge base and an industry standard clause library. Based on the professional domain knowledge base, the industry standard clause library, the allocation ratio, the semantic roles, and the semantic modifiers, simulate the perspectives of different contract parties, analyze the imbalance clauses, obtain evaluation results, and obtain biased clauses based on the evaluation results. Interpret the clause content from the perspectives of different contract parties, evaluate the different impacts of the same statement on different parties, and achieve a more comprehensive and neutral risk assessment. Specific examples can be:
[0127] 1. Identification of subject roles and construction of interest portraits:
[0128] Clarify the identities of the parties: Identify the parties defined in the contract, such as Party A, Party B, lessor, lessee, seller, buyer, service provider, and service recipient, and record their specific references in the contract (such as the full company name).
[0129] Pre-defined interests and concerns: By combining professional domain knowledge bases and industry standard clause libraries, we pre-defined the core interests and potential risks that different types of contract parties typically focus on in their respective contracts. For example: 1. For buyers / service recipients: typically, they focus on the quality, quantity, timely delivery, compliance with specifications, after-sales service, reasonableness of price, and adequacy of the other party's liability for breach of contract. 2. For sellers / service providers: typically, they focus on timely payment, reasonable limitations on their own liability, intellectual property protection, remedies for buyer breach of contract, and the effectiveness of exemption clauses. 3. For lessors: typically, they focus on timely and full collection of rent, the integrity of the leased property, and the lessee's compliant use. 4. For lessees: typically, they focus on the availability of the leased property, the stability of the lease term, and the reasonableness of the fees.
[0130] The aforementioned presupposed interests and concerns form the basis for simulating the perspective of a specific subject.
[0131] 2. Clause impact assessment (based on the subject’s perspective): When analyzing a specific clause (especially one that has been identified as a potential imbalance clause or requires special assessment), the perspective of each contracting party will be taken into account.
[0132] Rights assessment: From the perspective of the subject, analyze the rights granted by the clauses to determine whether the rights are sufficient, clear and enforceable, and whether there are any conflicts or restrictions with the rights granted by other clauses.
[0133] Obligation assessment: From the perspective of the subject, analyze the obligations stipulated in the terms and conditions to determine whether the obligations are clear and reasonable, the cost and difficulty of fulfilling them, and whether there are any unreasonable aggravated responsibilities.
[0134] Risk assessment: From the perspective of the entity, analyze the potential risks that the clauses may expose it to; for example, the losses that may be suffered due to the other party's breach of contract, the entity's own performance obstacles, the uncertainty of the scope of liability, etc.; whether the clauses provide sufficient risk avoidance or relief mechanisms.
[0135] Opportunity / Benefit Assessment: From the perspective of the entity, analyze the business opportunities or expected benefits that the clauses can bring to it; whether the provisions of the clauses will help it achieve the purpose of the contract.
[0136] 3. Use existing analysis results for simulation:
[0137] Reference for allocation ratios: When simulating perspectives, the previously calculated ratios of responsibility and power are considered. If a party's ratio is significantly lower, the system will focus on examining whether the terms are overly restrictive or insufficiently protective from their perspective.
[0138] Interpreting semantic roles and modifiers: From the perspective of a specific party, interpret the semantic roles (e.g., who is the agent and who is the recipient) and semantic modifiers (e.g., "only," "must," and "right to unilaterally") in the clause to determine the actual impact of these language elements on that party. For example, the statement that Party A has the right to unilaterally change the content of the service may be seen as flexibility from Party A's perspective, but may represent uncertainty and potential loss of interests from Party B's perspective.
[0139] 4. Comparative analysis and difference identification: The system compares the interpretation results of the same clause from different perspectives.
[0140] Identify which terms benefit one party but not the other, or which pose far greater risk to one party than the other.
[0141] Quantify or qualitatively describe the extent of this difference in impact. For example, when evaluating a payment clause, the payee's perspective focuses on the timeliness and certainty of payment; the payer's perspective focuses on whether the payment terms are reasonable and whether there is a sufficient grace period. If the clause stipulates that the payer must pay the full amount immediately upon receipt of the invoice, or pay a penalty of 5% of the total contract amount per day, then this is clearly a very demanding clause for the payer.
[0142] 5. Generate a bias assessment: Based on the above analysis, determine the overall bias of the clause, that is, whether the clause is designed to protect Party A or Party B, or is relatively balanced.
[0143] This multi-perspective simulation-based assessment helps uncover clauses that appear neutral but may actually be unfair to one party, providing deeper and more comprehensive risk insights. This simulation mechanism allows the system to move beyond a literal analysis of the text and instead empathize with the actual stakes of each contracting party, allowing for more accurate identification of potential unfairness and risks.
[0144] In this embodiment, the method may further include: using a word sense disambiguation algorithm, combined with knowledge bases in different professional fields, to accurately identify the specific meanings of professional terms in regulations and business contexts, achieve polysemy resolution, and reduce misjudgments caused by ordinary language understanding.
[0145] In this embodiment, the method may further include:
[0146] Antonym and negation review: This feature examines sentences based on the presence of antonyms and negations. For example, if the original sentence states that a rule must not be violated, it is considered logically correct. However, if the sentence is modified to allow a violation of the rule, it will be immediately considered abnormal.
[0147] Review of consistency of table amounts: Calculate the subtotals and totals of the tables in the contract and determine whether they are consistent based on the numerical values.
[0148] Check the consistency of uppercase and lowercase amounts: Check whether the uppercase and lowercase amounts are correct and consistent.
[0149] Logical review: Determine whether the contract terms comply with regulations.
[0150] Example 3
[0151] Based on the above embodiment, in this embodiment, the method further includes:
[0152] The semantic vector is matched with preset disagreement words to obtain clauses for resolving disagreements. In this embodiment, the preset disagreement words may include words such as disagreement and dispute.
[0153] Match the dispute resolution clause with preset resolution words (such as negotiation and mediation) to obtain a completeness result; match the dispute resolution clause with preset fuzzy words (such as fuzzy expressions that are difficult to enforce, such as friendly negotiation resolution) to obtain an enforceability result; match the dispute resolution clause with preset special words (such as statutory dispute resolution methods for specific types of contracts) to obtain a compliance result; and obtain a rationality result based on the completeness result, the enforceability result, and the compliance result.
[0154] By matching and comparing contract terms with the statutory or conventional dispute resolution methods for these specific types of contracts, compliance risks of dispute resolution clauses can be more effectively identified.
[0155] Extract the geographical location entity (such as a city, the name of a dispute resolution agency, etc.) and the dispute resolution method (such as negotiation and mediation) of the dispute resolution clause, obtain the association between the geographical location entity and the dispute resolution method, such as whether the geographical location entity is a negotiation location or a dispute resolution location, and determine the jurisdiction based on the association and the geographical location entity;
[0156] Obtaining address information of the contract subject of each party based on the semantic vector, and obtaining a degree of disadvantage of the contract subject of each party based on the address information and the jurisdiction;
[0157] Obtaining a jurisdiction compliance result of the jurisdiction based on the address information, the geographic location entity, and the jurisdiction; and verifying whether the specified jurisdiction complies with regulations.
[0158] A disagreement resolution result is obtained based on the reasonableness result, the degree of disadvantage result and the jurisdiction compliance result.
[0159] Example 4
[0160] refer to Figure 2 Based on the above embodiment, in this embodiment, the semantic vector of the contract clause is extracted based on the semantic model. In the calling order, the semantic model includes an encoding module, a multi-granularity self-attention module, a hierarchical progressive encoding module, and a semantic integration module. The encoding module is used to encode the contract document according to the hierarchical structure, the multi-granularity self-attention module is used to extract global semantic features, the hierarchical progressive encoding module is used to extract local semantic features, and the semantic integration module is used to fuse features to generate a semantic vector.
[0161] The multi-granularity self-attention module includes a nested self-attention layer, which includes a hierarchical attention mechanism and a relative position-aware attention mechanism.
[0162] The hierarchical progressive coding module includes a plurality of cross-layer connection layers, wherein the cross-layer connection layers include a cross-layer connection mechanism, and the cross-layer connection layers are connected by residual connections based on a preset number of interval layers;
[0163] The semantic integration module includes a gating mechanism.
[0164] Among them, the semantic model also includes a multi-level semantic feature extraction module, which is connected to the hierarchical progressive encoding module and the semantic integration module. The multi-level semantic feature extraction module includes a surface semantic extraction layer, a deep semantic fusion layer and a latent semantic reasoning layer. The surface semantic extraction layer is used to extract text expressions and keyword features, the deep semantic fusion layer is used to extract contextual semantics, and the latent semantic reasoning layer is used to construct a sentence relationship graph.
[0165] In this embodiment, the semantic model utilizes an improved Transformer architecture. This architecture, by introducing nested self-attention layers and cross-layer connection mechanisms, enables the model to simultaneously process local and global semantic features. Specifically, the architecture comprises three key modules: a multi-granularity self-attention module, a hierarchical progressive encoding module, and a semantic integration module. This architecture innovatively employs a hierarchical attention mechanism, segmenting long text into multiple semantic chunks and applying attention calculations with different parameters within and between chunks, significantly reducing computational complexity. For example, the computational power within a chunk can be calculated using a standard self-attention mechanism, such as scaled dot-product attention, to precisely capture local contextual information and short-range dependencies between tokens within each segmented semantic chunk. Inter-chunk computational power can be calculated by generating a generalized chunk representation vector for each semantic chunk processed through the intra-chunk attention process. This can be achieved by performing a pooling operation (e.g., average pooling or max pooling) on the final output representation vectors of all tokens within the chunk. The self-attention mechanism is then applied again between the generalized representation vectors of these different semantic chunks to calculate the mutual importance between the different semantic chunks and capture their global dependencies.
[0166] This nested approach, with different parameter configurations for intra- and inter-block attention calculations, allows us to first focus on understanding local details (within a block), then integrate these local understandings to grasp the overall structure and semantic flow of the text at a higher level (between blocks). This layered approach not only effectively addresses the quadratic computational complexity associated with long sequences in long text processing, but also enhances the model's ability to capture and integrate multi-granular semantic information.
[0167] This embodiment significantly improves the ability to process long texts by optimizing the attention mechanism and position encoding method.
[0168] In terms of attention mechanism, relative position-aware attention (RPAA) is introduced. Unlike the absolute position encoding used by traditional Transformer, RPAA considers the relative distance between tokens. Its core formula is:
[0169] ;
[0170] ;
[0171] in, represents the correlation function, represents the Softmax function, represents a value vector, represents the query vector, represents the key vector, represents transpose, represents the dimension of the key vector (or query vector), represents the relative position matrix, represents the mapping function, converting the relative distance into attention bias, Relative position matrix Rank Elements of the column, and Both represent position indexes, which are integers greater than or equal to 1.
[0172] By introducing the relative position matrix, the relative position-aware attention (RPAA) mechanism enables the model to not only rely on the similarity of word content when calculating attention, but also directly perceive and utilize the relative spatial relationship between words, which is crucial for understanding long-range references and contextual order between clauses in contract texts.
[0173] This mechanism can better capture long-distance dependencies and is particularly suitable for remote reference resolution between clauses in contract texts.
[0174] In terms of positional encoding, a piecewise recursive positional encoding is used to encode documents according to their hierarchical structure (section-clause-paragraph), preserving the relative position information between different levels. This effectively solves the vanishing gradient problem of traditional positional encoding when processing very long texts. Compared to the standard Transformer or its variants (such as BERT and RoBERTa) in existing technologies, these improvements can process longer sequences (up to 16K tokens) and improve the accuracy of long-distance dependency understanding.
[0175] This embodiment introduces a multi-layer semantic feature extraction mechanism consisting of three layers: a surface semantic extraction layer, a deep semantic fusion layer, and a latent semantic inference layer. The surface semantic extraction layer uses the bag-of-words model and word vector technology to capture direct textual representations; the deep semantic fusion layer uses a bidirectional LSTM combined with a self-attention mechanism to capture contextual semantics; and the latent semantic inference layer uses a graph neural network to construct a relationship graph between sentences and infer implicit regulatory concepts and relationships through a message passing mechanism. Features from each layer are selectively fused through a gating mechanism to form a comprehensive semantic representation.
[0176] In this embodiment, the semantic model also includes a contextual association analysis algorithm, which can accurately understand the logical relationship and dependency relationship between clauses. Its core formula is:
[0177] ;
[0178] in, Indicates the use of measurement and The value of the contextual association strength or a certain specific logical relationship between and Respectively represent and contract terms, and Respectively represent and The semantic vector representation of each contract clause, represents element-wise multiplication, represents the activation function, and Both represent learnable parameters.
[0179] The algorithm innovatively integrates knowledge constraints and logical rules specific to the regulatory domain, and adds explicit identification of logical relationship types (such as definitions, constraints, premises, and exceptions), enabling the system to accurately identify complex regulatory logical relationships between clauses.
[0180] In this embodiment, the semantic model can also introduce an adaptive blood deficiency mechanism, and by designing a feedback optimization loop, the model can continuously learn and improve from practical applications.
[0181] Example 5
[0182] On the basis of the above-mentioned embodiment, in this embodiment, the method can also introduce a dynamic rule update mechanism, and realize the dynamic update of rules through a closed-loop process of rule effectiveness evaluation-rule optimization-rule verification-rule deployment. The specific process includes: rule effectiveness evaluation: continuously monitor the triggering rate, accuracy and user feedback of each rule, and comprehensively evaluate the effectiveness of the rules; rule optimization: for inefficient rules, analyze the reasons for failure, and automatically generate optimization suggestions based on the latest cases and user feedback, and at the same time discover new rule patterns from historical data through machine learning algorithms; rule verification: new rules are backtested with historical data in a sandbox environment to verify their accuracy and generalization ability; rule deployment: through A / B testing, the verified rules are gradually deployed to the production environment, and the rule performance is continuously monitored. This dynamic update mechanism enables risk identification rules to be continuously optimized as new cases accumulate, and the risk identification ability of the system is also improved.
[0183] Example 6
[0184] Based on the above embodiment, in this embodiment, the method further includes: generating a modification suggestion;
[0185] Based on the identified risks and non-compliances, modification suggestions are generated through the following process: (1) Problem location: Accurately locate the problematic statements or missing content in the clauses; (2) Regulatory basis association: Extract regulatory basis or best practices related to the problem from the knowledge graph; (3) Alternative solution generation: Based on the large language model and historical case library, generate alternative clause texts that meet regulatory requirements and corporate interests; (4) Impact analysis: Evaluate the potential impact of the modification suggestions on other parts of the contract to ensure overall consistency after the modification; (5) Multi-solution comparison: For complex problems, generate multiple alternative modification solutions and conduct pros and cons analysis to support decision makers in selecting the optimal solution.
[0186] The modification suggestions are presented in a structured format of problem description-regulation basis-modification suggestions-expected effects, which is convenient for review and adoption.
[0187] Example 7
[0188] refer to Figure 3-Figure 5 Based on the above embodiments, in this embodiment, the contract documents are analyzed based on an intelligent review model. The intelligent review model includes a clause knowledge base, a review knowledge graph, a review checklist, a training annotation platform, a user interaction platform and a review model. The training annotation platform is used to annotate correct clause statements and incorrect clause statements for training the review model; the user interaction platform is used for users to operate contract terms, such as uploading, online editing and viewing, etc., and supports multi-person synchronous operations; the review model is used to identify risks in contract documents.
[0189] The review model consists of the RoBERTa module, the RAG-knowledge graph module, the MOE module, and an output module. The RoBERTa module extracts semantic vectors from contract documents, the RAG-knowledge graph module derives the first risk identification result based on the semantic vectors, the MOE module derives the second risk identification result based on the semantic vectors, and the output module outputs the overall risk identification result and revision suggestions.
[0190] The RoBERTa module converts the input contract text into a deep language semantic vector representation, providing high-quality semantic embedding for subsequent modules. The functions of the RoBERTa module are implemented by the semantic model.
[0191] The RAG-Knowledge Graph module retrieves the embedding vectors output by RoBerta and efficiently searches and matches them against a pre-built review knowledge graph, mapping them to high, medium, and low risk during analysis. This knowledge graph is supported by a rule base that includes multi-domain laws, industry standards, and expert annotations. Through vector similarity calculation and semantic diffusion based on the graph structure, the RAG-Knowledge Graph module retrieves review rule information closely associated with the text embedding, enhancing understanding of the regulatory applicability, industry standards, and risk points of specific contract clauses, forming an amplified embedded representation with external knowledge support.
[0192] The MOE module includes a type recognition model and a true / false discrimination model, and its expert layers are shown in Table 1 and Table 2 respectively:
[0193] Table 1 Example of the expert layer of the type recognition model
[0194]
[0195] Table 2 Example of the expert layer of the true and false judgment model
[0196]
[0197] The embedding vectors output by the RoBERTa module then enter the two models of the MOE module. The data first enters the Gating Network. Based on Top-k sparse gating, the Gating Network outputs a sparse weight vector (each value corresponds to the activation weight of an expert). These weights determine which experts participate in the calculation and the contribution of each expert. After the input data (embedding vectors) are combined, they are fed into each expert model (each expert model is a Transformer-based neural network). The resulting outputs are added to the residual connection and enter the expert combination layer. This layer integrates the outputs of all expert layers in each model based on the weights of the gating network to obtain the final output of each model.
[0198] The structure of the true or false judgment model is the same as that of the type recognition model. The main difference lies in the different training fields of the expert layer. The expert layer of the true or false judgment model mainly focuses on the true or false judgment of contract terms, while the type recognition model mainly focuses on the type judgment of contract terms.
[0199] After integrating the contract type identification results and clause correctness judgment results, the output module conducts a secondary search on the input review knowledge graph obtained through analysis to accurately locate the regulations, legal reminders, and compliance recommendations corresponding to the contract type and problematic clauses. This module reorganizes and condenses the professional information scattered within the knowledge graph and presents it in a structured and readable format, ensuring that the final review conclusions and recommendations have clear logical guidance, authentic legal basis, and actionable practical value. Through this module, users can directly obtain contract compliance review conclusions and improvement plans that meet actual business needs, significantly improving the efficiency and quality of review decisions.
[0200] Example 8
[0201] refer to Figure 1 Based on the above embodiment, this embodiment provides a contract risk intelligent identification system, which includes:
[0202] Preprocessing module: used for obtaining a file to be identified, and preprocessing the file to be identified to obtain a first file;
[0203] Extraction module: used for obtaining the contract terms of the first document and extracting the semantic vectors of the contract terms;
[0204] Construction module: used to build review knowledge graph and review checklist;
[0205] A first identification module: configured to obtain a first risk identification result based on the semantic vector and the review list;
[0206] A second recognition module: configured to obtain a contract category of the first document based on the semantic vector;
[0207] and for performing semantic matching between the semantic vector and the review knowledge graph to obtain potential risks; obtaining a first risk category based on the potential risks and the review knowledge graph, obtaining a preset risk judgment rule based on the first risk category, and obtaining a second risk category based on the preset risk judgment rule; and obtaining a second risk identification result based on the contract category and the second risk category;
[0208] A third recognition module is used to construct a clause rule base and obtain compliance results based on the semantic vector and the clause rule base;
[0209] and constructing different standardized contract template libraries according to different contract types, obtaining a complete result based on the semantic vector and the standardized contract template library, wherein the standardized contract template library includes a list of essential contract terms and the corresponding clause expressions of the list of essential contract terms;
[0210] Result module: used to obtain a total risk identification result based on the first risk identification result, the second risk identification result, the compliance result and the complete result.
[0211] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0212] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. A method for intelligent identification of contract risks, characterized in that: The method comprises: Acquire a file to be identified, pre-process the file to be identified to obtain a first file, obtain contract terms of the first file, and extract semantic vectors of the contract terms; Build review knowledge graph and review checklist; Obtaining a first risk identification result based on the semantic vector and the review checklist; obtaining a contract category of the first document based on the semantic vector; Perform semantic matching on the semantic vector and the review knowledge graph to obtain potential risks; obtain a first risk category based on the potential risks and the review knowledge graph, obtain a preset risk judgment rule based on the first risk category, and obtain a second risk category based on the preset risk judgment rule; obtain a second risk identification result based on the contract category and the second risk category; Building a clause rule base, and obtaining compliance results based on the semantic vector and the clause rule base; Building different standardized contract template libraries according to different contract types, and obtaining a complete result based on the semantic vector and the standardized contract template library, wherein the standardized contract template library includes a list of essential contract terms and the corresponding clause expressions of the list of essential contract terms; Obtaining an overall risk identification result based on the first risk identification result, the second risk identification result, the compliance result, and the complete result; The specific steps of obtaining a compliance result based on the semantic vector and the clause rule base include: obtaining a plurality of semantically identical clauses based on the semantic vectors, extracting key elements of each of the semantically identical clauses, determining whether the key elements are consistent, and if not consistent, obtaining conflict results based on the key elements and a conflict table; Matching the contract terms with the terms rule library to obtain a first compliance result; matching the contract terms with the preset industry rule library to obtain a second compliance result; matching the contract terms with the custom rule library to obtain a third compliance result; matching the contract terms with the historical dispute rule library to obtain a fourth compliance result; Obtaining the compliance result based on the first compliance result, the second compliance result, the third compliance result, and the fourth compliance result; The specific steps of obtaining a complete result based on the semantic vector and the standardized contract template library include: Based on the contract category, matching the contract terms with the standardized contract template library to obtain a first missing result; Calculating similarity between the contract clause and the standardized contract template library to obtain semantic similarity, and obtaining a second missing result based on the semantic similarity; Obtaining missing terms based on the first missing result and the second missing result; Obtain the historical contract position of the missing clause, and obtain the optimal revision position of the missing clause based on the historical contract position; and obtain the complete result based on the missing clause and the optimal revision position.
2. A method for intelligent identification of contract risks according to claim 1, characterized in that: If the first risk category is commercial risk, the preset risk judgment rules include: Obtaining business elements based on the semantic vector, the business elements including price terms, delivery terms, and restrictive terms, the restrictive terms including exclusivity terms, non-competition terms, confidentiality terms, and unilateral termination terms; Obtaining benefits and costs of different contract entities based on the business elements, and determining whether the benefits and costs are balanced based on a first preset range; Obtaining a contract price based on the commercial factors, and determining whether the contract price is abnormal based on a preset average price level and a preset floating range; Obtaining a payment node, a delivery node, a down payment ratio, and a payment period based on the business elements, determining whether the down payment ratio exceeds a preset ratio, determining whether the payment node and the delivery node match based on a preset time range, and determining whether the payment period is abnormal based on a preset period; Obtaining a form of liability for breach of contract based on the business elements, and determining whether the form of liability for breach of contract is abnormal based on a preset amount of breach of contract; Determining whether the restriction clause is abnormal based on preset restriction data; Potential restriction data is obtained based on the restriction clauses, a restriction score is obtained based on a preset restriction table and the potential restriction data, and whether an abnormality is detected is determined based on the restriction score and a preset threshold.
3. A method for intelligent identification of contract risks according to claim 1, characterized in that: The method further comprises: Analyze the contract terms based on the position identification model to obtain imbalance terms; Obtaining liability elements based on the imbalance clause, constructing a liability-obligation balance matrix based on the liability elements, and obtaining the distribution ratio of responsibilities and powers of different contract parties based on the liability-obligation balance matrix, wherein the liability elements include the responsible party, responsibilities and obligations, exercise of power, and power restrictions; Constructing an unfavorable clause indicator lexicon, matching the semantic vector with the unfavorable clause indicator lexicon to obtain a target vocabulary, obtaining a target text based on the target vocabulary and a preset text range, obtaining a conditional relationship based on the target text, and obtaining a semantic role of the target text based on the conditional relationship; Matching the target vocabulary with preset language elements to obtain semantic modifiers, wherein the preset language elements include qualifiers, modifiers, and exceptions; Construct a professional domain knowledge base and an industry standard clause library, simulate the perspectives of different contract parties based on the professional domain knowledge base, the industry standard clause library, the allocation ratio, the semantic roles and the semantic modifiers, analyze the unbalanced clauses, obtain evaluation results, and obtain preferred clauses based on the evaluation results.
4. A method for intelligently identifying contract risks according to claim 1, characterized in that: The method further comprises: Matching the semantic vector with preset divergence words to obtain divergence resolution clauses; Matching the dispute resolution clause with preset resolution words to obtain a completeness result; matching the dispute resolution clause with preset fuzzy words to obtain an enforceability result; matching the dispute resolution clause with preset special words to obtain a compliance result; and obtaining a rationality result based on the completeness result, the enforceability result, and the compliance result; Extracting the geographical location entity and the dispute resolution method of the dispute resolution clause, obtaining an association between the geographical location entity and the dispute resolution method, and obtaining a jurisdiction based on the association and the geographical location entity; Obtaining address information of the contracting parties based on the semantic vector, and obtaining disadvantageous degree results of the contracting parties based on the address information and the jurisdiction; Obtaining a jurisdictional compliance result of the jurisdiction based on the address information, the geographic location entity, and the jurisdiction; A disagreement resolution result is obtained based on the reasonableness result, the degree of disadvantage result and the jurisdiction compliance result.
5. A method for intelligently identifying contract risks according to claim 1, characterized in that: Extracting semantic vectors of the contract terms based on a semantic model. The semantic model includes, in order of calling, an encoding module, a multi-granularity self-attention module, a hierarchical progressive encoding module, and a semantic integration module. The encoding module is used to encode the contract document according to a hierarchical structure. The multi-granularity self-attention module is used to extract global semantic features. The hierarchical progressive encoding module is used to extract local semantic features. The semantic integration module is used to fuse features and generate a semantic vector. The multi-granularity self-attention module includes a nested self-attention layer, which includes a hierarchical attention mechanism and a relative position-aware attention mechanism.
6. A method for intelligent identification of contract risks according to claim 5, characterized in that: The hierarchical progressive coding module includes a plurality of cross-layer connection layers, each of which includes a cross-layer connection mechanism. Based on a preset number of interval layers, the cross-layer connection layers are connected through residual connections; the semantic integration module includes a gating mechanism.
7. A method for intelligently identifying contract risks according to claim 6, characterized in that: The semantic model also includes a multi-level semantic feature extraction module, which is connected to the hierarchical progressive encoding module and the semantic integration module. The multi-level semantic feature extraction module includes a surface semantic extraction layer, a deep semantic fusion layer and a latent semantic reasoning layer. The surface semantic extraction layer is used to extract text expressions and keyword features, the deep semantic fusion layer is used to extract contextual semantics, and the latent semantic reasoning layer is used to construct a sentence relationship graph.
8. A contract risk intelligent identification system, characterized by: The system comprises: Preprocessing module: used for obtaining a file to be identified, and preprocessing the file to be identified to obtain a first file; Extraction module: used for obtaining the contract terms of the first document and extracting the semantic vectors of the contract terms; Construction module: used to build review knowledge graph and review checklist; A first identification module: configured to obtain a first risk identification result based on the semantic vector and the review list; A second recognition module: configured to obtain a contract category of the first document based on the semantic vector; and for performing semantic matching between the semantic vector and the review knowledge graph to obtain potential risks; obtaining a first risk category based on the potential risks and the review knowledge graph, obtaining a preset risk judgment rule based on the first risk category, and obtaining a second risk category based on the preset risk judgment rule; and obtaining a second risk identification result based on the contract category and the second risk category; A third recognition module is used to construct a clause rule base and obtain compliance results based on the semantic vector and the clause rule base; and constructing different standardized contract template libraries according to different contract types, obtaining a complete result based on the semantic vector and the standardized contract template library, wherein the standardized contract template library includes a list of essential contract terms and the corresponding clause expressions of the list of essential contract terms; Result module: used for obtaining a total risk identification result based on the first risk identification result, the second risk identification result, the compliance result and the complete result; The specific steps of obtaining a compliance result based on the semantic vector and the clause rule base include: obtaining a plurality of semantically identical clauses based on the semantic vectors, extracting key elements of each of the semantically identical clauses, determining whether the key elements are consistent, and if not consistent, obtaining conflict results based on the key elements and a conflict table; Matching the contract terms with the terms rule library to obtain a first compliance result; matching the contract terms with the preset industry rule library to obtain a second compliance result; matching the contract terms with the custom rule library to obtain a third compliance result; matching the contract terms with the historical dispute rule library to obtain a fourth compliance result; Obtaining the compliance result based on the first compliance result, the second compliance result, the third compliance result, and the fourth compliance result; The specific steps of obtaining a complete result based on the semantic vector and the standardized contract template library include: Based on the contract category, matching the contract terms with the standardized contract template library to obtain a first missing result; Calculating similarity between the contract clause and the standardized contract template library to obtain semantic similarity, and obtaining a second missing result based on the semantic similarity; Obtaining missing terms based on the first missing result and the second missing result; Obtain the historical contract position of the missing clause, and obtain the optimal revision position of the missing clause based on the historical contract position; and obtain the complete result based on the missing clause and the optimal revision position.
Citation Information
Patent Citations
Contract text risk detection method, device and equipment and storage medium
CN109918635A
Building construction contract risk review method and system based on artificial intelligence
CN114187143A