Contract review result evaluation method and device, computer equipment and readable storage medium
By constructing a manually labeled benchmark set and semantic vector mapping, the problem of expression and classification differences in complex legal clause evaluation by large models was solved, and the accurate evaluation and quantitative assessment of contract review results were achieved.
Patent Information
- Application Number
- CN202512057484.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-02-10
AI Technical Summary
Existing technologies lack the ability to understand the semantics of complex legal clauses, and cannot meet the evaluation requirements of complex outputs from large models.
We construct a human-annotated benchmark set and eliminate the differences in description and classification between the output of large models and human annotations through semantic vector mapping and multi-dimensional evaluation methods, so as to achieve accurate comparison and evaluation.
It improves the accuracy and standardization of contract review results, provides reliable evaluation criteria, and can quantify the applicability of large models in actual contract review.
Smart Images

Figure CN121502432A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of document review technology, and in particular to a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for evaluating contract review results. Background Technology
[0002] With the development of document review technology, rule template matching and keyword retrieval technologies have emerged. These technologies determine whether risk statements contain core keywords by using a predefined risk terminology library, or determine whether the location text is consistent with the overlapping area marked by humans by using string matching. However, such methods are only applicable to contract texts with a high degree of structure and standardized expression. They have extremely weak semantic understanding capabilities for complex legal clauses and cannot meet the evaluation needs of complex outputs from large models. Summary of the Invention
[0003] Therefore, it is necessary to provide a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for evaluating contract review results that can meet the evaluation requirements of complex outputs of large models, in order to address the above-mentioned technical problems.
[0004] Firstly, this application provides a method for evaluating the results of contract review, including:
[0005] Obtain a set of manual annotation benchmarks for annotating contracts. The set of manual annotation benchmarks is obtained by annotating and verifying different types of anonymized contracts based on a rule document that guides annotation. The rule document includes the name of the review item, risk type, risk level, risk definition, and location rules for locating review items in the contract.
[0006] Map the names of manually reviewed items in the manually annotated benchmark set to the names of the large model review items output by the large model used for contract review, and obtain a mapping table between the names of large model review items and manually reviewed items.
[0007] The review results of the large model after reviewing the contract are obtained are compared with the annotation results in the manual annotation benchmark set to obtain the evaluation results of the review quality of the large model.
[0008] In one embodiment, the names of manually reviewed items in the manually annotated benchmark set are mapped to the names of large model review items output by the large model used for contract review, resulting in a mapping table between large model review item names and manually reviewed item names, including:
[0009] Input the names of manually reviewed items from the manually annotated benchmark set into the trained text embedding model to obtain the first semantic vector;
[0010] The names of the large model review items output by the large model used for contract review are cleaned, and the cleaned names of the large model review items are input into the trained text embedding model to obtain the second semantic vector.
[0011] Calculate the similarity between the first semantic vector and the second semantic vector. If the similarity is greater than a preset threshold, determine the manual review item name corresponding to the first semantic vector and the large model review item name corresponding to the second semantic vector as the same review item.
[0012] Based on the obtained judgment results, a mapping table is constructed between the names of the large model review items and the names of the manual review items.
[0013] In one embodiment, the manual annotation benchmark set includes the contract identifier of the annotated contract, the name of the manually reviewed item, the risk type, the risk level, the risk description, the modification suggestion, and the location of the review item in the contract; the evaluation result is obtained based on at least one of the following: a first score for characterizing the identification quality of the large model, a second score for characterizing the output quality of the risk description and modification suggestion of the large model, a third score for characterizing the degree of distance difference between the location output by the large model and the location of the same review item in the manual annotation benchmark set, the review position compliance rate of the large model in conducting biased review of the preset position, or the legal citation accuracy rate for characterizing the reliability of the large model in citing laws and regulations;
[0014] The first score, second score, third score, legal citation accuracy rate, and review stance compliance rate are obtained based on the review results of the contract after reviewing it using the aforementioned large model.
[0015] In one embodiment, the process of obtaining the second score includes:
[0016] The risk descriptions and modification suggestions from the manually annotated benchmark set, as well as the risk descriptions and modification suggestions for the same review item from the review results, are input back into the large model, and a second score is output to characterize the output quality of the risk descriptions and modification suggestions from the large model.
[0017] In one embodiment, the process of obtaining the third score includes:
[0018] The positioning positions in the manually annotated benchmark set are compared with the positioning positions for the same review item in the review results to obtain the positioning distance between the two for the first character and the positioning intersection-union ratio between the two for the positioning range.
[0019] The third score is obtained based on the positioning distance and the positioning intersection-union ratio.
[0020] In one embodiment, the process of obtaining the review position compliance rate includes:
[0021] Obtain stance judgment prompts, input the review results and stance judgment prompts back into the large model, and obtain the judgment results used to characterize whether the review results conform to the preset stance;
[0022] Based on the judgment results, the percentage of review items that conform to the preset stance is obtained as the review stance compliance rate.
[0023] Secondly, this application also provides an evaluation device for contract review results, comprising:
[0024] The acquisition module acquires a set of manual annotation benchmarks for annotating contracts. The set of manual annotation benchmarks is obtained by annotating and verifying different types of de-identified contracts based on a rule document used to guide annotation. The rule document includes the name of the review item, risk type, risk level, risk definition, and location rules for locating review items in the contract.
[0025] The mapping module is used to map the names of manual review items in the manual annotation benchmark set to the names of large model review items output by the large model used for contract review, and obtain a mapping relationship table between the names of large model review items and manual review items.
[0026] The evaluation module is used to obtain the review results of the contract after the large model reviews it, and compares the review results mapped to the same review item with the annotation results in the manual annotation benchmark set to obtain the evaluation results of the review quality of the large model.
[0027] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0028] Obtain a set of manual annotation benchmarks for annotating contracts. The set of manual annotation benchmarks is obtained by annotating and verifying different types of anonymized contracts based on a rule document that guides annotation. The rule document includes the name of the review item, risk type, risk level, risk definition, and location rules for locating review items in the contract.
[0029] Map the names of manually reviewed items in the manually annotated benchmark set to the names of the large model review items output by the large model used for contract review, and obtain a mapping table between the names of large model review items and manually reviewed items.
[0030] The review results of the large model after reviewing the contract are obtained are compared with the annotation results in the manual annotation benchmark set to obtain the evaluation results of the review quality of the large model.
[0031] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0032] Obtain a set of manual annotation benchmarks for annotating contracts. The set of manual annotation benchmarks is obtained by annotating and verifying different types of anonymized contracts based on a rule document that guides annotation. The rule document includes the name of the review item, risk type, risk level, risk definition, and location rules for locating review items in the contract.
[0033] Map the names of manually reviewed items in the manually annotated benchmark set to the names of the large model review items output by the large model used for contract review, and obtain a mapping table between the names of large model review items and manually reviewed items.
[0034] The review results of the large model after reviewing the contract are obtained are compared with the annotation results in the manual annotation benchmark set to obtain the evaluation results of the review quality of the large model.
[0035] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0036] Obtain a set of manual annotation benchmarks for annotating contracts. The set of manual annotation benchmarks is obtained by annotating and verifying different types of anonymized contracts based on a rule document that guides annotation. The rule document includes the name of the review item, risk type, risk level, risk definition, and location rules for locating review items in the contract.
[0037] Map the names of manually reviewed items in the manually annotated benchmark set to the names of the large model review items output by the large model used for contract review, and obtain a mapping table between the names of large model review items and manually reviewed items.
[0038] The review results of the large model after reviewing the contract are obtained are compared with the annotation results in the manual annotation benchmark set to obtain the evaluation results of the review quality of the large model.
[0039] The evaluation methods, devices, computer equipment, computer-readable storage media, and computer program products for the aforementioned contract review results first involve constructing a manual annotation benchmark set for contract annotation. Guided by a rule document containing review item names, risk types, risk levels, risk definitions, and contract review item positioning rules, annotation work is carried out on different types of anonymized contracts. After annotation, a verification process is required, ultimately forming the manual annotation benchmark set. Next, to achieve comparability between the large-scale model review results and the manual annotation results, the names of the manual review items in the manual annotation benchmark set are matched and associated with the names of the large-scale model review items output by the large-scale model used for contract review, generating a mapping relationship table between the two. This effectively eliminates comparison biases caused by differences in expression and classification between the large-scale model output review items and the manually annotated review items, ensuring accurate comparison based on the same review item and significantly improving evaluation accuracy. Finally, the results of the large-scale model's contract review are obtained. Based on the aforementioned mapping relationship table, the content in the large-scale model review results that is classified into the same review item is compared and analyzed with the annotation results corresponding to that review item in the manual annotation benchmark set. Through this process, an evaluation result that can measure the quality of the large-scale model's contract review is obtained. The standardized process of generating a manually labeled benchmark set, along with fixed mapping and comparison logic, reduces subjective human interference, making the evaluation process more standardized and reproducible. Furthermore, by using verified manually labeled results as the standard, the performance of the large model in review can be clearly quantified, providing a reliable basis for its optimization and iteration. At the same time, the benchmark set is closely aligned with the real business scenarios of various anonymized contracts, enabling the evaluation results to directly reflect the actual applicability of the large model and providing clear guidance for enterprises to determine whether it meets their contract review requirements. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 This is a diagram illustrating the application environment of a method for evaluating contract review results in one embodiment.
[0042] Figure 2 This is a flowchart illustrating a method for evaluating contract review results in one embodiment;
[0043] Figure 3 This is a structural block diagram of a device for evaluating contract review results in one embodiment;
[0044] Figure 4 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0046] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0047] The method for evaluating contract review results provided in this application embodiment can be applied to, for example... Figure 1 The application environment shown can involve only terminal 102, only server 104, or both terminal 102 and server 104, wherein terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated on server 104 or located on a cloud or other network server. Specifically, terminal 102 or server 104 completes a method for evaluating contract review results, which includes:
[0048] Obtain a set of manual annotation benchmarks for annotating contracts. The set of manual annotation benchmarks is obtained by annotating and verifying different types of anonymized contracts based on a rule document that guides annotation. The rule document includes the name of the review item, risk type, risk level, risk definition, and location rules for locating review items in the contract.
[0049] Map the names of manually reviewed items in the manually annotated benchmark set to the names of the large model review items output by the large model used for contract review, and obtain a mapping table between the names of large model review items and manually reviewed items.
[0050] The review results of the large model after reviewing the contract are obtained are compared with the annotation results in the manual annotation benchmark set to obtain the evaluation results of the review quality of the large model.
[0051] Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection equipment, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0052] In one exemplary embodiment, such as Figure 2 As shown, a method for evaluating contract review results is provided, which is then applied to... Figure 1 Taking the server in the example, the explanation includes the following steps 202 to 206. Wherein:
[0053] Step 202: Obtain the manual annotation benchmark set for annotating the contract. The manual annotation benchmark set is obtained by annotating and verifying different types of de-identified contracts based on the rule document used to guide the annotation. The rule document includes the review item name, risk type, risk level, risk definition, and location rules for locating review items in the contract.
[0054] The review item name corresponds to a specific review item category. Review item categories may differ depending on the contract. Taking labor contracts as an example, review item categories include eight categories: contract term, job position and responsibilities, job adjustments, other structural elements, benefits and subsidies, identification of employment form, handover and settlement upon departure, and number of contract copies and delivery. If a review item category is a probationary period clause, the review item name will include the probationary period duration and probationary period salary. Risk types include structural risk, legal risk, and transaction risk. Risk type is an attribute of the review item, indicating whether the risk stems from the contract structure, legal provisions, or transaction process. Risk levels are categorized as high, medium, and low.
[0055] Step 204: Map the names of the manually reviewed items in the manually labeled benchmark set to the names of the large model reviews output by the large model used for contract review, and obtain a mapping table between the names of the large model reviews and the names of the manually reviewed items.
[0056] Since the names of the review items output by the large model may have different expressions (such as "payment time unknown" vs "payment period not agreed"), it is necessary to achieve accurate matching with the review items of the large model through semantic mapping to ensure that the review items output by the large model correspond one-to-one with the manual review items and eliminate the comparison bias caused by differences in expression or classification.
[0057] Step 206: Obtain the review results of the contract after the large model reviews it. Compare the review results mapped to the same review item with the annotation results in the manual annotation benchmark set to obtain the evaluation results of the review quality of the large model.
[0058] The process involves first obtaining the review results output by the large model after reviewing the contract, then mapping these results to a unified review item system, and finally comparing the large model review results belonging to the same review item with the corresponding manual annotation results in the manual annotation benchmark set item by item. By comparing and analyzing the consistency between the two (such as whether they both identify the same risk point, whether the judgment on the compliance of the clauses is consistent, etc.), the final evaluation results (such as quantitative indicators such as accuracy and recall) that can measure the quality of the large model's contract review are obtained.
[0059] The evaluation method for the above contract review results first involves constructing a manual annotation benchmark set for contract annotation. Guided by a rule document containing review item names, risk types, risk levels, risk definitions, and contract review item positioning rules, annotation work is carried out on different types of anonymized contracts. After annotation, a verification process is required, ultimately forming the manual annotation benchmark set. Next, to ensure comparability between the large-scale model review results and the manual annotation results, the names of the manually annotated review items in the manual annotation benchmark set are matched and associated with the names of the large-scale model review items output by the large-scale model used for contract review, generating a mapping relationship table between the two. This effectively eliminates the comparison bias caused by differences in expression and classification between the large-scale model output review items and the manually annotated review items, ensuring accurate comparison based on the same review item and significantly improving evaluation accuracy. Finally, the results of the large-scale model's contract review are obtained. Based on the aforementioned mapping relationship table, the content in the large-scale model review results that is classified into the same review item is compared and analyzed with the corresponding annotation results of the review item in the manual annotation benchmark set. Through this process, an evaluation result that can measure the quality of the large-scale model's contract review is obtained. The standardized process of generating a manually labeled benchmark set, along with fixed mapping and comparison logic, reduces subjective human interference, making the evaluation process more standardized and reproducible. Furthermore, by using verified manually labeled results as the standard, the performance of the large model in review can be clearly quantified, providing a reliable basis for its optimization and iteration. At the same time, the benchmark set is closely aligned with the real business scenarios of various anonymized contracts, enabling the evaluation results to directly reflect the actual applicability of the large model and providing clear guidance for enterprises to determine whether it meets their contract review requirements.
[0060] In one embodiment, the names of manually reviewed items in the manually annotated benchmark set are mapped to the names of large model review items output by the large model used for contract review, resulting in a mapping table between large model review item names and manually reviewed item names, including:
[0061] Input the names of manually reviewed items from the manually annotated benchmark set into the trained text embedding model to obtain the first semantic vector;
[0062] The names of the large model review items output by the large model used for contract review are cleaned, and the cleaned names of the large model review items are input into the trained text embedding model to obtain the second semantic vector.
[0063] Calculate the similarity between the first semantic vector and the second semantic vector. If the similarity is greater than a preset threshold, determine the manual review item name corresponding to the first semantic vector and the large model review item name corresponding to the second semantic vector as the same review item.
[0064] Based on the obtained judgment results, a mapping table is constructed between the names of the large model review items and the names of the manual review items.
[0065] The specific steps are as follows:
[0066] S1: Construct a semantic vector library for review items
[0067] Input the manually labeled review item names (such as "subject qualification risk" and "payment risk") into a pre-trained legal semantic model (such as a BGE-Large model trained with legal specialization) to obtain the first semantic vector and construct a "manual review item-vector" mapping library.
[0068] S2: Calculate the semantic similarity between the large model review items and the human review items.
[0069] Text cleaning (removing punctuation and standardizing terminology) is performed on the names of the review items output by the large model. The second semantic vector is obtained through the legal semantic model. The cosine similarity between the first semantic vector (human review item vector) and the second semantic vector (large model review item vector) is calculated. A similarity threshold (e.g., 0.85) is set. If the similarity exceeds the threshold, it is determined to be the same review item. If there are multiple similar human review items, the one with the highest similarity is selected as the mapping result.
[0070] S3: Construct a mapping table for review items
[0071] Output a mapping table of {large model review item name, corresponding manual review item name, similarity score} to eliminate the impact of differences in review item descriptions on subsequent evaluations.
[0072] In the above embodiments, a mapping table is established between the names of the large model review items and the names of the manually reviewed items. This effectively eliminates the comparison deviation caused by differences in expression and classification between the large model output review items and the manually labeled review items, ensuring that the two are accurately compared based on the same review item, and greatly improving the accuracy of the evaluation.
[0073] In one embodiment, the manual annotation benchmark set includes the contract identifier of the annotated contract, the name of the manually reviewed item, the risk type, the risk level, the risk description, the modification suggestion, and the location of the review item in the contract; the evaluation result is obtained based on at least one of the following: a first score for characterizing the identification quality of the large model, a second score for characterizing the output quality of the risk description and modification suggestion of the large model, a third score for characterizing the degree of distance difference between the location output by the large model and the location of the same review item in the manual annotation benchmark set, the review position compliance rate of the large model in conducting biased review of the preset position, or the legal citation accuracy rate for characterizing the reliability of the large model in citing laws and regulations;
[0074] The first score, second score, third score, legal citation accuracy rate, and review stance compliance rate are obtained based on the review results of the contract after reviewing it using the aforementioned large model.
[0075] Specifically, based on the mapping relationship table and combined with the prompt-guided engineering model, the review results of the large model are automatically evaluated from the following five dimensions:
[0076] S1: First score, basic risk indicator assessment
[0077] The review results output by the large model (mapped human review items) are compared item by item with the human-annotated benchmark set, and the following calculations are performed:
[0078] 1) First ratio: Precision, the number of risk points correctly identified by the large model / the total number of risk points identified by the large model (correct identification means that the risk type, level and location text are consistent);
[0079] 2) Second ratio: Recall, the number of risk points correctly matched by human intervention / the total number of risk points manually labeled;
[0080] 3) First score = 2 × precision × recall / (precision + recall).
[0081] S2: Second score, risk description and modification suggestions quality rating
[0082] Design prompt word templates to guide the large model in semantically scoring the "risk statements" and "modification suggestions" output by the large model against those in the human benchmark set. The operability of the "modification suggestions" (whether the modification direction is clear, whether it is executable, and whether it conforms to legal terminology) is scored, and the average score of the two dimensions is calculated as the quality score result.
[0083] S3: Third score, location text evaluation
[0084] For the "location text" (the positions of the first and last characters in the contract text) output by the large model and the "location text" of the human baseline, calculate:
[0085] 1) Start Distance (SD): |Start position of text located by large model -start position of text located manually|, the smaller the SD, the higher the accuracy of the localization;
[0086] 2) Intersection over Union (IoU): (Number of characters in the large model location text ∩ the manually located text) / (Number of characters in the large model location text ∪ the manually located text). The closer the IoU is to 1, the higher the degree of location overlap.
[0087] S4: Compliance with review stance
[0088] Clearly define the pre-established stance (Party A or Party B) for contract review. Based on the core interests, key protections, and risk avoidance directions of this stance, design targeted stance judgment prompts. These prompts should clearly define the boundaries of the stance and guide the large model to evaluate its own output review results. The model should determine whether the results are consistent with the core interests of the pre-established stance and whether there are any issues such as stance deviation, misalignment of interests, or failure to adequately protect the rights and interests of the pre-established party.
[0089] Among all the review items output by the statistical model, the number of review items that are determined to meet the preset position requirements by the above assessment is then calculated as the proportion of this number to the total number of all review items. This proportion is the review position compliance rate.
[0090] S5: Accuracy of Legal Citations
[0091] Based on Retrieval-Augmented Generation (RAG), a legal and regulatory database (such as the Civil Code and related judicial interpretations) is constructed. The steps are as follows:
[0092] 1) Legal Retrieval: Based on the legal provisions mentioned in the large model (such as "Article 577 of the Civil Code"), retrieve the corresponding legal texts in the database;
[0093] 2) Judgment of the Reasonableness of Citation: The design prompts guide the large-scale model to judge a) whether the cited law is correct and b) whether the citation is appropriate and logical. The percentage of reasonably cited clauses is the legal citation accuracy rate.
[0094] In the above embodiments, a multi-dimensional and refined evaluation index system is constructed (including at least one score such as recognition quality, output quality, positioning accuracy, position compliance rate, and legal citation accuracy rate). This achieves a comprehensive evaluation of the large model's contract review performance—focusing not only on the accuracy of risk identification but also on the practicality of risk descriptions and modification suggestions, the precision of review item positioning, the compliance of position compliance, and the reliability of legal citations, avoiding the bias caused by a single index. Furthermore, relying on the detailed annotation information of the benchmark set and the quantitative calculation logic of each score, the evaluation results are more targeted and valuable for reference. This system can accurately identify the advantages and disadvantages of the large model in different review stages, providing a more comprehensive and reliable basis for the targeted optimization and iteration of the large model and its adaptation to the core needs of actual contract review scenarios.
[0095] In one embodiment, the process of obtaining the second score includes:
[0096] The risk descriptions and modification suggestions from the manually annotated benchmark set, as well as the risk descriptions and modification suggestions for the same review item from the review results, are input back into the large model, and a second score is output to characterize the output quality of the risk descriptions and modification suggestions from the large model.
[0097] The second score is the quality assessment of the risk description and modification suggestions: A prompt word template is designed to guide the large model in performing semantic quality assessments of the "risk description" and "modification suggestions" output by the large model and those from the human benchmark set. The operability of the "modification suggestions" (whether the modification direction is clear, whether it is executable, and whether it conforms to legal terminology) is also assessed, and the average score of the two dimensions is calculated as the quality score result.
[0098] In the above embodiments, authoritative and standardized risk descriptions and modification suggestions from manually annotated benchmarks are input into the large model along with the corresponding content output by the large model for the same review item for comparison and evaluation. This accurately quantifies the degree of matching between the two in terms of semantic consistency, logical rationality, and practicality, thereby obtaining an objective second score that fits the actual application needs. This avoids the problems of strong subjectivity, low efficiency, and inconsistent standards in manual evaluation, and focuses on core quality dimensions such as the accuracy of risk descriptions and the operability of modification suggestions. This makes the quality evaluation of the large model's output more targeted and reliable, providing accurate and efficient reference for subsequent optimization of the large model's ability to interpret risks and provide solutions in contract review.
[0099] In one embodiment, the process of obtaining the third score includes:
[0100] The positioning positions in the manually annotated benchmark set are compared with the positioning positions for the same review item in the review results to obtain the positioning distance between the two for the first character and the positioning intersection-union ratio between the two for the positioning range.
[0101] The third score is obtained based on the positioning distance and the positioning intersection-union ratio.
[0102] The third score is for location text evaluation: This is calculated by comparing the "location text" (the start-end character positions in the contract text) output by the large model with the "location text" of the human benchmark.
[0103] 1) Start Distance (SD): |Start position of text located by large model -start position of text located manually|, the smaller the SD, the higher the accuracy of the localization;
[0104] 2) Intersection over Union (IoU): (Number of characters in the large model location text ∩ the manually located text) / (Number of characters in the large model location text ∪ the manually located text). The closer the IoU is to 1, the higher the degree of location overlap.
[0105] In the above embodiments, by splitting the comparison dimensions of the positioning location—both calculating the positioning distance of the first character to quantify the deviation of the core positioning point, and measuring the degree of overlap of the positioning range through the positioning intersection-union ratio—a third score is derived based on two specific and quantifiable indicators, making the evaluation results of positioning accuracy more objective and traceable, and accurately reflecting the accuracy and completeness of the positioning review items in the contract text of the large model.
[0106] In one embodiment, the process of obtaining the review position compliance rate includes:
[0107] Obtain stance judgment prompts, input the review results and stance judgment prompts back into the large model, and obtain the judgment results used to characterize whether the review results conform to the preset stance;
[0108] Based on the judgment results, the percentage of review items that conform to the preset stance is obtained as the review stance compliance rate.
[0109] Review Position Compliance Rate: Clearly define the pre-set position (Party A or Party B) for contract review. Based on key dimensions such as the core interests, priority of rights protection, and risk avoidance of this position, design targeted position judgment prompts. These prompts must clearly define the boundaries of the position and guide the large-scale model to evaluate its own output review results. The model should determine whether these results align with the core demands of the pre-set position and whether there are any deviations, misalignments of interests, or insufficient protection of the pre-set party's rights. Subsequently, count the number of review items output by the large-scale model that meet the pre-set position requirements based on the above evaluation. Then calculate the proportion of this number to the total number of review items; this proportion is the review position compliance rate.
[0110] In the above embodiments, by designing targeted position judgment prompts, the large model is guided to autonomously assess the degree of conformity between its review results and the preset position (Party A / Party B). This not only leverages the semantic understanding capabilities of the large model to achieve efficient verification of position conformity, avoiding the subjectivity and inefficiency of manual assessment, but also makes position assessment more targeted by clearly focusing on the interest orientation and core demands of the preset position. At the same time, the position compliance rate is quantified by the "percentage of review items that meet the position requirements," which provides an intuitive and quantifiable presentation of the large model's position adaptation capability in contract review. This accurately reflects whether the large model aligns with the interests of a specific party, providing a reliable and easily implementable reference for optimizing the large model's position-oriented review capabilities and meeting the personalized contract review needs of different entities.
[0111] In one embodiment, a method for evaluating contract review results is presented. This method mainly includes three stages: manual annotation standardization, review item mapping, and multi-dimensional evaluation, as detailed below:
[0112] (1) Standardization of manual annotation and data preparation
[0113] S1: Establish unified labeling standards
[0114] Led by the Legal Affairs Committee, and based on laws and regulations such as the Civil Code and the Contract Law, as well as industry practices, the "Contract Risk Point Marking Standard" was formulated. This standard clarifies the classification of review items (such as "subject qualification clauses", "payment clauses", "breach of contract liability clauses", etc.), risk levels ("high", "medium", "low"), risk definitions (such as "payment agreement not specifying the time point" is defined as "payment risk"), and positioning rules (risk texts must be accurate to the sentence level), ensuring the authority and consistency of the marking standard.
[0115] S2: Manually marked risk points
[0116] A sample library of anonymized contracts covering multiple industries (such as finance, manufacturing, and services) and multiple types (such as procurement contracts, labor contracts, and lease contracts) was selected. Trained legal experts manually annotated the contracts according to the "Annotation Specifications" and output structured annotated data, including: {Contract ID, Review Item Name, Risk Type, Risk Level, Risk Description, Modification Suggestion, Location Text (Start Character Position - End Character Position)}.
[0117] The review item name corresponds to a specific review item category. If a review item category is a probationary period clause, the review item name will include the probationary period duration and probationary period salary. Risk types include structural risk, legal risk, and transaction risk. The risk description outlines the consequences that this risk point may cause.
[0118] S3: Review and Optimization of Annotated Data
[0119] The manual annotation results are reviewed through a "double review + legal committee final review" mechanism to remove annotation errors (such as positioning text deviations and misjudgments of risk types), supplement missing annotations, and form a high-quality, standardized "manual annotation benchmark set" as the standard for subsequent evaluation.
[0120] (2) Mapping between large model review items and manually labeled review items
[0121] Because the names of the review items output by the large model may have different expressions (e.g., "Payment time unknown" vs. "Payment period not agreed upon"), semantic mapping is needed to achieve accurate matching with the review items of the large model. The specific steps are as follows:
[0122] S1: Construct a semantic vector library for review items
[0123] Input manually labeled review item names (such as "subject qualification risk" and "payment risk") into a pre-trained legal semantic model (such as a BGE-Large model trained with legal specialization) to obtain high-dimensional semantic vectors and construct a "manual review item-vector" mapping library.
[0124] S2: Calculate the semantic similarity between the large model review items and the human review items.
[0125] Text cleaning (removing punctuation and standardizing terminology) is performed on the names of the review items output by the large model. Semantic vectors are also obtained through the legal semantic model. The cosine similarity between the large model review item vector and the manual review item vector is calculated. A similarity threshold (e.g., 0.85) is set. If the similarity exceeds the threshold, it is determined to be the same review item. If there are multiple similar manual review items, the one with the highest similarity is selected as the mapping result.
[0126] S3: Construct a mapping table for review items
[0127] Output a mapping table of {large model review item name, corresponding manual review item name, similarity score} to eliminate the impact of differences in review item descriptions on subsequent evaluations.
[0128] (3) Multi-dimensional automated evaluation
[0129] Based on the mapping table and combined with the prompt-guided engineering model, the review results of the large model are automatically evaluated from the following five dimensions:
[0130] S1: First score, basic risk indicator assessment
[0131] The review results output by the large model (mapped human review items) are compared item by item with the human-annotated benchmark set, and the following calculations are performed:
[0132] 1) First ratio: Precision, the number of risk points correctly identified by the large model / the total number of risk points identified by the large model (correct identification means that the risk type, level and location text are consistent);
[0133] 2) Second ratio: Recall, the number of risk points correctly matched by human intervention / the total number of risk points manually labeled;
[0134] 3) First score = 2 × precision × recall / (precision + recall).
[0135] S2: Second score, risk description and modification suggestions quality rating
[0136] Design prompt word templates to guide the large model in semantically scoring the "risk statements" and "modification suggestions" output by the large model against those in the human benchmark set. The operability of the "modification suggestions" (whether the modification direction is clear, whether it is executable, and whether it conforms to legal terminology) is scored, and the average score of the two dimensions is calculated as the quality score result.
[0137] S3: Third score, location text evaluation
[0138] For the "location text" (the positions of the first and last characters in the contract text) output by the large model and the "location text" of the human baseline, calculate:
[0139] 1) Start Distance (SD): |Start position of text located by large model -start position of text located manually|, the smaller the SD, the higher the accuracy of the localization;
[0140] 2) Intersection over Union (IoU): (Number of characters in the large model location text ∩ the manually located text) / (Number of characters in the large model location text ∪ the manually located text). The closer the IoU is to 1, the higher the degree of location overlap.
[0141] S4: Compliance with review stance
[0142] Clearly define the pre-established stance (Party A or Party B) for contract review. Based on the core interests, key protections, and risk avoidance directions of this stance, design targeted stance judgment prompts. These prompts should clearly define the boundaries of the stance and guide the large model to evaluate its own output review results. The model should determine whether the results are consistent with the core interests of the pre-established stance and whether there are any issues such as stance deviation, misalignment of interests, or failure to adequately protect the rights and interests of the pre-established party.
[0143] Among all the review items output by the statistical model, the number of review items that are determined to meet the preset position requirements by the above assessment is then calculated as the proportion of this number to the total number of all review items. This proportion is the review position compliance rate.
[0144] S5: Accuracy of Legal Citations
[0145] Based on Retrieval-Augmented Generation (RAG), a legal and regulatory database (such as the Civil Code and related judicial interpretations) is constructed. The steps are as follows:
[0146] 1) Legal Retrieval: Based on the legal provisions mentioned in the large model (such as "Article 577 of the Civil Code"), retrieve the corresponding legal texts in the database;
[0147] 2) Judgment of the Reasonableness of Citation: The design prompts guide the large-scale model to judge a) whether the cited law is correct and b) whether the citation is appropriate and logical. The percentage of reasonably cited clauses is the legal citation accuracy rate.
[0148] Finally, a table is generated, which allows you to view the scores in five dimensions, intuitively judge the compliance and accuracy of the contract review results, and also view the risk points, risk descriptions, and corresponding modification methods of the contract.
[0149] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0150] Based on the same inventive concept, this application also provides a contract review result evaluation device for implementing the contract review result evaluation method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations of one or more contract review result evaluation device embodiments provided below can be found in the limitations of the contract review result evaluation method described above, and will not be repeated here.
[0151] In one exemplary embodiment, such as Figure 3 As shown, a device for evaluating contract review results is provided, comprising: an acquisition module 302, a mapping module 304, and an evaluation module 306, wherein:
[0152] The module 302 obtains a set of manual annotation benchmarks for annotating contracts. The set of manual annotation benchmarks is obtained by annotating and verifying different types of de-identified contracts based on a rule document that guides annotation. The rule document includes the name of the review item, risk type, risk level, risk definition, and location rules for locating review items in the contract.
[0153] The mapping module 304 is used to map the names of manual review items in the manual annotation benchmark set to the names of large model review items output by the large model used for contract review, and obtain a mapping relationship table between the names of large model review items and manual review items.
[0154] Evaluation module 306 is used to obtain the review results of the contract after the large model reviews it, compare the review results mapped to the same review item with the annotation results in the manual annotation benchmark set, and obtain the evaluation results of the review quality of the large model.
[0155] In one embodiment, the mapping module 304 is specifically used to input the names of manually reviewed items from the manually labeled benchmark set into the trained text embedding model to obtain a first semantic vector; clean the names of the large model review items output by the large model used for contract review, input the cleaned large model review item names into the trained text embedding model to obtain a second semantic vector; calculate the similarity between the first semantic vector and the second semantic vector, and if the similarity is greater than a preset threshold, determine the manual review item name corresponding to the first semantic vector and the large model review item name corresponding to the second semantic vector as the same review item;
[0156] Based on the obtained judgment results, a mapping table is constructed between the names of the large model review items and the names of the manual review items.
[0157] In one embodiment, the manual annotation benchmark set includes the contract identifier of the annotated contract, the name of the manually reviewed item, the risk type, the risk level, the risk description, the modification suggestion, and the location of the review item in the contract; the evaluation result is obtained based on at least one of the following: a first score for characterizing the identification quality of the large model, a second score for characterizing the output quality of the risk description and modification suggestion of the large model, a third score for characterizing the degree of distance difference between the location output by the large model and the location of the same review item in the manual annotation benchmark set, the review position compliance rate of the large model in conducting biased review of the preset position, or the legal citation accuracy rate for characterizing the reliability of the large model in citing laws and regulations;
[0158] The first score, second score, third score, legal citation accuracy rate, and review stance compliance rate are obtained based on the review results of the contract after reviewing it using the aforementioned large model.
[0159] In one embodiment, the acquisition module 302 is further configured to input the risk descriptions and modification suggestions from the manually labeled benchmark set, as well as the risk descriptions and modification suggestions for the same review item from the review results, back into the large model, and output a second score to characterize the output quality of the risk descriptions and modification suggestions from the large model.
[0160] In one embodiment, the acquisition module 302 is further configured to compare the positioning position in the manual annotation benchmark set with the positioning position for the same review item in the review results, and obtain the positioning distance between the two for the positioning position of the first character, and the positioning intersection-union ratio between the two for the positioning range.
[0161] The third score is obtained based on the positioning distance and the positioning intersection-union ratio.
[0162] In one embodiment, the acquisition module 302 is further configured to acquire stance judgment prompts, input the review results and stance judgment prompts back into the large model to obtain a judgment result that characterizes whether the review results conform to a preset stance; based on the judgment result, the proportion of review items that conform to the preset stance is acquired as the review stance compliance rate.
[0163] Each module in the aforementioned contract review result evaluation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0164] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a method for evaluating contract review results. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0165] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0166] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0167] Obtain a set of manual annotation benchmarks for annotating contracts. The set of manual annotation benchmarks is obtained by annotating and verifying different types of anonymized contracts based on a rule document that guides annotation. The rule document includes the name of the review item, risk type, risk level, risk definition, and location rules for locating review items in the contract.
[0168] Map the names of manually reviewed items in the manually annotated benchmark set to the names of the large model review items output by the large model used for contract review, and obtain a mapping table between the names of large model review items and manually reviewed items.
[0169] The review results of the large model after reviewing the contract are obtained are compared with the annotation results in the manual annotation benchmark set to obtain the evaluation results of the review quality of the large model.
[0170] In one embodiment, the names of manually reviewed items in the manually annotated benchmark set are mapped to the names of large model review items output by the large model used for contract review, resulting in a mapping table between large model review item names and manually reviewed item names, including:
[0171] Input the names of manually reviewed items from the manually annotated benchmark set into the trained text embedding model to obtain the first semantic vector;
[0172] The names of the large model review items output by the large model used for contract review are cleaned, and the cleaned names of the large model review items are input into the trained text embedding model to obtain the second semantic vector.
[0173] Calculate the similarity between the first semantic vector and the second semantic vector. If the similarity is greater than a preset threshold, determine the manual review item name corresponding to the first semantic vector and the large model review item name corresponding to the second semantic vector as the same review item.
[0174] Based on the obtained judgment results, a mapping table is constructed between the names of the large model review items and the names of the manual review items.
[0175] In one embodiment, the manual annotation benchmark set includes the contract identifier of the annotated contract, the name of the manually reviewed item, the risk type, the risk level, the risk description, the modification suggestion, and the location of the review item in the contract; the evaluation result is obtained based on at least one of the following: a first score for characterizing the identification quality of the large model, a second score for characterizing the output quality of the risk description and modification suggestion of the large model, a third score for characterizing the degree of distance difference between the location output by the large model and the location of the same review item in the manual annotation benchmark set, the review position compliance rate of the large model in conducting biased review of the preset position, or the legal citation accuracy rate for characterizing the reliability of the large model in citing laws and regulations;
[0176] The first score, second score, third score, legal citation accuracy rate, and review stance compliance rate are obtained based on the review results of the contract after reviewing it using the aforementioned large model.
[0177] In one embodiment, the process of obtaining the second score includes:
[0178] The risk descriptions and modification suggestions from the manually annotated benchmark set, as well as the risk descriptions and modification suggestions for the same review item from the review results, are input back into the large model, and a second score is output to characterize the output quality of the risk descriptions and modification suggestions from the large model.
[0179] In one embodiment, the process of obtaining the third score includes:
[0180] The positioning positions in the manually annotated benchmark set are compared with the positioning positions for the same review item in the review results to obtain the positioning distance between the two for the first character and the positioning intersection-union ratio between the two for the positioning range.
[0181] The third score is obtained based on the positioning distance and the positioning intersection-union ratio.
[0182] In one embodiment, the process of obtaining the review position compliance rate includes:
[0183] Obtain stance judgment prompts, input the review results and stance judgment prompts back into the large model, and obtain the judgment results used to characterize whether the review results conform to the preset stance;
[0184] Based on the judgment results, the percentage of review items that conform to the preset stance is obtained as the review stance compliance rate.
[0185] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0186] Obtain a set of manual annotation benchmarks for annotating contracts. The set of manual annotation benchmarks is obtained by annotating and verifying different types of anonymized contracts based on a rule document that guides annotation. The rule document includes the name of the review item, risk type, risk level, risk definition, and location rules for locating review items in the contract.
[0187] Map the names of manually reviewed items in the manually annotated benchmark set to the names of the large model review items output by the large model used for contract review, and obtain a mapping table between the names of large model review items and manually reviewed items.
[0188] The review results of the large model after reviewing the contract are obtained are compared with the annotation results in the manual annotation benchmark set to obtain the evaluation results of the review quality of the large model.
[0189] In one embodiment, the names of manually reviewed items in the manually annotated benchmark set are mapped to the names of large model review items output by the large model used for contract review, resulting in a mapping table between large model review item names and manually reviewed item names, including:
[0190] Input the names of manually reviewed items from the manually annotated benchmark set into the trained text embedding model to obtain the first semantic vector;
[0191] The names of the large model review items output by the large model used for contract review are cleaned, and the cleaned names of the large model review items are input into the trained text embedding model to obtain the second semantic vector.
[0192] Calculate the similarity between the first semantic vector and the second semantic vector. If the similarity is greater than a preset threshold, determine the manual review item name corresponding to the first semantic vector and the large model review item name corresponding to the second semantic vector as the same review item.
[0193] Based on the obtained judgment results, a mapping table is constructed between the names of the large model review items and the names of the manual review items.
[0194] In one embodiment, the manual annotation benchmark set includes the contract identifier of the annotated contract, the name of the manually reviewed item, the risk type, the risk level, the risk description, the modification suggestion, and the location of the review item in the contract; the evaluation result is obtained based on at least one of the following: a first score for characterizing the identification quality of the large model, a second score for characterizing the output quality of the risk description and modification suggestion of the large model, a third score for characterizing the degree of distance difference between the location output by the large model and the location of the same review item in the manual annotation benchmark set, the review position compliance rate of the large model in conducting biased review of the preset position, or the legal citation accuracy rate for characterizing the reliability of the large model in citing laws and regulations;
[0195] The first score, second score, third score, legal citation accuracy rate, and review stance compliance rate are obtained based on the review results of the contract after reviewing it using the aforementioned large model.
[0196] In one embodiment, the process of obtaining the second score includes:
[0197] The risk descriptions and modification suggestions from the manually annotated benchmark set, as well as the risk descriptions and modification suggestions for the same review item from the review results, are input back into the large model, and a second score is output to characterize the output quality of the risk descriptions and modification suggestions from the large model.
[0198] In one embodiment, the process of obtaining the third score includes:
[0199] The positioning positions in the manually annotated benchmark set are compared with the positioning positions for the same review item in the review results to obtain the positioning distance between the two for the first character and the positioning intersection-union ratio between the two for the positioning range.
[0200] The third score is obtained based on the positioning distance and the positioning intersection-union ratio.
[0201] In one embodiment, the process of obtaining the review position compliance rate includes:
[0202] Obtain stance judgment prompts, input the review results and stance judgment prompts back into the large model, and obtain the judgment results used to characterize whether the review results conform to the preset stance;
[0203] Based on the judgment results, the percentage of review items that conform to the preset stance is obtained as the review stance compliance rate.
[0204] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:
[0205] Obtain a set of manual annotation benchmarks for annotating contracts. The set of manual annotation benchmarks is obtained by annotating and verifying different types of anonymized contracts based on a rule document that guides annotation. The rule document includes the name of the review item, risk type, risk level, risk definition, and location rules for locating review items in the contract.
[0206] Map the names of manually reviewed items in the manually annotated benchmark set to the names of the large model review items output by the large model used for contract review, and obtain a mapping table between the names of large model review items and manually reviewed items.
[0207] The review results of the large model after reviewing the contract are obtained are compared with the annotation results in the manual annotation benchmark set to obtain the evaluation results of the review quality of the large model.
[0208] In one embodiment, the names of manually reviewed items in the manually annotated benchmark set are mapped to the names of large model review items output by the large model used for contract review, resulting in a mapping table between large model review item names and manually reviewed item names, including:
[0209] Input the names of manually reviewed items from the manually annotated benchmark set into the trained text embedding model to obtain the first semantic vector;
[0210] The names of the large model review items output by the large model used for contract review are cleaned, and the cleaned names of the large model review items are input into the trained text embedding model to obtain the second semantic vector.
[0211] Calculate the similarity between the first semantic vector and the second semantic vector. If the similarity is greater than a preset threshold, determine the manual review item name corresponding to the first semantic vector and the large model review item name corresponding to the second semantic vector as the same review item.
[0212] Based on the obtained judgment results, a mapping table is constructed between the names of the large model review items and the names of the manual review items.
[0213] In one embodiment, the manual annotation benchmark set includes the contract identifier of the annotated contract, the name of the manually reviewed item, the risk type, the risk level, the risk description, the modification suggestion, and the location of the review item in the contract; the evaluation result is obtained based on at least one of the following: a first score for characterizing the identification quality of the large model, a second score for characterizing the output quality of the risk description and modification suggestion of the large model, a third score for characterizing the degree of distance difference between the location output by the large model and the location of the same review item in the manual annotation benchmark set, the review position compliance rate of the large model in conducting biased review of the preset position, or the legal citation accuracy rate for characterizing the reliability of the large model in citing laws and regulations;
[0214] The first score, second score, third score, legal citation accuracy rate, and review stance compliance rate are obtained based on the review results of the contract after reviewing it using the aforementioned large model.
[0215] In one embodiment, the process of obtaining the second score includes:
[0216] The risk descriptions and modification suggestions from the manually annotated benchmark set, as well as the risk descriptions and modification suggestions for the same review item from the review results, are input back into the large model, and a second score is output to characterize the output quality of the risk descriptions and modification suggestions from the large model.
[0217] In one embodiment, the process of obtaining the third score includes:
[0218] The positioning positions in the manually annotated benchmark set are compared with the positioning positions for the same review item in the review results to obtain the positioning distance between the two for the first character and the positioning intersection-union ratio between the two for the positioning range.
[0219] The third score is obtained based on the positioning distance and the positioning intersection-union ratio.
[0220] In one embodiment, the process of obtaining the review position compliance rate includes:
[0221] Obtain stance judgment prompts, input the review results and stance judgment prompts back into the large model, and obtain the judgment results used to characterize whether the review results conform to the preset stance;
[0222] Based on the judgment results, the percentage of review items that conform to the preset stance is obtained as the review stance compliance rate.
[0223] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0224] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0225] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0226] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for evaluating the results of contract review, characterized in that, The method includes: Obtain a set of manual annotation benchmarks for annotating contracts. The set of manual annotation benchmarks is obtained by annotating and verifying different types of de-identified contracts based on a rule document used to guide annotation. The rule document includes the name of the review item, risk type, risk level, risk definition, and location rules for locating review items in the contract. Map the names of manually reviewed items in the manually annotated benchmark set to the names of the large model review items output by the large model used for contract review, and obtain a mapping table between the names of large model review items and manually reviewed items. The review results of the contract reviewed by the large model are obtained, and the review results mapped to the same review item are compared with the annotation results in the manual annotation benchmark set to obtain the evaluation results of the review quality of the large model.
2. The method according to claim 1, characterized in that, The step of mapping the names of manually reviewed items in the manually annotated benchmark set to the names of large model review items output by the large model used for contract review, to obtain a mapping table between large model review item names and manually reviewed item names, includes: The names of manually reviewed items from the manually annotated benchmark set are input into the trained text embedding model to obtain the first semantic vector. The names of the large model review items output by the large model used for contract review are cleaned, and the cleaned names of the large model review items are input into the trained text embedding model to obtain the second semantic vector. Calculate the similarity between the first semantic vector and the second semantic vector. If the similarity is greater than a preset threshold, determine the manual review item name corresponding to the first semantic vector and the large model review item name corresponding to the second semantic vector as the same review item. Based on the obtained judgment results, a mapping table is constructed between the names of the large model review items and the names of the manual review items.
3. The method according to claim 1, characterized in that, The manually labeled benchmark set includes the contract identifier of the labeled contract, the name of the manually reviewed item, the risk type, the risk level, the risk description, the modification suggestion, and the location of the review item in the contract; the evaluation result is obtained based on at least one of the following: a first score characterizing the recognition quality of the large model, a second score characterizing the output quality of the risk description and modification suggestion of the large model, a third score characterizing the degree of distance difference between the location output by the large model and the location of the same review item in the manually labeled benchmark set, the review position compliance rate of the large model in conducting biased review of the preset position, or the legal citation accuracy rate characterizing the reliability of the laws and regulations cited by the large model. The first score, the second score, the third score, the legal citation accuracy rate, and the review stance compliance rate are obtained based on the review results of the contract after reviewing it using the large model.
4. The method according to claim 3, characterized in that, The process of obtaining the first score includes: The risk level and location of the large model review item name mapped to the same review item in the review results are compared with the corresponding risk level and location in the manually labeled benchmark set; wherein, the comparison result is consistent, indicating that the corresponding review item is correctly identified by the large model; Based on the comparison results, a first ratio is calculated between the number of correctly identified large models and the total number of correctly identified large models, and a second ratio is calculated between the number of correctly identified large models and the total number of annotations in the manually annotated benchmark set. The first score is calculated based on the first ratio and the second ratio.
5. The method according to claim 3, characterized in that, The process of obtaining the second score includes: The risk descriptions and modification suggestions from the manually labeled benchmark set, as well as the risk descriptions and modification suggestions for the same review item from the review results, are input again into the large model, and a second score is output to characterize the output quality of the risk descriptions and modification suggestions from the large model.
6. The method according to claim 3, characterized in that, The process of obtaining the third score includes: The positioning positions in the manual annotation benchmark set are compared with the positioning positions for the same review item in the review results to obtain the positioning distance between the two for the first character and the positioning intersection-union ratio between the two for the positioning range. The third score is obtained based on the positioning distance and the positioning intersection-union ratio.
7. The method according to claim 3, characterized in that, The process of obtaining the review stance compliance rate includes: Obtain stance judgment prompts, input the review results and the stance judgment prompts back into the large model, and obtain a judgment result that characterizes whether the review results conform to the preset stance; Based on the judgment result, the percentage of review items that conform to the preset position is obtained as the review position compliance rate.
8. A device for evaluating contract review results, characterized in that, The device includes: The acquisition module acquires a set of manual annotation benchmarks for annotating contracts. The set of manual annotation benchmarks is obtained by annotating and verifying different types of de-identified contracts based on a rule document used to guide annotation. The rule document includes the name of the review item, risk type, risk level, risk definition, and location rules for locating review items in the contract. The mapping module is used to map the names of manual review items in the manual annotation benchmark set to the names of large model review items output by the large model used for contract review, and obtain a mapping relationship table between the names of large model review items and manual review items. The evaluation module is used to obtain the review results of the contract after the large model reviews it, and compares the review results mapped to the same review item with the annotation results in the manual annotation benchmark set to obtain the evaluation results of the review quality of the large model.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.