Contract review method and related product
By performing sharding and multi-dimensional search on contracts, combining feature labels and contract review models, the problems of high costs and low accuracy in the existing technology are solved, and efficient and accurate contract review is achieved.
Patent Information
- Application Number
- CN202510377678.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-01
AI Technical Summary
The existing technology has high-cost, long-term customized training needs in contract review, and the vectorized search methods are interfered with by external factors, resulting in insufficient accuracy and reliability of the review results.
By sharding the contract to be reviewed and assigning feature tags to each shard, combining multiple search algorithms and contract review models, multi-dimensional search and review are carried out to generate accurate review results.
It improves the efficiency and accuracy of contract review, reduces the dependence on manual annotation data and expertise, and enhances scenario adaptability and accuracy of review results.
Smart Images

Figure CN120235582A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical field of digital data processing, and in particular, to contract review methods and related products. Background Art
[0002] In the field of contract review, with the rapid development of artificial intelligence technology, automated contract review has become an important means to improve work efficiency and reduce legal risks. Currently, after training a natural language processing (NLP) for specific tasks on a large number of contract texts, the trained model can be used to let the computer automatically identify and judge the risk points in the contract. In addition, the contract text can be vectorized into a knowledge base, and according to the review rules input by the user, relevant paragraphs can be retrieved in the knowledge base, and then the text paragraphs can be input into a large language model (LLM) in combination with specific prompt words to return the review results.
[0003] However, the method of using NLP for contract risk identification needs to be customized for different contract types and industry scenarios. It not only requires a large amount of high-quality labeled data, but also requires professional legal knowledge to guide model training, resulting in problems such as high cost, long cycle, and insufficient generalization ability; while the vectorized retrieval method can reduce the dependence on manual annotation, but its retrieval quality will be affected by external factors such as unclear user question expressions and cross-domain expressions, and there are also internal technical limitations such as information loss during the vectorization process and insufficient semantic retrieval accuracy, thus affecting the accuracy and reliability of the final review results. Summary of the Invention
[0004] Based on the above problems, embodiments of the present application provide contract review methods and related products, aiming to improve the efficiency and accuracy of contract review, while reducing the dependence on manually labeled data and professional knowledge.
[0005] In a first aspect, embodiments of the present application provide a contract review method, including:
[0006] Perform sharding processing on the contract to be reviewed, and assign at least one first feature label to each shard;
[0007] Determine the target review plan corresponding to the contract to be reviewed; the target review plan includes a plurality of review items;
[0008] Based on each shard and its first feature label in the contract to be reviewed, use a variety of retrieval algorithms to respectively retrieve the contract to be reviewed, and obtain a set of data to be reviewed corresponding to each review item; wherein, the variety of retrieval algorithms are retrieval algorithms based on different dimensions;
[0009] Input the set of data to be reviewed corresponding to each of the review items into the contract review model to obtain the review results corresponding to each of the review items.
[0010] Second, an embodiment of the present application further provides a contract review device, including:
[0011] A processing unit, configured to perform sharding processing on the contract to be reviewed and assign at least one first feature tag to each shard;
[0012] A determination unit, configured to determine a target review plan corresponding to the contract to be reviewed; the target review plan includes a plurality of review items;
[0013] A multi-dimensional retrieval unit, configured to respectively perform retrieval on the contract to be reviewed by using a variety of retrieval algorithms based on each shard and its first feature tag in the contract to be reviewed, so as to obtain a set of data to be reviewed corresponding to each of the review items; wherein, the variety of retrieval algorithms are respectively retrieval algorithms based on different dimensions;
[0014] A review result generation unit, configured to input the set of data to be reviewed corresponding to each of the review items into the contract review model to obtain the review results corresponding to each of the review items.
[0015] Third, an embodiment of the present application further provides a computer device, including:
[0016] A central processing unit, a memory, and an input / output interface;
[0017] The memory is a transient storage memory or a persistent storage memory;
[0018] The central processing unit is configured to communicate with the memory and execute the instruction operations in the memory to execute the contract review method described in any one of the above.
[0019] Fourth, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the contract review method described in any one of the above is executed.
[0020] Fifth, an embodiment of the present application further provides a computer program product, which includes computer instructions. When the computer instructions are executed by a processor, the contract review method described in any one of the above is implemented.
[0021] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0022] In the embodiments of the present application, by splitting the contract to be reviewed and allocating at least one first feature tag to each split, the entire content of the contract to be reviewed can be split into smaller structured units, and the accuracy and speed of subsequent retrieval can be improved through the first feature tag. By determining the target review scheme corresponding to the contract to be reviewed, the review items can be flexibly adjusted according to different contract types, improving applicability. Combining various retrieval algorithms with different dimensions can avoid missing key data to be reviewed and improve the accuracy of contract content matching. Since the content input into the contract review model has been screened and only contains the set of data to be reviewed related to the review items, the misjudgment of the contract review model can be reduced ultimately, and the accuracy and quality of the review results can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0024] Figure 1 Schematic diagram of the system architecture of a contract review method provided by the embodiments of the present application;
[0025] Figure 2 Schematic diagram of the flow of a contract review method provided by the embodiments of the present application;
[0026] Figure 3 Schematic diagram of the review result display interface provided by the embodiments of the present application;
[0027] Figure 4 Another schematic diagram of the flow of a contract review method provided by the embodiments of the present application;
[0028] Figure 5 Schematic diagram of the structure of a contract review device provided by the embodiments of the present application;
[0029] Figure 6 Schematic diagram of the structure of a computer device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0031] The terms "first", "second", "third", "fourth", etc. (if any) in the description, claims and drawings of this application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0032] In the following description, expressions such as "a specific embodiment" or "a specific example" are involved, which describe a subset of all possible embodiments. However, it can be understood that "a specific embodiment" or "a specific example" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict. In the following description, the term "plurality" refers to at least two. When it is said in this application that a certain value reaches a threshold (if any), in some specific examples, it may include the case where the former is greater than the latter. If terms such as "any" or "at least one" are mentioned, it may specifically refer to any one of the listed examples or any combination between these examples.
[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0034] The method provided by the embodiments of this application can be applied to a system architecture as Figure 1 shown. Among them, the terminal 102 communicates with the server 101 through the network, and the data storage system 100 can store the data that the server 101 needs to process or requires. The data storage system 100 can be integrated on the server 101, or can be placed in the cloud or on other network servers. The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets and portable wearable devices, and the portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The server 101 can be implemented by an independent server or a server cluster composed of multiple servers.
[0035] The terminal 102 can perform sharding processing on the contract to be reviewed, assign at least one first feature tag to each shard, and send each shard of the contract to be reviewed and its first feature tag to the server 101. The server 101 determines the corresponding target review scheme based on the received contract to be reviewed; the server 101 can also retrieve the contract to be reviewed respectively using a variety of retrieval algorithms based on each shard and its first feature tag in the contract to be reviewed, and obtain a set of data to be reviewed corresponding to each review item; in addition, the server 101 can also input the set of data to be reviewed corresponding to each review item into the contract review model, obtain the review results corresponding to each review item and return them to the terminal 102 for display. Through the combination of multi-dimensional retrieval algorithms and the contract review model, not only the accuracy and efficiency of contract review are improved, but also the dependence on large-scale manually labeled data and professional knowledge is effectively reduced. At the same time, it has good scene adaptation ability and can flexibly handle different types of contracts and diverse industry review requirements.
[0036] It should be noted that the method provided in the embodiments of the present application can be jointly implemented by the terminal device and the server as described above, or can be entirely implemented on the server side, or can also be entirely implemented on the terminal device side. It can be specifically determined according to the actual application scenario and is not limited here.
[0037] The following will elaborate on two ways to review contracts in the prior art. One is the method of risk identification based on NLP, that is, first training a large number of contract texts for specific tasks, such as training a model to identify key contract contents or potential risk points such as "payment terms", "liability for breach of contract", "jurisdiction", etc.; then, using the trained model to let the computer automatically identify and judge the risk points in the contract, so as to assist or partially replace manual review. This method highly depends on a large amount of labeled data and professional legal knowledge, resulting in high training costs and long cycles, and it is difficult to quickly adapt to the review requirements of different types of contracts. The other is the method of combining vector retrieval with large model review, that is, converting the contract content into a "vector" form that can be understood by the computer and storing it in a knowledge base. When receiving some review requirements or key points set by the user (such as "whether it contains a breach clause", "whether the payment method is clear", etc.), the system will search for contract paragraphs related to the review rules or key points input by the user in this knowledge base. After finding the relevant contract content, a set of designed prompts will be used to organize these contract paragraphs and input them into a large-scale pre-trained model, so that the large-scale pre-trained model can judge whether these contract paragraphs meet the review requirements or key points input by the user, and then output a review conclusion. Compared with the NLP-based solution, vector retrieval reduces the dependence on manual annotation and improves the degree of automation of review. However, the effect of the vector retrieval method is easily affected by external factors and internal factors, and there is a certain degree of uncertainty. Specific external factors include: different user query methods (such as the review requirements or key points input by the user are expressed vaguely), the user's question exceeds the scope of the knowledge base, etc.; internal factors include: some information may be lost during the vectorization process, and the accuracy of semantic search is limited by the performance of the vector retrieval algorithm, resulting in deviations in the review results.
[0038] To solve these problems in the prior art, the embodiments of the present application use a hybrid retrieval + model review method to review the risk points of contract texts, and thus can obtain contract review results more quickly and accurately, improving the efficiency and accuracy of contract review.
[0039] The following will further elaborate on the method of the present application and provide some specific possible implementation examples. In actual applications, the implementation contents between these examples can be combined or implemented separately according to the corresponding functional principles and application logics as needed. If combined, the execution order between the combined examples can be determined according to their respective processing logics, and can be specifically determined by the actual scenario.
[0040] Please refer to Figure 2 , the embodiments of the present application provide a contract review method, which includes steps S201 - S204.
[0041] S201: Segment the contract to be reviewed and assign at least one first feature label to each segment.
[0042] In the contract review process, it is first necessary to segment the contract to be reviewed and assign at least one first feature label to each segment. Segmenting the contract can split the contract into finer-grained segments, facilitating subsequent precise analysis and review.
[0043] Before segmenting, the basic data settings can be completed first to provide a structured reference basis for the contract review model engine. The basic data can include review items, review schemes, feature labels, etc.
[0044] The review scheme includes at least the following attributes:
[0045] Attribute Interpretation Scheme Name Unique Identifier of the Review Scheme Enabling Conditions Define When to Apply this Review Scheme Applicable Contract Types Such as Lease Contracts, Purchase Contracts, etc. Review Item Entries Include Review Items Associated with this Review Scheme
[0046] Each review item consists of at least the following elements:
[0047]
[0048] Feature labels are used to mark professional terms appearing in the contract, facilitating contract review identification and classification. Common labels include: "Receivable amount", "Contract parties", "Performance details", etc.
[0049] After setting the basic data, the obtained contract text to be reviewed can also be parsed first. Exemplarily, document processing tools such as Apache POI can be used to parse the text content from contracts to be reviewed in formats such as doc, docx, and pdf, remove irrelevant information or interfering text such as headers, footers, and annotations, and extract key text or table information, and then convert the original content into a format that is easy for the contract review model to understand and process (for example, plain text or structured data). It should be noted that the contract to be reviewed can be a contract in electronic format such as doc, docx, pdf, etc., or a collection of pdf files or picture files generated after scanning a paper contract.
[0050] After completing the contract text parsing, the contract to be reviewed can be segmented according to the preset segmentation rules, and the complete contract text can be split into multiple independent segments. Each segment represents a specific part of the contract, facilitating subsequent retrieval and review. Exemplarily, the preset segmentation rules include the following segmentation methods:
[0051] 1. Divide according to the chapter or clause numbers in the contract, for example: 1.1, 1.2, etc.
[0052] 2. Divide the text content of the contract to be reviewed according to its semantic structure, for example, dividing the content involving different topics such as payment, confidentiality, breach of contract, etc. into independent segments.
[0053] 3. Split the contract content to be reviewed into segments according to at least one of natural paragraphs, sentences or preset identifiers. The preset identifiers here can be specific symbols in the contract to be reviewed, such as semicolons, periods, line breaks, etc.
[0054] When reviewing contract segments, the large language model LLM can also be used to assign at least one first feature tag to each segment. Here, the first feature tag refers to a tag that semantically classifies or summarizes the content of each segment, which is mainly used for subsequent review and location retrieval operations. The source of the first feature tag can be the feature tag set during the early basic data setting.
[0055] By splitting the contract to be reviewed into multiple shards, the contract terms can be reviewed more accurately to avoid misjudgment of the entire text. Subsequent retrieval only needs to be performed in the shards related to the review rules, thereby greatly improving the efficiency of the contract review model.
[0056] S202: Determine a target review plan corresponding to the contract to be reviewed; the target review plan includes multiple review items.
[0057] The contracts to be reviewed mentioned in the embodiments of this application may include sales contracts, construction contracts, house rental contracts, purchase contracts, etc. It is understandable that contracts in different fields may have different review items. For example, sales contracts may focus more on the review of subject matter description, delivery time, delivery method, quality acceptance, breach of contract liability, etc.; house rental contracts may focus more on the review of lease subject matter, rental amount, payment method, lease term, maintenance responsibility, etc.
[0058] The target review scheme can be determined based on the user's input instructions or based on the contract title, keywords or semantic analysis to select the most matching target review scheme. Furthermore, when adding a new contract type, it is only necessary to add a review scheme corresponding to the new contract type without modifying the underlying logic. Therefore, by applying the target review scheme corresponding to the contract to be reviewed, the efficiency and accuracy of the contract review can be improved and the manual workload can be reduced.
[0059] It should be noted that the execution order of step S201 and step S202 is not limited and they may be executed simultaneously, depending on the specific situation.
[0060] S203: Based on each shard in the contract to be reviewed and its first characteristic tag, the contracts to be reviewed are searched respectively using a plurality of search algorithms to obtain a set of data to be reviewed corresponding to each review item;
[0061] Among them, the multiple retrieval algorithms are retrieval algorithms based on different dimensions, that is, the multiple retrieval algorithms are different from each other in retrieval logic and / or matching mechanism. Exemplarily, the first retrieval algorithm may be a retrieval algorithm based on keyword matching, the second retrieval algorithm may be a retrieval algorithm based on semantic similarity calculation, and the third retrieval algorithm may be a retrieval algorithm for recognition based on a rule engine (such as regular expressions, sentence templates).
[0062] In the embodiment of the present application, by combining multiple different types of retrieval algorithms, the contract content to be reviewed is retrieved and analyzed from multiple dimensions, effectively avoiding the problem of missing contract terms that may be caused by a single retrieval method. For example, when only keyword matching is used, if the contract terms are expressed using synonyms or different language expressions, it may lead to inaccurate hitting of the target content. By introducing retrieval methods from other dimensions, the content can be further identified and extracted from the semantic level and the structural level, realizing the accurate matching of the contract content to be reviewed under different expression methods, thereby extracting the contract fragments corresponding to each review item and finally forming a complete data set to be reviewed, improving the comprehensiveness and accuracy of the review.
[0063] S204: Input the data set to be reviewed corresponding to each review item into the contract review model to obtain the review results corresponding to each review item.
[0064] Exemplarily, assume that a review item is "whether the liability for breach of contract clause is complete". Then the data set to be reviewed may include all relevant clauses and their contexts in the contract related to liability for breach of contract. After inputting it into the contract review model, the contract review model can judge whether the clause meets the standard through rule matching, semantic understanding or machine learning methods.
[0065] Please refer to Figure 3 , Figure 3 which is a schematic diagram of a review result display interface provided by the embodiment of the present application. As Figure 3 shown, when the review is completed, the review results of the contract to be reviewed can be displayed, and the review results can also be divided into risk levels, such as high risk, medium risk and low risk. Specifically, reference can be made to Figure 3 the red box and the green box in the intelligent review module on the right side of the contract. For example, the review result under a certain review item is displayed in the red box, indicating that the contract clause corresponding to the review item is a high-risk clause; the orange box and the yellow box respectively represent medium-risk clauses and low-risk clauses, and the review results displayed in the green box are the contract clauses of the review items that have passed the review. The review results can be displayed from the aspects of risk description and modification suggestions. For the review results of each review item, the risk description and modification suggestions can be displayed by clicking on the corresponding review item control.
[0066] In summary, in the embodiments of the present application, by performing sharding processing on the contract to be reviewed and allocating at least one first feature tag to each shard, the entire content of the contract to be reviewed can be split into smaller structured units, and the accuracy and speed of subsequent retrieval can be improved through the first feature tag. By determining the target review plan corresponding to the contract to be reviewed, the review items can be flexibly adjusted according to different contract types, improving applicability. Combining different-dimensional retrieval algorithms can avoid missing the data to be reviewed related to the review items and improve the accuracy of contract content matching. Since the content input into the contract review model has been screened and only contains the data set to be reviewed related to the review items, the misjudgment of the contract review model can be finally reduced, and the accuracy and quality of the review results can be improved.
[0067] To ensure that the data set to be reviewed for contract review has high accuracy and integrity, thereby guaranteeing the reliability of the subsequent output review results, based on Figure 2 the example content of, in some specific examples, step S203 of the embodiments of the present application specifically includes: based on each shard of the contract to be reviewed and its first feature tag, using a variety of retrieval algorithms to respectively retrieve the contract to be reviewed, and obtaining a data set to be reviewed corresponding to each review item, including: based on each shard of the contract to be reviewed and its first feature tag, using a variety of retrieval algorithms to respectively retrieve the contract to be reviewed, and obtaining multiple retrieval sets corresponding to different retrieval algorithms; performing a merging and optimization process on the multiple retrieval sets to obtain a data set to be reviewed corresponding to each review item.
[0068] As mentioned in the foregoing embodiments, by using a variety of retrieval algorithms, the retrieval blind spots caused by different expression forms, semantic variants or structural differences in the contract terms can be covered, thereby improving the accuracy and coverage of clause recognition.
[0069] Taking into account that the search sets retrieved by multiple search algorithms may have some duplicate content, a merge optimization process is required in the embodiment of the present application. The merge optimization process here refers to the fusion, screening, and enhancement of the search results (i.e., multiple search sets) respectively produced by multiple search algorithms, and finally generating a more accurate and reliable set of data to be reviewed. Specifically, the merge optimization process can optimize the strategies such as deduplication, confidence scoring, intersection / union analysis, etc. of the results output by different search algorithms, so as to improve the overall matching quality and review efficiency of the data set to be reviewed. For example, for similar content fragments that are hit in multiple search sets, results with higher semantic consistency or stronger correlation with the review items can be retained first; for conflicting or duplicated content fragments, they can be screened based on contextual information. A weighted score is given to each hit content fragment based on its matching method (such as keyword matching, semantic similarity matching, rule template matching, etc.), the degree of association with the review item, contextual semantics and other factors to form a confidence or relevance index, and multiple search sets under the same review item are sorted accordingly, retaining only the top K items or content with a relevance higher than a preset threshold. The K value and preset threshold here can be configured on demand according to the actual review requirements of different review items, and the embodiments of the present application do not limit this.
[0070] Through the above-mentioned merge optimization processing, the complementary advantages of different retrieval algorithms can be fully utilized. While ensuring the recall rate of relevant clauses in the contract review process, the matching accuracy can be improved and the interference of irrelevant content can be reduced, thereby improving the intelligence and reliability of the contract review process.
[0071] In order to improve the adaptability of contract review in diverse expressions and complex semantic environments and achieve the effect of extracting data to be reviewed with both high accuracy and high recall, a new contract review algorithm based on the CNN is proposed. Figure 2 In some specific examples, the target review scheme also includes review rules corresponding to each review item; before step S203, the method of the embodiment of the present application may also include: using a large language model to generate keywords, second feature tags and text vectors corresponding to each review rule;
[0072] In the embodiments of the present application, multiple retrieval algorithms include a first retrieval algorithm based on keywords, a second retrieval algorithm based on feature tags, and a third retrieval algorithm based on text vectors. Then, in the embodiments of the present application, the "using multiple retrieval algorithms to separately retrieve the contract under review based on each slice and its first feature tag in the contract under review to obtain multiple retrieval sets corresponding to different retrieval algorithms" in step S203 specifically includes: using the first retrieval algorithm to perform keyword retrieval on each slice based on the keywords of each review rule to obtain a first retrieval set; using the second retrieval algorithm to retrieve each slice based on the second feature tag of each review rule and the first feature tag of each slice to obtain a second retrieval set; using the third retrieval algorithm to perform vectorized retrieval on each slice based on the text vector of each review rule to obtain a third retrieval set.
[0073] After determining the target review plan corresponding to the contract under review, the large language model LLM can be used to generate the keywords of each review rule, the second feature tags of each review rule, and the text vectors of each review rule. Among them, the keywords can directly take the original core words of the review rules, such as "payment", "default", "term", etc.; the second feature tags are used to assist in identifying the categories of review rules, such as "payment terms", "contract term", etc. The text vector obtained by converting the text information into vector form can provide support for semantic search, enabling similar terms to be identified even if different contracts use different wordings.
[0074] In the embodiments of the present application, after obtaining the retrieval information, for each review item, the contract under review is separately retrieved through the first retrieval algorithm, the second retrieval algorithm, and the third retrieval algorithm. For example, the first retrieval algorithm retrieves the contract under review based on the Elastic search search engine to obtain a first retrieval set; exemplarily, if the review rule requires checking "payment terms", then the system will search for text slices in the contract that contain keywords such as "payment", "amount", "transfer", etc. The second retrieval algorithm obtains a second retrieval set by retrieving the matching degree between the second feature tag of each review rule and the first feature tag of each slice; exemplarily, the first feature tag of a certain payment term a belongs to the financial term category, and the second feature tag generated by the LLM for the review rule b under a certain review item based on the existing feature tag library is also the financial term category. Therefore, this payment term a can be matched with the review rule b, and subsequently, the review rule b will be used to review the payment term a. The third retrieval method retrieves the contract under review through a vectorized retrieval method to obtain a third retrieval set. For example, although a certain contract does not clearly mention the specific payment date, but uses the wording "shall be paid within 30 days after the contract is signed", the semantic similarity can be identified through the vector retrieval of the third retrieval method.
[0075] Through the above embodiments, the embodiments of the present application can effectively solve the problems of inconsistent contract expression methods and difficult identification of key terms. By leveraging the capabilities of large language models in vectorization, tag generation, and keyword expansion, combining the complementary advantages of various retrieval algorithms with different dimensions, and through merging and optimization processing, not only can accurate matching of target review items be achieved, but also the clause recall ability in complex contract texts can be significantly improved, thereby comprehensively enhancing the accuracy and reliability of the contract review model.
[0076] Optionally, in order to prevent text semantic truncation caused by a large span of contract content expression or overly fine fragmentation granularity, and to ensure both precision and comprehensiveness during contract review, based on Figure 2 the example content, in some specific examples, "merging and optimizing multiple retrieval sets to obtain a set of data to be reviewed corresponding to each review item" in step S203 of the embodiments of the present application specifically includes: merging and removing duplicates from multiple retrieval sets to obtain the target paragraphs corresponding to each review item in the contract to be reviewed; for the target paragraphs corresponding to each review item, using various retrieval algorithms to retrieve the context of the target paragraphs to obtain associated context paragraphs; combining the target paragraphs corresponding to each review item and their associated context paragraphs to obtain a set of data to be reviewed corresponding to each review item.
[0077] Exemplarily, keyword retrieval found paragraph d of "payment terms", feature tag retrieval also found the same paragraph d, and text vector retrieval found a paragraph p with different wording but similar semantics. Since paragraph d is repeated among these results, duplicate removal is required, and finally the target paragraphs are determined to be paragraph d and paragraph p.
[0078] However, reviewing a single target paragraph d or target paragraph p alone may not be sufficient to accurately judge compliance. The complete semantics of some clauses may span multiple paragraphs or even appendix content and need to be understood in combination with the context. Therefore, in order to improve the accuracy of contract review, for each target paragraph, various retrieval algorithms such as keyword retrieval, feature tag retrieval, and text vector retrieval are further used to expand the retrieval of its context to obtain context paragraphs that are semantically or structurally related to it. The scope of context retrieval can include several adjacent fragments before and after the target paragraph, other content within the same clause, or content that is logically in a reference relationship with the target paragraph, such as contract appendices or reference clauses. The setting of the context scope can be configured according to specific business scenarios, and the present application does not limit this. Similarly, for the context part of the target paragraph, the results separately hit by various retrieval algorithms can also be de-duplicated to ensure the accuracy and uniqueness of the final context content.
[0079] Specifically, please understand the above semantic truncation problem in combination with the following specific examples: For example, in the payment terms, the main text mentions that "the balance shall be paid after acceptance", but the specific payment method is listed in the appendix. Without considering the context, it may lead to inaccurate review conclusions. Or, there is the following content in the contract: "The project is delivered in two phases: the first delivery time is June 30, 2024, and the second delivery time is December 31, 2024." If the system processes each sentence as a slice and is split into slice 1: "The project is delivered in two phases: the first delivery time is June 30, 2024." Slice 2: "The second delivery time is December 31, 2024." At this time, if a review item needs to identify whether the "delivery plan is complete" and only slice 1 is hit, it may be impossible to know that there is also a "second phase", resulting in missing key information. Therefore, to avoid the above situation, it is necessary to further expand the context on the basis of determining the target paragraph. In this way, during the review, while ensuring information accuracy, the semantic context of the terms can be effectively retained, avoiding misinterpretation caused by the isolated analysis of a single paragraph, preventing text semantic truncation, and further enhancing the robustness and reliability of contract review at the levels of semantic restoration and rule matching, and improving the accuracy of contract review.
[0080] Based on Figure 2 the example content, in some specific examples, to prevent retrieval omissions, the target review scheme further includes review rules corresponding to each review item; each review rule generates a corresponding retrieval rule through a large language model; the retrieval rule is used to locate the set of data to be reviewed corresponding to the review item; after step S203, the method of the embodiment of the present application may further include: performing similarity detection and / or hit detection on the retrieval rule and the set of data to be reviewed under the same review item to obtain a detection result; if the detection result indicates that the similarity between the review rule and the set of data to be reviewed under the same review item is lower than a preset threshold, and / or, the detection result indicates that the set of data to be reviewed that matches the review rule is not hit, then use the large language model to perform a full-text search on the contract to be reviewed according to the retrieval rule to obtain a new set of data to be reviewed.
[0081] The review rule defines the core content that needs to be concerned during contract review. For example, "the payment terms must clearly define the payment cycle" is a review rule. The retrieval rule is a specific retrieval strategy generated by the LLM according to the review rule, which is used to guide the specific way to retrieve relevant terms. For example, for a certain review rule, the generated retrieval rule may include: specifying a vocabulary set for keyword matching (such as "payment", "pay", "transfer", etc.), specifying a label for feature tag matching (such as the feature tag is "payment terms"), or specifying a sentence-level expression for semantic vector retrieval (such as "identifying semantic paragraphs related to payment obligations"), etc.
[0082] After completing the multi-dimensional search in S203, for the purpose of quality control of the data set to be reviewed, it is necessary to evaluate the completeness and accuracy of the data set to be reviewed that is currently retrieved. Therefore, it is necessary to perform similarity detection and / or hit detection on the search rules and the data set to be reviewed under the same review item. For example, the semantic similarity between the search rule (such as its vector representation) and the data set to be reviewed can be calculated; for example, if a certain review rule is intended to identify content related to "payment method", and the contract fragment currently retrieved is more inclined to "liquidated damages" related content, then the similarity between the two may be lower than the preset threshold, indicating that the current data set does not meet the review intention. In the hit detection, it can be determined whether the data set to be reviewed has the key content features specified by the search rule, such as whether the keywords appear. Taking the review of "payment terms" as an example, if the search rule requires hitting keywords such as "payment", "transfer", and "payment", and no relevant words appear in the current data set, it is considered that the search rule is not hit.
[0083] When the similarity represented by the above detection results is lower than the preset threshold and / or the search rules are not hit, the embodiment of the present application can trigger a backup strategy: calling a large language model to perform a full-text semantic search on the entire contract. During this retrieval process, LLM can not only perform high-level semantic matching within the entire contract, but also adjust keywords and expand the scope of semantic matching to obtain more relevant contract content. For example, "payment method" is expanded to expressions such as "payment arrangement", "transfer rules", and "fund payment path" to obtain content fragments that are more consistent with the semantics of the review item, and finally form a new and more matching set of data to be reviewed.
[0084] In the embodiment of the present application, step S203 adopts a mixed retrieval method of multiple retrieval algorithms, which can quickly find the data to be reviewed corresponding to the review item and achieve efficient preliminary positioning. Subsequently, a matching detection mechanism (including similarity detection and hit detection) between the retrieval rules and the data set to be reviewed is combined with the LLM-driven full-text supplementary inspection mechanism to construct a set of multi-layer protection strategies. In this way, a supplementary search can be performed when there is a risk of omission, effectively reducing the probability of omissions in the contract review process, thereby improving the comprehensiveness and accuracy of the review results and improving the quality of contract review.
[0085] based on Figure 2 According to the example content, in some specific examples, step S204 of the embodiment of the present application specifically includes: inputting the data set to be reviewed corresponding to each review item into the contract review model corresponding to the review item, and obtaining the review result of each review item; wherein the contract review model includes a contract review model in the form of a plug-in and a contract review model in the form of a large language model (LLM).
[0086] In the embodiments of the present application, a corresponding contract review model form can be set for each review item in the target review plan. The contract review model can adopt an LLM model or a plug-in rule model. If the LLM model form is used, to ensure the security of data processing, an open-source LLM model can be selected for the contract review model. For an open-source LLM, the advantage is that it can be deployed in a local or private cloud environment, so that there is no need to worry about uploading internal data to an external server or a third-party interface, avoiding the risk of data leakage and ensuring the security of internal sensitive data. The open-source LLM already has strong semantic understanding capabilities on a large amount of diverse pre-trained data, enabling it to have good clause recognition and semantic review capabilities, directly execute review tasks on the dataset to be reviewed, and output contract review results with relatively high accuracy.
[0087] In practical applications, the LLM can be guided to execute review tasks through the Prompt template set in the target review plan. This Prompt template can be dynamically filled in combination with the contract type and review rules to form an injection instruction for a specific review item, thereby driving the LLM to perform semantic review on the input dataset to be reviewed and output the corresponding review judgment results.
[0088] In addition, it should be noted that when using the LLM for contract review, the limitation of its text processing length needs to be considered. The text length that the LLM can process is limited. If the full text of an overly long contract is directly input into the LLM for review, it may exceed its processing capacity, resulting in errors or performance degradation of the contract review model. Therefore, in the embodiments of the present application, the contract to be reviewed can be reasonably split according to preset rules to ensure the stability of LLM processing and review performance.
[0089] Please refer to Figure 4 , the text division rule can be based on the maximum processing length of the LLM, that is, the token length limit of the LLM to divide the text. For example, if the maximum processing length of a certain LLM is 12K words, in a feasible implementation, the length of a single piece of text can be set not to exceed 12K words. Further, considering the attention mechanism of the LLM, when processing overly long text, problems such as information loss and focus deviation may occur. Therefore, in practical applications, an optimized division length standard can be reset on the basis of the maximum processing length of the LLM, such as 8K words, to ensure that the LLM can more accurately focus on the core content of contract compliance review when processing shards. After the initial division in units of 8K, the content of the contract to be reviewed can also be structurally sharded in combination with the preset sharding rules in the foregoing step S201 to improve the LLM's understanding ability and review effect at the shard level. The specific preset sharding rules and division basis can refer to the content of the foregoing S201 and will not be elaborated here.
[0090] The various contract text segmentation methods provided in the embodiments of this application can not only avoid model performance fluctuations caused by overlong input, but also improve the model's ability to focus attention, thereby enhancing the accuracy and stability of the review results. Among them, the specific segment length and segmentation rules can be flexibly configured according to different application scenarios, and this application does not limit this.
[0091] In addition to the above-mentioned LLM-style contract review model, the contract review model can also exist in the form of a plug-in. The plug-in-style contract review model is mainly used to handle tasks such as red bar review, subject review, sensitive word review, and template review. The plug-in-style contract review model is suitable for preliminary screening of common content with clear rules and not easy to change, and can be deployed through application programming interfaces (APIs), cloud services, internal office system integration, local file editor plug-ins, etc.
[0092] For example, for the review of the main information in the contract, such as the company name, address, contact information, legal person information, etc., a contract review model in the form of a plug-in can be set up in the target review plan. When conducting the plug-in review, the verification and comparison of the main information can be completed by connecting to an external enterprise database or a third-party information interface.
[0093] Red line review (also known as "red line clause review") is used to identify high-risk sensitive clauses in contracts. The review rules are usually preset by the corporate legal or risk control team. For example, "non-refundable deposit" and "unlimited liability for breach of contract" are typical red line clauses. If they appear in the contract, the plug-in model can automatically identify and mark them as high-risk content for subsequent key review and processing.
[0094] Both red bar review and sensitive word review can detect sensitive terms and clauses in the entire contract text through keyword matching and natural language processing (NLP) semantic analysis, and identify inappropriate expressions or clauses that may trigger legal risks.
[0095] Template comparison refers to comparing the contract to be reviewed with the standard contract template preset within the enterprise, and automatically identifying the differences in additions, deletions, and changes. For example, if the current contract adds a new payment method that is unfavorable to us in the "payment terms" section or deletes the payment protection clause in the original template, the plug-in contract review model can mark these changes and identify the risks of inconsistency with the template standard.
[0096] The embodiment of the present application integrates the strong semantic understanding ability of the LLM model and the efficient rule review ability of the plug-in model. It can not only handle complex semantic review tasks, but also quickly complete the efficient screening of structured clauses, meet the needs of automated review in various business scenarios, while also taking into account the review quality and data security, and significantly improving the efficiency of contract compliance review.
[0097] Based on Figure 2 Based on the example content, in some specific examples, after step S201, the method of the embodiment of the present application may further include: desensitizing the target fields in the contract to be reviewed based on a preset desensitization rule, and generating a data desensitization record; the data desensitization record is used for the mapping relationship between the target field and the original value of the target field; after step 204, the method of the embodiment of the present application may further include: obtaining the target field to be restored according to the data desensitization record; performing an original value restoration process on the target field according to the data desensitization record, so that the original value of the target field is used as the field content indicated by the target field.
[0098] Desensitization refers to encrypting, replacing, or masking sensitive information in a contract to prevent unauthorized access or data leakage. During the contract review process, it is necessary to prevent unauthorized reviewers from seeing sensitive information, but still be able to conduct review and analysis. Therefore, the fields that need to be desensitized in the contract can be automatically identified through a preset desensitization rule, such as personal information (name, ID number, contact information), company information (bank account number, tax number), etc.
[0099] The desensitization can be achieved through character replacement. For example, the ID number 123456789012345678 → 1234****5678; or it can also be hashed and encrypted, such as ABCD → 7f6a8c9d; or a data identifier can be used: replace the original content with MASK_ID_001. In this way, during the contract review process, the LLM does not need to know the real sensitive information when reviewing the contract text, and can still complete the normal contract review process, while also ensuring data security.
[0100] To ensure traceability, each desensitized field corresponds to a desensitization record, which is used to store the mapping relationship between the desensitized value and the original value. For example: MASK_ID_001 → "Zhang San"; MASK_ID_002 → "(company account) 12345678", so that the original data can be safely restored later when presenting the review results. When restoring the desensitized information, it is necessary to first obtain the field to be restored (such as "name of Party B") and its corresponding desensitized value (such as MASK_ID_001) in the desensitization record, and then replace MASK_ID_001 with the original value (such as "Zhang San"), and finally output the complete contract review result.
[0101] The embodiment of the present application can prevent unauthorized contract reviewers from obtaining sensitive information through the method of desensitizing before review and restoring after review. The desensitized contract can still be parsed and processed by the LLM, making the contract review both safe and efficient, meeting the enterprise compliance requirements, and at the same time not affecting the integrity of the contract and the accuracy of the review.
[0102] Based onFigure 2 In some specific examples of the example content, to enable users to intuitively view the review results, after step S204, the method of the embodiment of the present application may further include: according to each data set to be reviewed, using a similarity algorithm to retrieve in the contract to be reviewed, obtaining the source text corresponding to the data set to be reviewed and the coordinate range of the area where the source text is located; according to the coordinate range of the area where the source text is located, positioning and marking the source text of the contract to be reviewed, and displaying the review results corresponding to the marked source text on the review result display interface.
[0103] It should be noted that although in step S204, the LLM has output the review results corresponding to each review item and the data sets to be reviewed that match them, due to the possible "hallucination" problem in the generation process of the LLM, the content it returns may have slight differences in expression from the original contract, such as omitting some words, changing punctuation, or reorganizing the content. To ensure that the final display results accurately reflect the true content of the contract, in the embodiment of the present application, a similarity algorithm (such as semantic matching based on text vectors, fuzzy matching, etc.) can be used to locate the most similar original text fragment in the full text of the original contract.
[0104] Exemplarily, assume that the review item is "payment terms", and the text in the data set to be reviewed is "The payment amount should be 50% of the total contract amount, and the balance should be paid within 30 days after delivery.", but the corresponding original contract text is "Party A needs to pay 50% of the amount after the contract comes into effect, and the remaining amount will be paid within 30 days after the delivery is completed.". At this time, using the similarity algorithm to match can find a fragment with a relatively high similarity to the text in the data set to be reviewed in the full contract text and determine its specific position in the full contract text (such as PDF coordinates, Word document page number / line number, etc.).
[0105] In practical applications, in the WPS scenario, the Range coordinate interval of the target text in the document can be calculated, and the interface provided by the WPS control can be used to achieve precise positioning of the original contract text. After the positioning is completed, in the review result display interface, key content in the original contract text can be marked and displayed through visualization means such as color highlighting, icon marking, and annotation description.
[0106] Through the above processing method, the embodiment of the present application can accurately find and highlight the original clause content corresponding to each review result in the full text of the contract, significantly improving the readability and interpretability of the review results. Reviewers can intuitively see the one-to-one correspondence between the key contract content and the system's judgment results (such as risk level, modification suggestions, review conclusions) on the review result display interface, which is convenient for quickly identifying potential risks and making targeted modifications, thereby improving the efficiency and quality of contract review.
[0107] Further, reference can be made toFigure 4 The "result scoring" step. The embodiments of the present application can also receive the scores or feedback information corresponding to the review results for each review item, which are used to evaluate the accuracy and rationality of the current review results. The scores or feedback information may include, but is not limited to: the high or low scores given by users to the review results, the revised suggestions for risk judgments, the supplementary explanations for the omitted or misjudged content in the review results, etc.
[0108] Based on the above scores or feedback information, the contract review model can implement a self-learning mechanism to dynamically optimize the existing review strategies. For example, in the contract review model in the form of a large language model (LLM), the understanding ability and adaptability of the model to specific review scenarios can be continuously enhanced through methods such as incremental learning, prompt adjustment, knowledge supplementation, or fine-tuning updates; in the contract review model in the form of a plugin, the subsequent review accuracy can be improved by updating the keyword library, review rules, or template comparison logic.
[0109] By introducing a model optimization process driven by scores and feedback information, the adaptability of the contract review model to industry-specific contract types, review habits, and risk preferences can be effectively improved, thereby enhancing the sustainability and self-adaptive ability of the contract review system.
[0110] Correspondingly, in order to implement the contract review method of the embodiments of the present application, the embodiments of the present application also provide a contract review device, as Figure 5 shown. The device includes:
[0111] A processing unit 501, configured to perform sharding processing on the contract to be reviewed and assign at least one first feature label to each shard;
[0112] A determination unit 502, configured to determine the target review plan corresponding to the contract to be reviewed; the target review plan includes a plurality of review items;
[0113] A multi-dimensional retrieval unit 503, configured to respectively perform retrieval on the contract to be reviewed by using a variety of retrieval algorithms based on each shard and its first feature label in the contract to be reviewed, and obtain a set of data to be reviewed corresponding to each review item; wherein, the variety of retrieval algorithms are respectively retrieval algorithms based on different dimensions;
[0114] A review result generation unit 504, configured to input the set of data to be reviewed corresponding to each review item into the contract review model to obtain the review results corresponding to each review item.
[0115] In one embodiment, the multi-dimensional retrieval unit 503 is specifically configured to:
[0116] Based on each of the shards and their first feature tags in the contract to be reviewed, use various retrieval algorithms to retrieve the contract to be reviewed respectively, and obtain multiple retrieval sets corresponding to different retrieval algorithms;
[0117] Perform a merging and optimization process on the multiple retrieval sets to obtain a set of data to be reviewed corresponding to each review item.
[0118] In one embodiment, the target review plan further includes review rules corresponding to each review item; the processing unit 501 is specifically configured to:
[0119] Use a large language model to generate keywords, second feature tags, and text vectors corresponding to each review rule;
[0120] The various retrieval algorithms include a first retrieval algorithm based on keywords, a second retrieval algorithm based on feature tags, and a third retrieval algorithm based on text vectors; the multi-dimensional retrieval unit 503 is specifically configured to:
[0121] Based on the keywords of each review rule, use the first retrieval algorithm to perform keyword retrieval on each shard to obtain a first retrieval set;
[0122] Based on the second feature tags of each review rule and the first feature tags of each shard, use the second retrieval algorithm to retrieve each shard to obtain a second retrieval set;
[0123] Based on the text vectors of each review rule, use the third retrieval algorithm to perform vectorized retrieval on each shard to obtain a third retrieval set.
[0124] In one embodiment, the multi-dimensional retrieval unit 503 is specifically configured to:
[0125] Merge and deduplicate the multiple retrieval sets to obtain the target paragraphs corresponding to each review item in the contract to be reviewed;
[0126] For the target paragraphs corresponding to each review item, use the various retrieval algorithms to retrieve the context of the target paragraphs to obtain associated context paragraphs;
[0127] Combine the target paragraphs corresponding to each review item and their associated context paragraphs to obtain a set of data to be reviewed corresponding to each review item.
[0128] In one embodiment, the target review plan further includes review rules corresponding to each review item; each review rule generates a corresponding retrieval rule through a large language model; the retrieval rule is used to locate the set of data to be reviewed corresponding to the review item; the processing unit 501 is specifically configured to:
[0129] Perform similarity detection and / or hit detection on the retrieval rule and the data set to be reviewed under the same review item, and obtain a detection result;
[0130] If the detection result indicates that the similarity between the retrieval rule and the data set to be reviewed under the same review item is lower than a preset threshold, and / or, the detection result indicates that the data set to be reviewed that matches the retrieval rule is not hit, then use the large language model to perform a full-text search on the contract to be reviewed according to the retrieval rule, and obtain a new data set to be reviewed.
[0131] In one embodiment, the review result generation unit 504 is specifically configured to:
[0132] Input the data set to be reviewed corresponding to each review item into the contract review model corresponding to the review item, and obtain the review result for each review item; wherein, the contract review model includes a contract review model in the form of a plug-in and a contract review model in the form of a large language model.
[0133] In one embodiment, the processing unit 501 is specifically configured to:
[0134] Based on a preset desensitization rule, desensitize the target field in the contract to be reviewed, and generate a data desensitization record; the data desensitization record is used for the mapping relationship between the target field and the original value of the target field;
[0135] The processing unit 501 is specifically configured to:
[0136] Obtain the target field to be restored according to the data desensitization record;
[0137] Perform original value restoration processing on the target field according to the data desensitization record, so that the original value of the target field is used as the field content indicated by the target field.
[0138] In one embodiment, the processing unit 501 is specifically configured to:
[0139] According to each data set to be reviewed, retrieve in the contract to be reviewed using a similarity algorithm, and obtain the source text corresponding to the data set to be reviewed and the coordinate range of the area where the source text is located;
[0140] Locate and mark the source text of the contract to be reviewed according to the coordinate range of the area where the source text is located, and display the review result corresponding to the marked source text on the review result display interface.
[0141] In actual application, the processing unit 501 can be implemented by a processor in a computer device in combination with a communication interface, and the determination unit 502, the multi-dimensional retrieval unit 503, and the review result generation unit 504 can be implemented by the communication interface in the contract review device.
[0142] It should be noted that: when the above-mentioned embodiment provides a contract review device for contract review, only the division of the above-mentioned program modules is used for illustration. In actual application, the above-mentioned processing can be allocated to different program modules according to needs, that is, the internal structure of the device is divided into different program modules to complete all or part of the processing described above. In addition, the contract review device provided in the above-mentioned embodiment and the contract review method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0143] Based on the hardware implementation of the above program modules, and in order to implement a contract review method provided by an embodiment of the present application, an embodiment of the present application also provides a computer device, as Figure 6 shown, the computer device 600 includes:
[0144] A central processing unit 601, a memory 602, and an input / output interface 603;
[0145] The memory 602 is a transient storage memory or a persistent storage memory;
[0146] The central processing unit 601 is configured to communicate with the memory 602 and execute the instruction operations in the memory 602 to execute any one of the above contract review methods.
[0147] Of course, in actual application, the various components in the computer device 600 are coupled together through a bus system 604. It can be understood that the bus system 604 is used to realize the connection and communication between these components. The bus system 604 includes not only a data bus, but also a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 6 all kinds of buses are labeled as the bus system 604.
[0148] The memory 602 in the embodiment of the present application is used to store various types of data to support the operation of the computer device 600. Examples of these data include: any computer program for operating on the computer device 600.
[0149] It can be understood that when the processor in the computer device described above executes the computer program, it can also implement the functions of each unit in the corresponding device embodiments described above, which will not be elaborated here. Exemplarily, the computer program can be divided into one or more modules / units, and one or more modules / units are stored in the memory and executed by the processor to complete the various embodiments of the present application. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the computer device. For example, the computer program can be divided into the various units in the above computer device, and each unit can implement the specific functions as described in the corresponding computer device above.
[0150] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the processor and the memory are only examples of the computer device and do not constitute a limitation on the computer device. It may include more or fewer components, or combine certain components, or different components. For example, the computer device may also include input / output devices, network access devices, a bus, etc.
[0151] The processor can be a central processing unit (CPU, Central Processing Unit), or other general-purpose processors, digital signal processors (DSP, Digital Signal Processor), application specific integrated circuits (ASIC, Application Specific Integrated Circuit), field-programmable gate arrays (FPGA, Field-Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The processor is the control center of the computer device and connects various parts of the entire computer device through various interfaces and lines.
[0152] The memory can be used to store computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory, the processor can implement various functions of the computer device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.
[0153] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it executes the contract review method described in any one of the above.
[0154] An embodiment of the present application also provides a computer program product, on which a computer program / instructions are stored. When the computer program / instructions are executed by a processor, they are used to implement the contract review method described in the first aspect or any specific implementation manner of the first aspect of the embodiments of the present application.
[0155] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0156] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0157] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0158] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0159] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
Claims
1. A contract review method, characterized in that: include: Processing the contract to be reviewed in shards, and assigning at least one first characteristic tag to each shard; Determine the target review plan corresponding to the contract to be reviewed; The target review scheme includes a plurality of review items; Based on each of the shards in the contract to be reviewed and the first characteristic tag thereof, the contract to be reviewed is searched using a plurality of search algorithms to obtain a set of data to be reviewed corresponding to each of the review items; wherein the plurality of search algorithms are search algorithms based on different dimensions respectively; The data set to be reviewed corresponding to each review item is input into the contract review model to obtain the review result of each review item.
2. The method according to claim 1, characterized in that The method of searching the contracts to be reviewed respectively based on each of the fragments and the first feature tags in the contracts to be reviewed by using a plurality of search algorithms to obtain a set of data to be reviewed corresponding to each of the review items includes: Based on each of the fragments in the contract to be reviewed and the first characteristic tag thereof, the contract to be reviewed is searched respectively using multiple search algorithms to obtain multiple search sets corresponding to different search algorithms; The multiple search sets are merged and optimized to obtain a set of data to be reviewed corresponding to each review item.
3. The method according to claim 2, characterized in that The target review scheme also includes a review rule corresponding to each review item; before searching the contract to be reviewed using multiple search algorithms based on each of the fragments in the contract to be reviewed and the first feature tag thereof to obtain multiple search sets corresponding to different search algorithms, the method also includes: Generate keywords, second feature labels and text vectors for each of the review rules using a large language model; The multiple search algorithms include a first search algorithm based on keywords, a second search algorithm based on feature tags, and a third search algorithm based on text vectors; based on each of the fragments in the contract to be reviewed and its first feature tag, the contract to be reviewed is searched using multiple search algorithms to obtain multiple search sets corresponding to different search algorithms, including: Based on the keywords of each of the review rules, a keyword search is performed on each of the slices using a first search algorithm to obtain a first search set; Based on the second feature tag of each of the review rules and the first feature tag of each of the slices, using the second search algorithm to search each of the slices to obtain a second search set; Based on the text vector of each of the review rules, a third search algorithm is used to perform vectorized search on each of the fragments to obtain a third search set.
4. The method according to claim 2, characterized in that: The merging and optimizing processing of the multiple search sets to obtain a set of data to be reviewed corresponding to each review item includes: Merging and removing duplicates from the multiple search sets to obtain a target paragraph corresponding to each review item in the contract to be reviewed; For each target paragraph corresponding to the review item, the context of the target paragraph is searched using the multiple search algorithms to obtain associated context paragraphs; The target paragraph corresponding to each review item and its associated context paragraph are combined to obtain a set of data to be reviewed corresponding to each review item.
5. The method according to any one of claims 1 to 4, characterized in that: The target review scheme further includes review rules corresponding to each review item; each review rule generates a corresponding retrieval rule through a large language model; the retrieval rule is used to locate the to-be-reviewed data set corresponding to the review item; after the to-be-reviewed contract is retrieved using multiple retrieval algorithms based on each of the shards in the to-be-reviewed contract and its first feature tag to obtain the to-be-reviewed data set corresponding to each review item, the method further includes: Performing similarity detection and / or hit detection on the search rule under the same review item and the data set to be reviewed to obtain a detection result; If the detection result indicates that the similarity between the retrieval rule and the data set to be reviewed under the same review item is lower than a preset threshold, and / or the detection result indicates that the data set to be reviewed that matches the retrieval rule is not hit, then the large language model is used to perform a full-text search on the contract to be reviewed according to the retrieval rule to obtain a new data set to be reviewed.
6. The method according to any one of claims 1 to 4, characterized in that: The data set to be reviewed corresponding to each review item is input into the contract review model to obtain the review result of each review item, including: The data set to be reviewed corresponding to each review item is input into the contract review model corresponding to the review item to obtain the review result of each review item; wherein the contract review model includes a contract review model in the form of a plug-in and a contract review model in the form of a large language model.
7. The method according to any one of claims 1 to 4, characterized in that: After the contract to be reviewed is segmented, the method further includes: Based on the preset desensitization rules, desensitize the target field in the contract to be reviewed, and generate a data desensitization record; the data desensitization record is used for the mapping relationship between the target field and the original value of the target field; After inputting the to-be-reviewed data set corresponding to each of the review items into the contract review model to obtain the review result of each of the review items, the method further includes: Acquire the target field to be restored according to the data desensitization record; The target field is restored to its original value according to the data desensitization record, so that the original value of the target field serves as the field content indicated by the target field.
8. The method according to any one of claims 1 to 4, characterized in that: After inputting the to-be-reviewed data set corresponding to each of the review items into the contract review model to obtain the review result of each of the review items, the method further includes: According to each of the data sets to be reviewed, a similarity algorithm is used to search in the contracts to be reviewed to obtain the source text corresponding to the data set to be reviewed and the coordinate range of the area where the source text is located; The source text of the contract to be reviewed is located and marked according to the coordinate range of the area where the source text is located, and the review result corresponding to the marked source text is displayed on the review result display interface.
9. A contract review device, characterized in that: include: A processing unit, used for processing the contract to be reviewed in slices and assigning at least one first characteristic tag to each slice; A determination unit, configured to determine a target review scheme corresponding to the contract to be reviewed; the target review scheme includes a plurality of review items; A multi-dimensional retrieval unit, configured to retrieve the contract to be reviewed respectively using a plurality of retrieval algorithms based on each of the slices in the contract to be reviewed and the first feature tag thereof, to obtain a set of data to be reviewed corresponding to each review item; wherein the plurality of retrieval algorithms are retrieval algorithms based on different dimensions respectively; The review result generating unit is used to input the to-be-reviewed data set corresponding to each of the review items into the contract review model to obtain the review result of each of the review items.
10. A computer device, characterized in that: include: CPU, memory and input / output interface; The memory is a short-term storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute instruction operations in the memory to perform the contract review method described in any one of claims 1 to 8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the contract review method as described in any one of claims 1 to 8 is performed.
12. A computer program product, characterized in that The computer program product includes computer instructions, which, when executed by a processor, implement the contract review method as described in any one of claims 1 to 8.
Citation Information
Cited By
Intelligent contract review system based on multi-agent collaborative tool enhanced orchestration
CN121685203A
Legal document risk identification method and system and storage medium
CN122492402A