A method and system for obtaining key information of a contract template and a storage medium

By acquiring contract titles and clauses, utilizing contract classification and topic models, and combining them with a data segmentation model of judicial reasoning, the problem of fragmented and repetitive editing content in traditional contract text generation is solved. This achieves efficient classification and clustering of contract templates and provides accurate preset templates.

CN115409013BActive Publication Date: 2026-04-14GUANGDONG BOWEI CHUANGYUAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGDONG BOWEI CHUANGYUAN TECH CO LTD
Filing Date
2022-09-01
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Traditional contract text generation methods result in fragmented content, are prone to errors, and are slow and inaccurate. Furthermore, contract templates are repetitive and cannot be effectively categorized or clustered.

Method used

By obtaining the contract title and clauses, and using contract classification models, chapter theme models, and clause theme models, the differences from the preset contract templates are compared. Users are prompted to add missing chapters and clause themes. Key information is extracted by combining data on judicial reasoning and judicial opinion segmentation models to classify and cluster the contract templates.

Benefits of technology

It improves the accuracy and efficiency of contract processing, ensures that key information is not omitted from contract templates, simplifies the contract generation process, and provides accurate preset contract templates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115409013B_ABST
    Figure CN115409013B_ABST
Patent Text Reader

Abstract

A contract template key information acquisition method and system and a storage medium, the method comprising: obtaining a to-be-processed contract, obtaining a contract title and contract terms according to the to-be-processed contract, inputting the contract title into a contract classification model to obtain a contract category; inputting the body of the to-be-processed contract into a chapter theme model to obtain a chapter theme of the to-be-processed contract, comparing the chapter theme with preset chapter themes of a preset contract template corresponding to the contract category, and prompting to add the preset chapter corresponding to the missing preset chapter theme; inputting the contract terms into a term theme model to obtain the term theme of the to-be-processed contract, comparing the term theme with preset term themes of a preset contract template corresponding to the contract category, and prompting to add the preset term corresponding to the missing preset term theme. The user can directly upload the contract and obtain the missing part in the uploaded contract and the preset contract template, so that the contract will not miss the key information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data processing technology, specifically relating to a method, system, and storage medium for obtaining key information from a contract template. Background Technology

[0002] Due to the diverse nature of the business's target audience and content, various contracts need to be developed, each with different content.

[0003] However, the traditional method of generating contract texts usually involves users modifying pre-defined contract templates according to business needs. However, this editing method requires a lot of fragmented content to be edited, which is prone to errors during editing or modification. This results in slow contract processing speed and low accuracy, failing to meet business needs.

[0004] Numerous contract templates are available from various sources, but the limited number of contract types leads to significant duplication. Therefore, it's necessary to organize these templates. Organizing contract templates involves extracting key information for categorization and clause clustering, ultimately resulting in new contract templates. Summary of the Invention

[0005] To address the technical deficiencies in the background art, this invention proposes a method, system, and computer-readable storage medium for obtaining key information from contract templates, solving the aforementioned technical problems and meeting practical needs. The specific technical solution is as follows:

[0006] A method for obtaining key information of a contract template according to a first aspect of the present invention includes the following steps:

[0007] Obtain pending contracts, and based on the pending contracts, obtain the contract title and contract terms;

[0008] Input the contract title into the contract classification model to obtain the contract category;

[0009] Input the main text of the contract to be processed into the chapter topic model to obtain the chapter topic of the contract to be processed;

[0010] The preset chapter topics are compared with the preset chapter topics of the preset contract templates corresponding to the contract category. If there are preset chapter topics that exist in the preset contract template but not in the contract to be processed, a prompt is made to add the preset chapter corresponding to the missing preset chapter topic.

[0011] Input the contract terms into the terms subject model to obtain the terms subject of the contract to be processed;

[0012] The default clauses are compared with the default clauses of the default contract template corresponding to the contract category. If there are default clauses that exist in the default contract template but not in the contract to be processed, a prompt is made to add the default clauses corresponding to the missing default clauses.

[0013] The beneficial effects of the first aspect of the present invention are as follows: when a user uploads a contract, the contract title and contract terms of the uploaded contract are obtained, and the contract category of the uploaded contract is obtained by inputting the contract title into the contract classification model. The chapter theme model and the clause theme model can be used to obtain the chapter theme and clause theme of the contract to be processed. By comparing with the preset contract template of the corresponding contract category, the user is prompted to add the preset chapter corresponding to the missing preset chapter theme and the preset clause corresponding to the missing preset clause theme.

[0014] According to some embodiments of the present invention, the preset contract template is determined by the following steps:

[0015] Obtain a contract template training dataset, which includes at least one contract template data, and the contract template data includes at least contract title data and corresponding contract clause data.

[0016] Obtain the classification tree of contract categories. Based on the classification tree of contract categories, use the example-based text title classification method to classify each contract title data to obtain the contract category corresponding to each contract title data, the contract clause data corresponding to each contract title data, the preset clauses of the preset contract template corresponding to the corresponding contract category, and the preset chapters corresponding to the preset clauses.

[0017] Based on the chapter theme model and clause theme model, we obtain the clause theme data corresponding to each contract category and the chapter theme data corresponding to each clause theme data.

[0018] The clause theme data and chapter theme data are the preset clause theme and preset chapter theme of the preset contract template corresponding to the contract category, respectively;

[0019] According to the contract category, obtain the judgment reasoning dataset corresponding to the contract category. The judgment reasoning dataset includes at least one judgment reasoning data and at least two legal relationship subject data and corresponding subject name data for each judgment reasoning data. Each judgment reasoning data includes at least one judicial opinion data.

[0020] Each set of judgment reasoning data is input into the invalid reasoning identification model to obtain the identification result corresponding to each set of judgment reasoning data. Judgment reasoning data whose identification result is invalid reasoning is deleted to obtain the valid reasoning dataset.

[0021] The judicial opinion segmentation model is used to segment the judgment reasoning data in the valid reasoning dataset to obtain the judicial opinion training dataset.

[0022] The judicial opinion data, at least two legal relationship subject data and corresponding subject name data in the judicial opinion training dataset are input into the key information extraction model to obtain the key information dataset, which includes key information data.

[0023] Input the key information data in the key information dataset into the behavior consequence differentiation model to obtain the behavior pattern data and legal consequence data corresponding to each key information data, and add the corresponding behavior pattern data and legal consequence data into the key information dataset.

[0024] Cluster the behavioral pattern data in the key information dataset to obtain at least one behavioral pattern set, and the behavioral pattern data in each behavioral pattern set are similar.

[0025] Calculate the probability that the contract clause data belongs to each set of behavioral patterns, and take the behavioral pattern data in the set of behavioral patterns with the highest probability as the behavioral pattern data corresponding to the contract clause data;

[0026] Obtain missing behavior pattern data in the behavior pattern set where the number of elements exceeds a preset threshold and there is no corresponding contract clause data. Obtain missing contract clause data based on the missing behavior pattern data. Utilize the semantic similarity between the missing contract clause data and contract clause data in different preset chapters of the preset contract template, and add the missing contract clause data to the preset chapter with the highest semantic similarity.

[0027] According to some embodiments of the present invention, obtaining clause theme data corresponding to each contract category and chapter theme data corresponding to each clause theme data based on the chapter theme model and the clause theme model includes:

[0028] For each contract category, chapter theme data is obtained through the chapter theme model;

[0029] The frequency of the chapter topic data is sorted, and a preset number of chapter topic data with the highest frequency are selected as candidate chapter topics.

[0030] For each contract category, contract terms are divided, and the contract terms are input into the terms theme model to obtain the terms theme data corresponding to the contract category;

[0031] For each clause topic data, calculate the frequency of the clause topic data appearing in each candidate chapter topic, and take the chapter topic data with the highest frequency of each clause topic data as the chapter topic data corresponding to that clause topic.

[0032] According to some embodiments of the present invention, the prompt to add the missing portion of the user-uploaded contract relative to the preset contract template includes:

[0033] When a user uploads a contract, the contract title and terms are retrieved.

[0034] Input the contract title into the contract classification model to obtain the contract category;

[0035] Input the main text into the chapter theme model to obtain the chapter theme. Compare the chapter theme with the chapter theme of the preset contract template. If there is a chapter theme that exists in the preset contract template but not in the contract uploaded by the user, prompt the user to add the preset chapter corresponding to the missing preset chapter theme.

[0036] The contract terms are input into the terms theme model to obtain the terms theme of the user-uploaded contract. The terms theme of the user-uploaded contract is compared with the preset terms theme of the preset contract template. If there are terms the preset contract template has but the user-uploaded contract does not, a prompt is made to add the preset terms corresponding to the missing preset terms theme.

[0037] According to some embodiments of the present invention, the subject matter model of the terms is determined by the following steps:

[0038] Obtain a training dataset for contract terms, which includes training data for the main text of contract terms and training data for the labels of contract terms.

[0039] Based on the training data of the contract terms body text and the training data of the contract terms labels, contract terms body-label training data is constructed. The lightweight transformer-based bidirectional encoder representation model-bidirectional long short-term memory network-capsule network model is trained using the contract terms body-label training data to obtain the terms topic model.

[0040] According to some embodiments of the present invention, the invalidity reason identification model is determined as follows:

[0041] Obtain a training dataset of judges' reasons, which includes training data of judges' reasons, including the text of the judges' reasons and the corresponding valid or invalid labels;

[0042] The training dataset of the judge's reasoning is divided into a training set and a validation set. The training set includes training data and training labels, and the validation set includes validation data and validation labels.

[0043] The ALBERT-DPCNN model is trained using the training set data and validated and optimized using the validation set to obtain a judgment reason classification model. The judgment reason classification model is used to calculate the probability that a judgment reason belongs to a valid reason.

[0044] Based on the training data and training labels in the training set, the string length of the adjudication reason after removing commas and periods is obtained. The string length is divided into a first preset number of string length intervals, and the probability that the adjudication reason data corresponding to each string length interval is a valid reason is calculated based on the number of words.

[0045] Based on the training data and training labels in the training set, the number of commas and periods is obtained, and the number of commas and periods is divided into a second preset number of comma and period number intervals. The probability that the punctuation count of the judge's reason data corresponding to each of the comma and period number intervals is a valid reason is calculated.

[0046] For each piece of verification data in the verification set, the verification data is input into the judgment reason classification model to obtain the probability of a valid reason, the probability of the number of words being valid and the probability of the number of punctuation marks being valid are calculated, and the verification probability that each piece of verification data is a valid reason is calculated based on the probability of a valid reason, the probability of the number of words being valid and the probability of the number of punctuation marks being valid, and the parameters of the calculation formula are adjusted based on the verification probability and the verification label.

[0047] According to some embodiments of the present invention, the judicial opinion segmentation model is determined as follows:

[0048] A segmented training dataset of judicial opinions is obtained. The segmented training dataset of judicial opinions contains several reasons for judgment. Each reason for judgment includes at least one judicial opinion. Each judicial opinion is labeled using the BEMS labeling method to obtain labeled training data. The judicial opinion is labeled as B at the beginning, E at the end, M in the middle, and S at the other parts.

[0049] Based on the labeled training data obtained using the BEMS annotation method, training data conforming to the reading comprehension model is obtained. A named entity recognition model composed of a transducer-based bidirectional encoder representation model with attention decoupling and enhanced decoding, a long short-term memory network with vector quantization, and a machine reading comprehension model is trained. The judicial opinion segmentation model can divide the judgment reasoning into at least one judicial opinion.

[0050] According to some embodiments of the present invention, the key information extraction model is determined as follows:

[0051] A key information extraction training dataset is obtained, which includes the original text of paired judicial opinions and the key information corresponding to the judicial opinions.

[0052] The original text of the judicial opinion is input into the generator of the generative adversarial network to obtain the key prediction information;

[0053] The predicted key information and the key information corresponding to the judicial opinion are input into the discriminator of the generative adversarial network to obtain the similarity between the predicted key information and the key information corresponding to the judicial opinion.

[0054] The parameters of the generator are adjusted based on the similarity, so that the generator can generate key information data whose text similarity to the key information corresponding to the judicial opinion is greater than a preset threshold.

[0055] According to some embodiments of the present invention, the behavioral consequence differentiation model is determined as follows:

[0056] Obtain a behavioral consequences training dataset, which includes text and corresponding labels for behavioral patterns or legal consequences;

[0057] Using the text as input and the corresponding label as output, train the ALBERT-DPCNN model to obtain the behavior consequence discrimination model.

[0058] A system for obtaining key information of a contract template according to a second aspect of the present invention includes:

[0059] Input unit, through which the user uploads the contract;

[0060] The processing unit is used to implement the method for obtaining key information of the contract template as described in the first aspect embodiment, and to calculate the part of the contract uploaded by the user that is missing from the contract template.

[0061] The output unit is used to prompt the user to add the missing parts of the contract uploaded by the user relative to the contract template.

[0062] The beneficial effects of the second aspect of the present invention are as follows: the contract template key information acquisition system can acquire key information of the contract template, classify the contract template, and cluster the clauses to finally obtain a preset contract template. Users can directly upload the contract to the contract template key information acquisition system and obtain the uploaded contract and the missing parts in the preset contract template. It is simple and convenient, and ensures that no key information is omitted from the contract.

[0063] According to a third aspect of the present invention, a computer-readable storage medium stores computer-executable instructions for causing a computer to perform a method for obtaining key information of a contract template as described in a first aspect of the present invention. Attached Figure Description

[0064] Figure 1 This is a flowchart of a method for obtaining key information from a contract template according to an embodiment of the present invention.

[0065] Figure 2 This is a flowchart illustrating the determination of a clause subject model provided in one embodiment of the present invention.

[0066] Figure 3 This is a flowchart illustrating the determination of an invalidity reason identification model provided in one embodiment of the present invention.

[0067] Figure 4 This is a flowchart illustrating the determination of a judicial opinion segmentation model provided in one embodiment of the present invention.

[0068] Figure 5 This is a flowchart illustrating the determination of a key information extraction model provided in one embodiment of the present invention.

[0069] Figure 6 This is a flowchart illustrating the determination of a behavioral consequence differentiation model provided in one embodiment of the present invention. Detailed Implementation

[0070] The embodiments of the present invention will be described below with reference to the accompanying drawings and related examples. The embodiments of the present invention are not limited to the following examples, and the present invention relates to the relevant necessary components in this technical field, which should be regarded as well-known technology in this technical field and can be known and mastered by those skilled in this technical field.

[0071] The following is for reference. Figures 1 to 6 A method for obtaining key information of a contract template according to an embodiment of the first aspect of the present invention is described.

[0072] like Figure 1 As shown, Figure 1 This is a flowchart illustrating a method for obtaining key information from a contract template according to an embodiment of the present invention. The method for obtaining key information from a contract template according to an embodiment of the present invention includes, but is not limited to, steps S100, S200, S300, S400, S500, and S600.

[0073] Step S100: Obtain the contract to be processed; based on the contract to be processed, obtain the contract title and contract terms.

[0074] Step S200: Input the contract title into the contract classification model to obtain the contract category;

[0075] Step S300: Input the main text of the contract to be processed into the chapter topic model to obtain the chapter topic of the contract to be processed;

[0076] Step S400: Compare the preset chapter topics of the chapter topics with the preset chapter topics of the preset contract templates corresponding to the contract categories. If there are preset chapter topics that exist in the preset contract templates but not in the contract to be processed, prompt the user to add the preset chapter corresponding to the missing preset chapter topic.

[0077] Step S500: Input the contract terms into the terms subject model to obtain the terms subject of the contract to be processed;

[0078] Step S600: Compare the default clause topics of the clause topics with the default clause topics of the default contract templates corresponding to the contract categories. If there are default clause topics that exist in the default contract templates but not in the contract to be processed, prompt the user to add the default clauses corresponding to the missing default clause topics.

[0079] When a user uploads a contract, the system retrieves the contract title and terms. The contract title is input into a contract classification model to determine the contract category. The main text of the contract is input into a chapter theme model to obtain the chapter themes. These chapter themes are compared to those in a pre-defined contract template. If a chapter theme exists in the template but not in the uploaded contract, the system prompts the user to add the corresponding chapter. Similarly, the contract terms are input into a clause theme model to obtain the clause themes. These clause themes are compared to those in the template. If a clause theme exists in the template but not in the uploaded contract, the system prompts the user to add the corresponding clause. This system compares user-uploaded contracts with pre-defined contract templates and prompts the user to add any missing chapters or clauses.

[0080] In some embodiments, the contract classification model obtains a classification tree of contract types and then uses an example-based text title classification method to classify contract titles.

[0081] In some embodiments, when a user uploads a contract titled "House Rental Contract," the contract is input into the contract classification model, resulting in a contract category of "Lease Contract." The text of the uploaded contract is input into the chapter theme model to obtain the chapter themes "Contract Parties, Subject Matter, Price and Payment, Rights and Obligations, Liability for Breach of Contract, Contract Effectiveness." The preset chapter themes for the contract category "Lease Contract" are "Contract Parties, Subject Matter, Price and Payment, Rights and Obligations, Liability for Breach of Contract, Contract Effectiveness, Dispute Resolution." After comparison, if there is a chapter theme "Dispute Resolution" present in the preset contract template but absent in the user's uploaded contract, a prompt is made to add the corresponding preset chapter "Dispute Resolution." The contract terms of the user's uploaded contract are input into the clause theme model to obtain the clause themes. The clause themes of the user's uploaded contract are compared with the preset clause themes of the preset contract template. If the preset contract template has a "Priority Lease Clause" but the user's uploaded contract does not, a prompt is made to add the corresponding preset clause: "Upon expiration of the contract, if the lessor continues to lease, the lessee has the priority lease right under the same conditions."

[0082] In some embodiments, determining the chapter topic model includes, but is not limited to, the following steps:

[0083] Obtain the contract training dataset, which includes contract type training data and contract text training data.

[0084] Based on the contract type training data and contract text training data, chapter topic training data and chapter text training data are obtained, and the chapter topic training data are merged.

[0085] The BEMS annotation method is used to annotate the text of the chapter training data, where the beginning of the chapter is labeled as B, the end as E, and the middle part as M, to obtain the annotated training data.

[0086] Using the chapter text training data as input and the labeled training data obtained by the BEMS annotation method as output, the BERT-BILSTM-CRF model is trained to obtain the chapter segmentation model.

[0087] The BERT-BILSTM-ATTENTION-RCNN model is trained using chapter text training data as input and chapter topic training data as output, resulting in a text topic model.

[0088] Obtain the contract training dataset. Based on the contract type training data and contract text training data in the contract training dataset, obtain chapter topic training data and chapter text training data, and merge the chapter topic training data. For the text in the chapter text training data, use the BEMS annotation method to annotate it, labeling the beginning of the chapter as B, the end as E, and the middle part as M, to obtain annotated training data. Then, using the chapter text training data as input and the annotated training data obtained by the BEMS annotation method as output, train the BERT-BILSTM-CRF model to obtain the chapter segmentation model. Using the chapter text training data as input and the chapter topic training data as output, train the BERT-BILSTM-ATTENTION-RCNN model to obtain the text topic model. Using the contract training dataset, we obtained the chapter segmentation model and the text topic model. The chapter segmentation model is used to divide a certain amount of text into a chapter based on its content. BERT-BILSTM-CRF can divide multiple sentences of text into different chapters so that the contract text can be reordered after adding new clauses. The text topic model is used to classify a certain amount of text into a chapter topic. The BERT layer generates vectors containing contextual information. These trained vectors are then used as input to a BiLSTM layer. The BiLSTM consists of two LSTM layers with different orientations. After computation, the predictions are concatenated, and this concatenated result is used as input to the next CRF layer. The constraints imposed by the CRF layer on the output sequence effectively prevent erroneous information from the BiLSTM layer's output. First, a chapter segmentation model is used to divide a certain amount of text into chapters. Then, a main topic model is used to classify the text into a chapter topic, thus determining the chapter topic.

[0089] In some embodiments of the present invention, the determination of the preset contract template includes, but is not limited to, the following steps:

[0090] Obtain a contract template training dataset. The contract template training dataset includes at least one contract template data, and the contract template data includes at least contract title data and corresponding contract clause data.

[0091] Obtain the classification tree of contract categories. Based on the classification tree of contract categories, use the example-based text title classification method to classify each contract title data to obtain the contract category corresponding to each contract title data, the contract clause data corresponding to each contract title data, the preset clauses of the preset contract template corresponding to the corresponding contract category, and the preset chapters corresponding to the preset clauses.

[0092] Based on the chapter theme model and clause theme model, clause theme data corresponding to each contract category and chapter theme data corresponding to each clause theme data are obtained; the clause theme data and chapter theme data are the preset clause theme and preset chapter theme of the preset contract template corresponding to the contract category, respectively.

[0093] Based on the contract category, obtain the corresponding dataset of judgment reasons for that contract category. The dataset of judgment reasons includes at least one data point of judgment reason and at least two data points of legal relationship subjects and corresponding subject name data for each data point of judgment reason. Each data point of judgment reason includes at least one data point of judicial opinion.

[0094] Each set of judgment reasoning data is input into the invalid reasoning identification model to obtain the identification result corresponding to each set of judgment reasoning data. The judgment reasoning data whose identification result is invalid reasoning is deleted to obtain the valid reasoning dataset.

[0095] The judicial opinion segmentation model is used to segment the judgment reasoning data in the valid reasoning dataset to obtain the judicial opinion training dataset.

[0096] Input the judicial opinion data, at least two legal relationship subject data and corresponding subject name data from the judicial opinion training dataset into the key information extraction model to obtain the key information dataset, which includes key information data.

[0097] Input the key information data from the key information dataset into the behavior consequence differentiation model to obtain the behavior pattern data and legal consequence data corresponding to each key information data, and add the corresponding behavior pattern data and legal consequence data into the key information dataset.

[0098] Cluster the behavioral pattern data in the key information dataset to obtain at least one behavioral pattern set, and the behavioral pattern data in each behavioral pattern set are similar.

[0099] Calculate the probability that the contract clause data belongs to each set of behavioral patterns, and take the behavioral pattern data in the set of behavioral patterns with the highest probability as the behavioral pattern data corresponding to the contract clause data;

[0100] Get missing behavior pattern data where the number of elements in the behavior pattern set exceeds a preset threshold and there is no corresponding contract clause data. Get missing contract clause data based on the missing behavior pattern data. Use the semantic similarity between the missing contract clause data and the contract clause data in different preset chapters of the preset contract template to add the missing contract clause data to the preset chapter with the highest semantic similarity.

[0101] Obtain the contract template training dataset and the classification tree for contract categories. Use an example-based text title classification method to classify the contract title data in the contract template training dataset, obtaining the contract category corresponding to each contract title data. The contract clause data corresponding to each contract title data consists of the preset clauses and preset chapters of the preset contract template corresponding to the corresponding contract category. Based on the chapter theme model and clause theme model, obtain the clause theme and the chapter theme corresponding to each contract category. The clause theme data and chapter theme data are the preset clause theme and preset chapter theme of the preset contract template corresponding to the contract category, respectively. Based on the contract category, obtain the corresponding judgment reasoning dataset. The judgment reasoning dataset includes at least one judgment reasoning data and at least two legal relationship subject data and corresponding subject name data for each judgment reasoning data. Each judgment reasoning data in the judgment reasoning dataset includes at least one judicial opinion data. Invalid grounds identification model is used to identify and delete invalid grounds in the judgment grounds dataset, resulting in a valid grounds dataset. Then, a judicial opinion segmentation model is used to segment the judgment grounds data in the valid grounds dataset, resulting in a judicial opinion training dataset. The judicial opinion data, at least two legal relationship subject data, and corresponding subject name data from the judicial opinion training dataset are input into a key information extraction model to obtain a key information dataset. The key information data from the key information dataset is input into a behavior consequence differentiation model to obtain the behavior pattern data and legal consequence data corresponding to each key information data. The corresponding behavior pattern data and legal consequence data are added to the key information dataset, and the behavior pattern data in the key information dataset is clustered to obtain at least one behavior pattern set. The behavior pattern data in each behavior pattern set is similar. Calculate the probability that contract clause data belongs to each set of behavioral patterns, take the set of behavioral patterns with the highest probability as the behavioral pattern data to which the contract clause data belongs, and obtain the missing behavioral pattern data in the behavioral pattern set where the number of elements exceeds a preset threshold and there is no corresponding contract clause data. Based on the missing behavioral pattern data, increase the acquisition of missing contract clause data. Utilize the semantic similarity between the missing contract clause data and the contract clause data in different preset chapters of the preset contract template, and add the missing contract clause data to the preset chapter with the highest semantic similarity.

[0102] Contract titles are characterized by their ability to accurately reflect the contract type, but their brevity makes them unsuitable for deep learning or machine learning. Therefore, an example-based text title classification mechanism leverages contextual information to determine the relevance of the title to the category, rather than solely relying on co-occurrence information, thus improving the accuracy of contract title classification. Furthermore, the judgment reasoning data obtained based on contract categories typically includes a large amount of invalid content. An invalid reasoning identification model can be used to identify invalid reasons within the judgment reasoning dataset, yielding a dataset of valid reasons. Courts often present multiple judicial opinions in the judgment reasoning paragraphs, addressing different issues. Directly performing subsequent calculations without segmenting these opinions leads to inaccurate extraction of key information. Therefore, using a judicial opinion segmentation model to segment judicial opinions improves the accuracy of key information extraction. In judgment documents, behavioral patterns are often the source of controversy. Key information can be used to derive the corresponding behavioral patterns and legal consequences, and clustering these patterns can identify frequently controversial situations. By analyzing behavioral patterns in court judgments, potentially contentious issues can be identified. However, these contentious issues may not be universally applicable. Therefore, it's necessary to select common controversial matters from these issues and add corresponding clauses. Furthermore, contract clauses should be added based on behavioral patterns lacking corresponding contractual clauses. This results in a pre-defined contract template that includes not only common chapters and clauses from the contract template dataset but also frequently found in court judgments and is prone to controversy. Key information from the contract templates is then used for classification and clause clustering, ultimately yielding a pre-defined contract template. Users can directly upload their contracts and receive both the uploaded contract and any missing parts of the pre-defined template—a simple and convenient process that ensures no crucial information is omitted.

[0103] In some embodiments, semantic similarity can be calculated based on a tree hierarchy.

[0104] In some embodiments, when the contract type is a lease contract, the judgment reasoning dataset includes the following judgment reasoning: "The leased property delivered by Party A to Party B is defective, and Party A should bear the guarantee responsibility to Party B. Therefore, this court supports Party B's claim for compensation from Party A for losses caused by the defective leased property. A lease contract is established between Party A and Party B, and both parties agree that Party B should return the leased property to Party A upon the expiration of the contract. Therefore, this court supports Party A's claim for the return of the leased property by Party B. Party A's claim for payment of the deposit by Party B is without legal basis and is not supported by this court." Each judgment reasoning corresponds to at least two legal relationship entities and their corresponding names: "Party A corresponds to the lessor, and Party B corresponds to the lessee." The invalid grounds identification model identified the invalid grounds in the judgment reasoning dataset as "Party A's claim for Party B to pay the deposit is without legal basis and is not supported by this court." These invalid grounds were then removed from the judgment reasoning dataset, resulting in the valid grounds dataset: "The leased property delivered by Party A to Party B is defective, and Party A should bear the guarantee responsibility to Party B. Therefore, this court supports Party B's claim for compensation from Party A for losses caused by the defective leased property. A lease contract was established between Party A and Party B, and both parties agreed that Party B should return the leased property to Party A upon the expiration of the contract. Therefore, this court supports Party A's claim for compensation from Party B for losses caused by the defective leased property. A lease contract was established between Party A and Party B, and both parties agreed that Party B should return the leased property to Party A upon the expiration of the contract. Therefore, this court supports Party A's claim for compensation from Party B for the return of the leased property." The court supports the plaintiff's claim for the return of the leased property. Using a judicial opinion segmentation model, the court segmented the legal arguments in the valid grounds dataset, resulting in the judicial opinion dataset: "Article 1: The leased property delivered by Party A to Party B is defective. Party A shall bear the guarantee responsibility to Party B. Therefore, the court supports Party B's claim for compensation from Party A for losses caused by the defective leased property. Article 2: A lease contract is established between Party A and Party B, and both parties agree that Party B shall return the leased property to Party A upon the expiration of the contract. Therefore, the court supports Party A's claim for the return of the leased property to Party B." Inputting the judicial opinion data, at least two legal relationship subject data, and corresponding subject name data from the judicial opinion training dataset into the key information extraction model, the court obtained the key information dataset: "Article 1: The leased property delivered by the lessor to the lessee is defective. The lessor shall bear the guarantee responsibility to the lessee. Therefore, the lessee has the right to claim compensation from the lessor for losses caused by the defective leased property. Article 2: A lease contract is established, stipulating that the lessee shall return the leased property to the lessor upon the expiration of the contract. Therefore, the lessor has the right to claim the return of the leased property from the lessee." Using the behavioral consequence differentiation model, the behavioral pattern corresponding to each key piece of information is: "Article 1: The leased property delivered by the lessor to the lessee has defects, and the lessor shall bear the guarantee responsibility to the lessee. Article 2: The lease contract is established, stipulating that the lessee shall return the leased property to the lessor upon the expiration of the contract." and the legal consequences are: "Article 1: The lessee has the right to claim compensation from the lessor for losses caused by the defects of the leased property. Article 2: The lessor has the right to claim the return of the leased property from the lessee."

[0105] In some embodiments, if the contract clause is “Party A shall guarantee that the leased property delivered to Party B is in normal working order,” the probability that it is a warranty against defects is 0.9, and the probability that it is a rights and obligations clause is 0.1. Therefore, this clause is a warranty against defects.

[0106] In some embodiments, the preset threshold is 20 times. If the corresponding contract clause behavior pattern is "priority lease", then when the number of elements in the obtained behavior pattern set exceeds the preset threshold and there is no behavior pattern with a corresponding contract clause, a contract clause is added according to the behavior pattern "priority lease". The added contract clause is "After the contract expires, if the lessor continues to lease, the lessee has the priority lease right under the same conditions". Based on the semantic similarity between the contract clause and the contract clause in the preset chapter, the contract clause is added to the preset chapter "Rights and Obligations" to obtain the preset contract template.

[0107] In some embodiments of the present invention, based on the chapter theme model and the clause theme model, clause theme data corresponding to each contract category and chapter theme data corresponding to each clause theme data are obtained, including but not limited to the following steps:

[0108] For each contract category, chapter theme data is obtained through the chapter theme model;

[0109] Sort the chapter topic data by frequency, and select the preset number of chapter topics with the highest frequency as candidate chapter topics;

[0110] For each contract category, the contract terms are divided, and the contract terms are input into the terms subject model to obtain the terms subject data corresponding to the contract category.

[0111] For each clause topic data, calculate the frequency of the clause topic data appearing in each candidate chapter topic, and take the chapter topic data with the highest frequency of each clause topic data as the chapter topic data corresponding to that clause topic.

[0112] Chapter-theme models are used to obtain chapter-theme data for different contract categories. These chapter-theme data are then sorted by frequency of occurrence, and the most frequent chapter-theme data is selected as candidate chapter-themes. Contract clauses are further categorized for different contract categories, and clause-theme models are used to obtain the corresponding clause-theme data. For each clause-theme, the frequency of its appearance in each candidate chapter-theme is calculated, and the most frequent chapter-theme data for that clause-theme is selected as its corresponding chapter-theme data. Because each contract has unique characteristics, chapters are added based on these characteristics, resulting in inconsistencies in chapter-theme data across contract types. Therefore, the chapter-theme data in the dataset is statistically analyzed, and the most frequent chapter-themes are selected to identify the most common chapter-themes for each contract type. Furthermore, different contract templates within the same type generally include similar clause-themes, but due to template specificity, the same clause-themes may be distributed across different chapter-themes. Therefore, contract clauses can be assigned to the most common chapters.

[0113] like Figure 2 As shown, in some embodiments of the present invention, the determination of the subject matter model includes, but is not limited to, steps S510 and S520.

[0114] Step S510: Obtain the contract terms training dataset, which includes training data of the main text of the contract terms and training data of the labels of the contract terms.

[0115] Step S520: Construct contract text-label training data based on the contract text training data and the contract label training data, and use the contract text-label training data to train a lightweight transformer-based bidirectional encoder representation model-bidirectional long short-term memory network-capsule network model to obtain the clause topic model.

[0116] In some embodiments, constructing contract text-label training data based on the contract terms training data and the contract terms label training data includes:

[0117] The contract text-label training data is constructed in the form of "[cls]label[sep]text[sep]", where "label" is the contract clause label training data and "text" is the contract clause text training data.

[0118] In some embodiments, the lightweight transducer-based bidirectional encoder representation model-bidirectional long short-term memory network-capsule network model includes: a lightweight transducer-based bidirectional encoder representation model, a bidirectional long short-term memory network, and a capsule network, wherein the capsule network includes: a primary capsule layer, a convolutional capsule layer, and a fully connected capsule layer.

[0119] A training dataset for contract terms is obtained, comprising training data for the main text of contract terms and training data for contract terms labels. Then, a contract text-label training dataset is constructed based on the main text and label training data. A lightweight, transformer-based bidirectional encoder representation model-bidirectional long short-term memory network-capsule network model is trained using this dataset to obtain a clause topic model. The clause label model is used to determine the clause category based on the main text of the contract. ALBERT layers offer fast training speed and excellent language representation, capturing specific information at specified locations in the context. The bidirectional long short-term memory network can acquire all forward and backward semantic information in the text. The capsule network achieves good results in text multi-classification tasks even with limited training data. The resulting clause label model enables the determination of the contract clause category based on the main text of the contract.

[0120] like Figure 3 As shown, in some embodiments of the present invention, the determination of the invalidity reason identification model includes, but is not limited to, steps S411, S412, S413, S414, S415 and S416.

[0121] Step S411: Obtain the training dataset of the judge's reasons. The training dataset of the judge's reasons includes training data of the judge's reasons, which includes the text of the judge's reasons and the corresponding valid or invalid labels.

[0122] Step S412: Divide the training dataset of the judge's reasoning into a training set and a validation set. The training set includes training data and training labels, and the validation set includes validation data and validation labels.

[0123] Step S413: Train the ALBERT-DPCNN model using the training set data, and validate and optimize it using the validation set to obtain the judgment reason classification model. The judgment reason classification model is used to calculate the probability that the judgment reason belongs to a valid reason.

[0124] Step S414: Based on the training data and training labels in the training set, obtain the string length of the adjudication reason after removing commas and periods, divide the string length into a first preset number of string length intervals, and calculate the number of words and the probability of the adjudication reason data corresponding to each string length interval being a valid reason.

[0125] Step S415: Based on the training data and training labels in the training set, obtain the number of commas and periods, divide the number of commas and periods into a second preset number of comma and period number intervals, and calculate the probability that the punctuation count of the referee's reason data corresponding to each comma and period number interval is a valid reason.

[0126] Step S416: For each verification data in the verification set, input the verification data into the adjudication reason classification model to obtain the probability of a valid reason, calculate the probability of validity based on the number of words and the number of punctuation marks, and calculate the verification probability of each verification data being a valid reason based on the probability of validity based on the number of words and the number of punctuation marks, and adjust the parameters of the calculation formula based on the verification probability and the verification label.

[0127] Obtain the training dataset for the judge's reasons. The training dataset for the judge's reasons includes the text of the judge's reasons and the corresponding valid or invalid labels. Divide the judge's reasons training model to obtain a training set including training data and training labels and a validation set including validation data and validation labels. Use the training set data to train the ALBERT-DPCNN model, and use the validation set to validate and optimize it to obtain the judge's reasons classification model. Based on the training data and labels in the training set, the string length of the adjudication reason (excluding commas and periods) is obtained. This string length is divided into a first preset number of string length intervals. The probability of a valid adjudication reason being a valid reason based on the number of characters within each string length interval is calculated. Similarly, based on the training data and labels in the training set, the number of commas and periods is obtained. This number of commas and periods is divided into a second preset number of comma and period number intervals. The probability of a valid adjudication reason being a valid reason based on the number of punctuation marks within each comma and period number interval is calculated. The validation data from the validation set is input into the adjudication reason classification model to obtain the probability of a valid reason. The probability of a valid reason based on the number of characters and the probability of a valid reason based on the number of punctuation marks are calculated. Based on these probabilities, the validation probability of each validation data point being a valid reason is calculated. The parameters of the calculation formula are adjusted based on the validation probabilities and validation labels. The invalid reason identification model can determine the probability of a judgment reason being a valid reason by calculation. Invalid reason datasets can be deleted from the adjudication reason dataset based on the probabilities calculated by the invalid reason identification model, thus obtaining a dataset of valid reasons and improving the accuracy of valid reason identification.

[0128] In some embodiments, the probability that a text with a length of 12 is a valid reason is 0.2, and the probability that a text with three commas and periods is a valid reason is 0.1.

[0129] like Figure 4 As shown, in some embodiments of the present invention, the determination of the judicial viewpoint segmentation model includes, but is not limited to, steps S421 and S422.

[0130] Step S421: Obtain the training dataset of judicial opinions segmentation. The training dataset of judicial opinions segmentation contains several judgment reasons. Each judgment reason includes at least one judicial opinion. Use the BEMS annotation method to annotate each judicial opinion to obtain annotated training data. The beginning of the judicial opinion is labeled as B, the end as E, the middle part as M, and the other parts as S.

[0131] Step S422: Based on the labeled training data obtained using the BEMS annotation method, training data conforming to the reading comprehension model is obtained. The named entity recognition model, composed of an attention-decoupled enhanced decoding-based bidirectional encoder representation model, a long short-term memory network with vector quantization, and a machine reading comprehension model, is trained to obtain a judicial opinion segmentation model. The judicial opinion segmentation model can divide the judgment reasoning into at least one judicial opinion.

[0132] A training dataset for segmenting judicial opinions is obtained. This dataset contains several judgment reasons, each including at least one judicial opinion. Each judicial opinion is labeled using the BEMS annotation method: the beginning of the judicial opinion is labeled B, the end E, the middle part M, and the remaining parts S. Based on the labeled training data obtained using the BEMS annotation method, training data conforming to a reading comprehension model is obtained. This training model consists of a bidirectional encoder representation model based on a transformer and enhanced decoding with attention decoupling, a long short-term memory network with vector quantization, and a machine reading comprehension model. This results in a judicial opinion segmentation model, which can divide judgment reasons into at least one judicial opinion. Segmenting judicial opinions improves the accuracy of extracting key information from them.

[0133] In some embodiments, the named entity recognition model, composed of an attention-decoupled, enhanced-decoding, transformer-based bidirectional encoder representation model, a vector-quantized long short-term memory network, and a machine reading comprehension model, includes the attention-decoupled, enhanced-decoding, transformer-based bidirectional encoder representation model, the vector-quantized long short-term memory network, and the machine reading comprehension model. The attention-decoupled, enhanced-decoding, transformer-based bidirectional encoder representation model includes: a first word embedding layer for generating word content embedding vectors and position vectors; a transformer layer for calculating attention weights between words based on word content and relative position; a second word embedding layer for generating absolute word position vectors; and an enhanced decoding layer for decoding masked words based on aggregated context embeddings of word content and position. The vector-quantized long short-term memory network includes: two layers of disoriented long short-term memory networks, a vector quantization module corresponding to each long short-term memory network, and a decoding module corresponding to each vector quantization module. The machine reading comprehension model includes: a binary classifier for predicting start position labels, a binary classifier for predicting end position labels, and a probability matrix classifier. The attention-decoupled, augmented decoding-based transformer-based bidirectional encoder representation model generates vectors containing contextual information from the input text. These vectors are then used as input to a vector-quantized Long Short-Term Memory (LSTM) network. This LSM network consists of two layers of dissimilar LSMs, each with its corresponding vector quantization module and decoding module. The predictions are concatenated and used as input to the next layer of the machine reading comprehension model. By constraining the output sequence, the model effectively avoids errors in the LSM network's output, improving accuracy. This attention-decoupled, augmented decoding-based transformer-based bidirectional encoder representation model can represent text content, relative text position, and absolute text position using vectors. Since contract performance information often requires both relative and absolute text positions for determination, incorporating these into vector representation improves model accuracy. Furthermore, the vector-quantized LSM network compresses the data volume generated by the LSM network, increasing the computation speed of subsequent models. Machine reading comprehension models encode prior knowledge, which can reduce the impact of sparsity in training data caused by a lack of labeled data.

[0134] like Figure 5 As shown, in some embodiments of the present invention, the determination of the key information extraction model includes, but is not limited to, steps S431, S432, S433 and S434.

[0135] Step S431: Obtain the key information extraction training dataset, which includes the original text of paired judicial opinions and the key information corresponding to the judicial opinions.

[0136] Step S432: Input the original text of the judicial opinion into the generator of the generative adversarial network to obtain the predicted key information;

[0137] Step S433: Input the predicted key information and the key information corresponding to the judicial opinion into the discriminator of the generative adversarial network to obtain the similarity between the predicted key information and the key information corresponding to the judicial opinion.

[0138] Step S434: Adjust the parameters of the generator according to the similarity, so that the generator can generate key information data with a text similarity greater than a preset threshold between the key information corresponding to the judicial opinion.

[0139] A training dataset for key information extraction is obtained. The legal relationship subjects in the judicial opinions are replaced with their corresponding names. The original text of the judicial opinions is input into the generator of a generative adversarial network (GAN) to obtain predicted key information. The predicted key information and the corresponding key information in the judicial opinions are then input into the discriminator of the GAN to obtain the similarity between the two. Based on this similarity, the generator parameters are adjusted so that the generator can generate key information data with a text similarity greater than a preset threshold to the key information corresponding to the judicial opinions. Using the key information extraction model, a key information dataset containing key information data can be obtained.

[0140] like Figure 6 As shown, in some embodiments of the present invention, the determination of the behavioral consequence differentiation model includes, but is not limited to, steps S441 and S442.

[0141] Step S441: Obtain the behavioral consequences training dataset, which includes text and corresponding behavioral patterns or legal consequences labels.

[0142] Step S442: Using text as input and corresponding labels as output, train the ALBERT-DPCNN model to obtain the behavior consequence discrimination model.

[0143] Obtain a training dataset of behavioral consequences, using text as input and the corresponding labels of behavioral patterns or legal consequences as output, to train an ALBERT-DPCNN model, obtaining a behavioral consequence discrimination model. Then, use a validation set to validate and optimize the behavioral consequence discrimination model. The behavioral consequence discrimination model is used to calculate the probability that a text belongs to a behavioral pattern or legal consequence.

[0144] In some embodiments, a validation set is used to validate and optimize the behavioral consequence differentiation model.

[0145] The following is for reference. Figures 1 to 6 A system for obtaining key information of a contract template according to an embodiment of the second aspect of the present invention is described.

[0146] The system for obtaining key information from a contract template includes an input unit, a processing unit, and an output unit. The input unit is used by the user to upload a contract; the processing unit is used to implement the method for obtaining key information from a contract template in any embodiment of the first aspect, calculating the missing parts of the user-uploaded contract relative to the contract template; and the output unit is used to prompt the user to add the missing parts of the user-uploaded contract relative to the contract template.

[0147] When a user uploads a contract, the system identifies the type of contract based on the uploaded contract, finds the corresponding preset contract template, compares the user's uploaded contract with the preset contract template, prompts the user to add the missing parts of the uploaded contract relative to the preset contract template, and outputs the added contract portion.

[0148] The system for obtaining key information of a contract template acquires the contract uploaded by the user through an input unit. Then, through a processing unit, it uses the method described in the first aspect of the embodiment to obtain the contract category based on the contract title and the chapter theme. This is then compared with preset chapter themes of the same contract category. If a chapter theme exists in the preset contract template but not in the user-uploaded contract, the output unit prompts the user to add the corresponding preset chapter. The processing unit obtains the clause theme of the user-uploaded contract using the contract clauses. It compares the clause theme of the user-uploaded contract with the preset clause themes of the preset contract template. If the preset contract template contains a clause theme but the user-uploaded contract does not, the output unit prompts the user to add the corresponding preset clause theme.

[0149] The system for obtaining key information from contract templates can acquire key information from contract templates, classify contract templates, and cluster clauses to obtain preset contract templates. Users can directly upload contracts to the system and obtain the uploaded contracts and the missing parts of the preset contract templates. It is simple and convenient, ensuring that no key information is missed in the contracts.

[0150] According to a third aspect embodiment of the present invention, a computer-readable storage medium stores computer-executable instructions that are executed by a processor or controller, for example, by a processor in the above-described apparatus embodiments, causing the processor to perform the method for obtaining key information of the contract template in any of the above embodiments, for example, performing the above... Figures 1 to 6 Chinese method.

[0151] The above description is only a preferred embodiment of the present invention. It should be noted that those skilled in the art can make several improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for obtaining key information from a contract template, characterized in that, include: Obtain pending contracts, and based on the pending contracts, obtain the contract title and contract terms; Input the contract title into the contract classification model to obtain the contract category; Input the main text of the contract to be processed into the chapter topic model to obtain the chapter topic of the contract to be processed; The preset chapter topics of the chapter topics and the preset chapter topics of the preset contract templates corresponding to the contract categories are compared. If there are preset chapter topics that the preset contract templates have but the contract to be processed does not have, a prompt is made to add the preset chapter corresponding to the missing preset chapter topic. Input the contract terms into the terms subject model to obtain the terms subject of the contract to be processed; The default clause topics of the default clause topics and the default clause topics of the default contract templates corresponding to the contract categories are compared. If there are default clause topics that exist in the default contract templates but not in the contract to be processed, a prompt is made to add the default clauses corresponding to the missing default clause topics. The chapter topic model is trained using the BERT-BILSTM-CRF model, and the clause topic model is trained using a lightweight transformer-based bidirectional encoder representation model-bidirectional long short-term memory network-capsule network model. The preset contract template is determined by the following steps: Obtain the contract template training dataset and the classification tree of contract categories. Use the example-based text title classification method to classify the contract title data in the contract template training dataset to obtain the contract category corresponding to each contract title data. The contract clause data corresponding to each contract title data is the preset clause and the preset chapter of the preset contract template corresponding to the corresponding contract category. Based on the chapter theme model and the clause theme model, obtain the clause theme corresponding to each contract category and the chapter theme corresponding to each clause theme. The clause theme data and the chapter theme data are the preset clause theme and the preset chapter theme of the preset contract template corresponding to the contract category, respectively. Based on the contract category, obtain the corresponding dataset of judgment reasons for that contract category. The dataset of judgment reasons includes at least one data of judgment reasons and at least two data of legal relationship subjects and corresponding subject name data for each data of judgment reasons. Each data of judgment reasons in the dataset of judgment reasons includes at least one data of judicial opinion. Invalid grounds identification model is used to identify and delete invalid grounds in the judgment grounds dataset to obtain valid grounds dataset. Then, the judgment grounds data in the valid grounds dataset is segmented using the judicial opinion segmentation model to obtain judicial opinion training dataset. The judicial opinion data, at least two legal relationship subject data and corresponding subject name data in the judicial opinion training dataset are input into the key information extraction model to obtain key information dataset. The key information data in the key information dataset is input into the behavior consequence differentiation model to obtain the behavior pattern data and legal consequence data corresponding to each key information data. The corresponding behavior pattern data and legal consequence data are added to the key information dataset, and the behavior pattern data in the key information dataset is clustered to obtain at least one behavior pattern set. The behavior pattern data in each behavior pattern set is similar. Calculate the probability that contract clause data belongs to each set of behavioral patterns, take the set of behavioral patterns with the highest probability as the behavioral pattern data to which the contract clause data belongs, and obtain the missing behavioral pattern data in the set of behavioral patterns where the number of elements exceeds a preset threshold and there is no corresponding contract clause data. Based on the missing behavioral pattern data, increase the acquisition of missing contract clause data. Utilize the semantic similarity between the missing contract clause data and the contract clause data in different preset chapters of the preset contract template, add the missing contract clause data to the preset chapter with the highest semantic similarity, and obtain the preset contract template.

2. The method for obtaining key information from a contract template according to claim 1, characterized in that, Based on the aforementioned chapter theme model and clause theme model, clause theme data corresponding to each contract category and chapter theme data corresponding to each clause theme data are obtained, including: For each contract category, chapter theme data is obtained through the chapter theme model; The frequency of the chapter topic data is sorted, and a preset number of chapter topic data with the highest frequency are selected as candidate chapter topics. For each contract category, contract terms are divided, and the contract terms are input into the terms theme model to obtain the terms theme data corresponding to the contract category; For each clause topic data, calculate the frequency of the clause topic data appearing in each candidate chapter topic, and take the chapter topic data with the highest frequency of each clause topic data as the corresponding chapter topic data for that clause topic.

3. The method for obtaining key information from a contract template according to claim 1, characterized in that, The subject matter model of the terms is determined by the following steps: Obtain a training dataset for contract terms, which includes training data for the main text of contract terms and training data for the labels of contract terms. Based on the training data of the contract terms body and the training data of the contract terms labels, contract terms body-label training data is constructed. The lightweight transducer-based bidirectional encoder representation model-bidirectional long short-term memory network-capsule network model is trained using the contract terms body-label training data to obtain the terms topic model.

4. The method for obtaining key information from a contract template according to claim 1, characterized in that, The invalidity reason identification model is determined as follows: Obtain a training dataset of judges' reasons, which includes training data of judges' reasons, including the text of the judges' reasons and the corresponding valid or invalid labels; The training dataset of the judge's reasoning is divided into a training set and a validation set. The training set includes training data and training labels, and the validation set includes validation data and validation labels. The ALBERT-DPCNN model is trained using the training set data and validated and optimized using the validation set to obtain a judgment reason classification model. The judgment reason classification model is used to calculate the probability that a judgment reason belongs to a valid reason. Based on the training data and training labels in the training set, the string length of the adjudication reason after removing commas and periods is obtained. The string length is divided into a first preset number of string length intervals, and the number of words and the probability of the adjudication reason data corresponding to each string length interval being a valid reason are calculated. Based on the training data and training labels in the training set, the number of commas and periods is obtained, and the number of commas and periods is divided into a second preset number of comma and period number intervals. The probability that the punctuation count of the adjudication reason data corresponding to each of the comma and period number intervals is a valid reason is calculated. For each verification data in the verification set, the verification data is input into the adjudication reason classification model to obtain the probability of a valid reason. The probability of validity based on the number of words and the probability of validity based on the number of punctuation marks are calculated. Based on the probability of validity based on the number of words and the probability of validity based on the number of punctuation marks, the verification probability of each verification data being a valid reason is calculated. The parameters of the calculation formula are adjusted based on the verification probability and the verification label.

5. The method for obtaining key information from a contract template according to claim 1, characterized in that, The segmentation model of the judicial opinion is determined as follows: A segmented training dataset of judicial opinions is obtained. The segmented training dataset of judicial opinions contains several reasons for judgment. Each reason for judgment includes at least one judicial opinion. Each judicial opinion is labeled using the BEMS labeling method to obtain labeled training data. The judicial opinion is labeled as B at the beginning, E at the end, M in the middle, and S at the other parts. Based on the labeled training data obtained using the BEMS annotation method, training data conforming to the reading comprehension model is obtained. A named entity recognition model composed of an attention-decoupled enhanced decoding-based bidirectional encoder representation model, a long short-term memory network with vector quantization, and a machine reading comprehension model is trained to obtain the judicial opinion segmentation model. The judicial opinion segmentation model can divide the judgment reasoning into at least one judicial opinion.

6. The method for obtaining key information from a contract template according to claim 1, characterized in that, The key information extraction model is determined as follows: A key information extraction training dataset is obtained, which includes the original text of paired judicial opinions and the key information corresponding to the judicial opinions. The original text of the judicial opinion is input into the generator of the generative adversarial network to obtain the key prediction information; The predicted key information and the key information corresponding to the judicial opinion are input into the discriminator of the generative adversarial network to obtain the similarity between the predicted key information and the key information corresponding to the judicial opinion. The parameters of the generator are adjusted based on the similarity, so that the generator can generate key information data whose text similarity to the key information corresponding to the judicial opinion is greater than a preset threshold.

7. The method for obtaining key information from a contract template according to claim 1, characterized in that, The behavioral consequence differentiation model is determined as follows: Obtain a behavioral consequences training dataset, which includes text and corresponding labels for behavioral patterns or legal consequences; Using the text as input and the corresponding label as output, train the ALBERT-DPCNN model to obtain the behavior consequence discrimination model.

8. A system for obtaining key information from a contract template, characterized in that, include: Input unit, through which the user uploads the contract; The processing unit is configured to implement the method for obtaining key information of the contract template as described in any one of claims 1 to 7, and to calculate the part of the contract uploaded by the user that is missing from the contract template; The output unit is used to prompt the user to add the missing parts of the contract uploaded by the user relative to the contract template.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to execute the method for obtaining key information of the contract template as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Contract risk analysis method and device and storage medium

    CN110211006A

  • File comparison method and device, electronic equipment and storage medium

    CN113901780A