A method and system for generating a contract abstract
By analyzing and analyzing the terms and related tags of the contract document, combining user and system attributes, key terms are extracted to generate a contract summary, solving the problem that the existing technology cannot accurately grasp the core information of the contract, and achieving efficient and accurate contract summary generation.
Patent Information
- Application Number
- CN202510254877.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-03-05
AI Technical Summary
The existing general text summary technology cannot effectively grasp the essence of the contract and the key terms of the transaction, resulting in the generated contract text summary that cannot accurately reflect the core information of the contract.
By obtaining and analyzing contract documents, determining the contract type, and determining the importance of the terms based on the terms content and associated tags, extracting key terms and generating summary.
It realizes intelligent, efficient and accurate extraction of key information in the contract, and generates concise contract summary, improving users' work efficiency in contract drafting, consultation, approval and other links.
Smart Images

Figure CN119760129B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data intelligent processing, and specifically relates to a method and system for generating contract summaries. Background Art
[0002] Contracts often face a challenge in the processes of drafting, negotiation, approval, etc. within an enterprise: contract texts are often long and complex. This not only requires processors to spend a lot of time carefully reading, but also makes it difficult to quickly grasp the essence of the contract. To solve this problem, contract summaries have become an effective tool. Generating summaries by extracting key information from contracts can help processors quickly browse and understand the main points of the contract. Contract summaries not only save valuable time, but also improve work efficiency, ensuring that processors can focus on those parts that are crucial for decision-making. By refining and simplifying information, complex contract texts become easier to manage and understand.
[0003] However, different from ordinary texts, contract texts have extremely strong rigor, professionalism, systematicness, interference, and applicability. Some contract terms are simple and clear, while others are extremely complex, involving numerous details, conditions, and legal terms, etc. These characteristics determine that contract texts must take into account all aspects of the transactions they involve, including both important terms and non-important terms, and there may be correlations between the terms within the contract. Some terms are core terms, and some other terms may be supplementary or restrictive conditions around the core terms.
[0004] Existing general text summarization technologies aim to compress the number of text words, and they cannot grasp the essence of the transactions and key terms of contracts. General text generation summary schemes often focus on all text contents and do not distinguish and process the essence of contract terms, easily resulting in the situation of failing to distinguish between primary and secondary. That is, there are important terms not included in the summary, and there are terms in the summary that are non-important terms or even common sense terms, thus leading to the situation that the generated contract text summary fails to grasp the key points and cannot achieve the purpose of the contract summary. Therefore, existing general text summary generation schemes are not applicable to generating contract text summaries. Summary of the Invention
[0005] Based on this, in view of the above problems, this application proposes a method and system for generating contract summaries, aiming to intelligently, efficiently, and accurately extract key information from contracts and generate contract summaries to facilitate users to quickly grasp the core information of contracts in various links such as contract drafting, negotiation, and approval, and improve work efficiency.
[0006] This application provides a method for generating contract summaries on the one hand. The method includes:
[0007] Obtain a contract document to be processed and determine the contract type of the contract document;
[0008] Analyze the content of each clause in the contract document and assign associated tags to each clause;
[0009] Determine the importance of each clause based on the content of each clause, the contract type, and the associated tags of each clause;
[0010] Extract one or more target clauses according to the user attribute information and / or the current system attribute information and the importance of each clause, and form a target text according to a preset rule;
[0011] Generate an abstract text of the contract document based on the target text.
[0012] Preferably, the method further includes:
[0013] Determine a corresponding subset of clause tags according to the contract type;
[0014] Assign associated tags to each clause based on the content of each clause and the subset of clause tags.
[0015] Furthermore, the method further includes:
[0016] Input the content of each clause in the contract document into a preset large language model and combine it with a pre-constructed first target prompt to assign one or more associated tags to each clause; and / or,
[0017] Assign one or more associated tags to each clause based on a first model pre-trained by a large language model and / or an NLP model.
[0018] Preferably, the determining the importance of each clause according to the contract type and the associated tags of each clause includes:
[0019] Train a historical contract text based on a preset large language model and / or an NLP model to obtain a trained second model;
[0020] Input the clause content of the contract, the associated tags of each clause, and the contract type into the trained second model for importance grading to obtain the importance level and / or quantization value of each clause;
[0021] and / or,
[0022] Query the initial importance level and / or quantization value of each preset clause tag in the contract type, and extract the key feature information of the clause content for strengthening or weakening to obtain the importance level and / or quantization value of each clause.
[0023] Preferably, the training a historical contract text based on a preset large language model and / or an NLP model to obtain a trained second model includes:
[0024] Construct a second initial model based on a preset large language model and / or NLP model;
[0025] Encode the contract type and clause tags to obtain feature vectors;
[0026] Generate clause text vectors according to the clause content, and numerically process the key feature information in the clause content;
[0027] Fuse the clause text vectors, feature vectors and numerically processed feature information to obtain comprehensive feature vectors;
[0028] Divide the preprocessed feature vector data into a training set, a validation set and a test set;
[0029] Input the training set data into the second initial model in batches, and perform multiple rounds of iterative training according to the preset loss function and optimization algorithm, and update the model parameters;
[0030] Adjust the hyperparameters according to the validation set performance; obtain the finally trained second model.
[0031] Preferably, the method further includes:
[0032] Obtain the user's role information, identity information and / or the interface information of the current system, and screen one or more labeled clauses associated with the user and / or the current system interface;
[0033] Determine one or more target clauses according to the importance of each clause, and summarize them in a preset order to obtain the target text.
[0034] Preferably, the method further includes:
[0035] If the current contract is a framework contract, extract the participating objects, terms, signing purposes and / or additional information on the signing purposes and terms in each target clause to obtain the summary element information of each target clause;
[0036] If the current contract is not a framework contract, extract the participating objects, business subjects / transaction essence and / or delivery / settlement information in each target clause to obtain the summary element information of each target clause;
[0037] Summarize the summary element information of each target clause into a summary according to the preset text expression structure.
[0038] Further, the method further includes:
[0039] Input the target text into a preset large language model combined with a pre-constructed second target prompt word to generate corresponding summaries for each clause and generate corresponding summary titles;
[0040] Input the content of each abstract and each contract clause into the large language model and the pre - constructed third - target prompt for semantic verification to obtain the final abstract text.
[0041] The second aspect of this application provides a contract abstract generation system, and the system includes:
[0042] A contract type judgment unit, configured to obtain a contract document to be processed and determine the contract type of the contract document;
[0043] A clause label recognition unit, configured to parse the content of each clause of the contract document and assign associated labels to each clause;
[0044] A clause importance determination unit, configured to determine the importance of each clause according to the content of each clause, the contract type, and the associated labels of each clause;
[0045] A contract abstract generation unit, configured to extract one or more target clauses according to user attribute information and / or current system attribute information and the importance of each clause, form a target text according to a preset rule; and generate an abstract text of the contract document based on the target text.
[0046] The third aspect of this application provides a computer - readable storage medium storing a computer program, which when executed by a processor causes the processor to execute the steps of any one of the above - mentioned methods.
[0047] The fourth aspect of this application provides a computer terminal device, including a memory and a processor, where the memory stores a computer program, and when the computer program is executed by the processor, it causes the processor to execute the steps of any one of the above - mentioned methods.
[0048] The contract abstract generation solution provided above in this application determines the type of the contract text, assigns associated labels according to the content of the contract clauses, then determines the importance of each clause according to the content of each clause, the contract type, and the associated labels of each clause, and further combines user attributes and / or current system attributes to screen out relevant clauses as target texts for generating the contract abstract. The solution of this application can specifically identify the importance of each clause in the current contract text, thereby specifically extracting the corresponding key clause content as the text basis for abstract generation, and further incorporating the contract type judgment, clause label recognition, clause importance judgment, and clause importance screening links into the contract abstract technical solution, achieving the purpose of making the contract abstract concise and to the point, and greatly improving the accuracy of abstract generation.
[0049] Meanwhile, the solution of the present application further combines the attribute information of the user and / or the system to generate corresponding relevant contract summary information, which helps different users quickly, efficiently, and accurately obtain and understand the key contract information associated therewith, and improves the work efficiency of various links such as contract drafting, negotiation, and approval.
[0050] Furthermore, the solution of the present application also uses a pre-trained model algorithm to intelligently identify the importance of each clause. The algorithm not only considers the information of the clause content, but also incorporates information from other dimensions such as the contract type and clause tags. The contract type can reflect the business background and general framework of the clause from a macro level, and the criteria for judging the importance of key clauses often vary for different types of contracts. Clause tags classify and summarize clauses from a micro perspective, which helps to more accurately grasp the functional attributes of the clauses. By comprehensively using input data from multiple dimensions, the model can more comprehensively and accurately identify and extract richer semantic and associated information in the contract, improving the accuracy of contract importance identification. Further, the present application also specifically extracts based on the specific characteristics of the contract type, and combines these characteristics to further enable the model to more comprehensively understand the connotation of clause importance.
[0051] In addition, the present application also determines the essential content and key information of the corresponding summary generation by judging the essence of the contract attributes, and realizes the automatic generation of contract clause summaries through a preset text expression structure and target information extraction technology, which can efficiently and accurately generate standardized summary texts. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0053] Among them:
[0054] Figure 1 is a flowchart of a contract summary generation method in an embodiment;
[0055] Figure 2 is an example diagram of partial data input and output of a clause importance training model in an embodiment;
[0056] Figure 3 is a schematic diagram of intelligent summary generation for a housing lease contract in an embodiment;
[0057] Figure 4 is a structural block diagram of a contract summary generation system in an embodiment;
[0058] Figure 5 It is a structural block diagram of a computer device in an embodiment. Specific implementation manners
[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the protection scope of the present invention.
[0060] The terms "including", "comprising", "having" and any variations thereof in the specification and claims of this application and the above-mentioned accompanying drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices. In the terms in the claims, specification and specification drawings of this application, relational terms such as "first" and "second" are only used to distinguish one entity / operation / object from another entity / operation / object, and do not necessarily require or imply any such actual relationship or order between these entities / operations / objects.
[0061] Referring to "embodiments" herein means that specific features, structures or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase does not necessarily refer to the same embodiment each time it appears in the specification, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0062] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0063] In one embodiment, as Figure 1 shown is a flowchart of a contract summary method of this application, and the method includes:
[0064] S10. Obtain a contract document to be processed and determine the contract type of the contract document.
[0065] The importance of contract terms is not only related to the content of the terms, but also to the type of contract. For example, in a contract for the sale of tangible goods, the price of the goods, delivery time, warranty, etc. are important terms and should appear in the summary so that they are the focus of attention. However, confidentiality clauses and integrity clauses are non-important terms and should not appear in the summary. When the contract type is an intellectual property ownership and confidentiality agreement, the confidentiality clause is an important clause and should appear in the summary so that it is the focus of attention. Therefore, different clauses in the same contract type and the same clauses in different contract types are of different importance and have different priorities for inclusion in the summary.
[0066] Specifically, the contract document to be processed in the present application solution (i.e., the contract document for which a summary needs to be generated) can be a contract document being drafted by the system, or a contract document that has been drafted, or a contract document uploaded to the system during modification, approval, etc. The contract document can be a text file, PDF document, Word document, etc. After obtaining the contract document to be processed, the system first identifies and determines the contract type through the contract type judgment submodule to determine the contract type to which the contract document belongs.
[0067] Preferably, the present application scheme presets a full set of contract types (such as sales, procurement, leasing, services, frameworks, construction projects, entrustment, confidentiality, etc.), and can further subdivide the types under each major category. For example, sales contracts can be divided into goods sales contracts, real estate sales contracts, etc.
[0068] Furthermore, in order to accurately determine the contract type to which the current contract document belongs, in a preferred embodiment of the present application, the contract document is input into a preset large language model, combined with pre-built judgment prompt words, based on the basic text understanding ability and prompt words of the large language model, so that the large language model selects the most appropriate specific contract type from the entire set of contract types for output.
[0069] Preferably, in one embodiment, the present application scheme determines the contract type according to preset text rules or regular expressions: traverse the key nouns, verbs and other vocabulary of the contract, extract one or more high-frequency keywords of the contract, and perform matching and judgment based on the high-frequency keywords and the full set of contract types to identify the contract type of the current contract document. For example, when keywords such as "rent, rent, lessee" appear frequently, it is judged as a lease type contract; when keywords such as "warranty, return, sale, delivery of goods" appear frequently, it is judged as a goods sale contract, etc. It is also possible to combine a pre-built dictionary of contract types and corresponding keywords to determine the contract type through text scanning and matching; or determine the contract type based on the overall framework structure of the contract text.
[0070] In another preferred embodiment, the solution of the present application uses a trained classification model (such as a support vector machine, a decision tree, a convolutional neural network in deep learning, an NLP model, a large model with a decoder-only architecture, etc.), and inputs the contract text data into the model for classification prediction. The training data of the model can be a large number of contract texts with pre-labeled contract types, and the model is allowed to learn the semantic feature patterns corresponding to different contract types through supervised learning. According to the probability or category label output by the classification model, the specific type of the contract is determined. For example, the contract type corresponding to the category with the highest output probability is determined. For example, if the highest probability is "technical service contract", it is determined that the current contract is of this type.
[0071] In addition, in other embodiments, the solution of the present application can also determine the current contract type by or in combination with methods such as title extraction and subject matter type judgment, or by the user's selection.
[0072] S11. Analyze the content of each clause in the contract document and assign associated labels to each clause.
[0073] Specifically, in an embodiment of the present application, after determining the contract type of the contract document, analyze the content of each clause in the contract document, and determine what label should be assigned to each clause of the contract according to the specific clause content and in combination with the contract type. The assigned label can be one or more specific labels, or can be certain class labels (such as legal attribute classes, transaction attribute classes, etc.). Preferably, the assigning of associated labels to each clause further includes:
[0074] S110. Determine the corresponding clause label subset according to the contract type.
[0075] Specifically, the solution of the present application can preset a complete set of contract clause labels (such as subject matter description, subject matter price, delivery method, payment method, liability for breach of contract, etc.), and then screen the complete set of labels according to the contract type to obtain the corresponding clause label subset, or preset the corresponding label subset for each contract type.
[0076] For example, an example of a clause label subset for a commodity sales contract (partial labels): [business subject matter, subject matter price, payment method, delivery method, acceptance method, quality assurance, return and exchange, liability for breach of contract (Party A), liability for breach of contract (Party B), signing method, applicable law, intellectual property, confidentiality...]
[0077] An example of a clause label subset for a house rental contract (partial labels): [rental subject matter, rent, deposit, payment method, delivery method, renewal method, refund method, decoration, liability for breach of contract (Party A), liability for breach of contract (Party B), signing method, applicable law, intellectual property, confidentiality...].
[0078] S111. Assign associated tags to each clause based on the content of each clause and the subset of clause tags.
[0079] Preferably, in one embodiment of the present application, for each clause of the contract, using the basic understanding ability of the large language model combined with the corresponding pre-constructed prompt words, the tags of the clauses are output, and the tags may be one or a combination of multiple ones. Specifically, it includes: inputting the content of each clause of the contract document into a preset large language model and combining the pre-constructed first target prompt words to assign one or more associated tags to each clause. Among them, the first target prompt words relate to assigning associated tags based on the subset of clause tags.
[0080] And / or
[0081] Based on the first model trained by the large language model and / or the NLP model and the subset of clause tags, assign one or more associated tags to each clause. Specifically, based on the first initial model pre-trained by the large language model and / or the NLP model, use the labeled historical clause data to train the first initial model to obtain the trained first model, and then input the content of the current contract clause into the trained first model to assign associated tags to each clause.
[0082] S12. Determine the importance of each clause according to the content of each clause, the contract type, and the associated tags of each clause.
[0083] Specifically, after assigning corresponding tags to each clause, the solution of the present application further grades the importance of each clause according to the specific tags of the clause and the contract type, that is, it can output specific scores (and then conduct importance grading based on the scores), or directly output levels (such as high, medium-high, medium, low, etc.). The determining the importance of each clause according to the contract type and the associated tags of each clause includes:
[0084] Preferably, in one embodiment of the present application, by querying the initial importance level and / or quantization value of each clause tag in the contract type preset, and extracting the key feature information of the clause content to strengthen or weaken, the importance level and / or quantization value of each clause are obtained.
[0085] Specifically, organize the clause tags and contract type data into a structured data set or a preset queryable database, and set corresponding strengthening or weakening weights for the key feature information. The data record includes three input features: the content of the contract clause, the contract type, and the clause tag, as well as the corresponding importance score or quantization weight. Initially, professional legal personnel can score according to experience and relevant regulations, etc., and the range of scores or quantization weights can be set by themselves, or converted into importance levels such as high, medium, low, etc. Then, by matching and querying the corresponding data set and combining the preset calculation, the importance level and / or quantization value of each clause are obtained.
[0086] and / or
[0087] Preferably, in another embodiment of the present application, a traditional NLP model or a large language model is used for training. The model type is a classification model, so that the trained model can distinguish the importance or quantification value of the clauses. Specifically, it includes:
[0088] S121. Train the historical contract text based on a preset large language model and / or NLP model to obtain a pre-trained second model;
[0089] S122. Input the clause content, each clause association label and the contract type of the contract into the pre-trained second model for importance grading to obtain the importance level and / or quantification value of each clause.
[0090] Preferably, the training of the historical contract text based on a preset large language model and / or NLP model to obtain a pre-trained second model includes:
[0091] S1221. Construct a second initial model based on a preset large language model and / or NLP model.
[0092] Specifically, in an embodiment of the present application, a pre-trained large language model and / or NPL model based on the Transformer architecture is used as the basic model, such as BERT, RoBERTa, GPT, etc. This large language model has powerful text semantic understanding ability and has been pre-trained on a large scale of text data, and can capture rich language features and semantic information, which helps to analyze the connotation of contract clauses.
[0093] Alternatively, in the case of certain restrictions on computing resources and training time and not particularly large data scale, some lightweight neural network architectures designed specifically for sequence labeling and text classification tasks are selected, such as bidirectional long short-term memory network (BiLSTM) combined with conditional random field (CRF), etc.
[0094] After constructing the second initial model, initialize the model parameters. After determining the model, set the parameters of the model according to actual needs, including but not limited to one or more of the following parameters: learning rate, batch size, number of iterations, network, number of layers, number of neurons, temperature and other parameters. Specifically, if a pre-trained initial model is selected, its publicly available pre-trained weights can be downloaded and fine-tuned based on these weights to accelerate the convergence speed and utilize the general language knowledge it has learned.
[0095] Further, in an embodiment of the present application, historical data of different contract types and different clause contents with assigned tags are collected, and multi-dimensional evaluation is performed in combination with the actual situation to determine the corresponding importance grading or quantitative score value. Then, preprocessing such as text cleaning and format conversion is performed, and a corresponding training data set is formed. For example, the form of the training data set (partial data) is as follows Figure 2 As shown, the model input data includes: contract clause content, contract clause tags, contract type and other information, and the model output is the importance / score of each clause.
[0096] S1222. Encode the contract text type and clause tags to obtain feature vectors.
[0097] Preferably, in an embodiment of the present application, the tags of each contract text type and each clause in the historical samples are obtained, and one-hot encoding (One-Hot Encoding) or other suitable encoding methods are performed on the contract type and clause tags so that they can be effectively processed by the model.
[0098] S1223. Generate clause text vectors according to one or more words in the clause content, extract key feature information in the clause content, and perform numerical processing.
[0099] Specifically, the contract clause content is text-cleaned to remove interfering information such as extra spaces, abnormal punctuation marks, and special characters. The text is uniformly converted into a suitable encoding format (such as UTF-8), and necessary truncation or padding operations are performed according to the input requirements of the model. For example, if the model has a limit on the input text length, too long clause content is appropriately truncated, and too short content can be padded by adding placeholders, etc.
[0100] Furthermore, a word vector tool, such as Word2Vec, GloVe, etc., which is trained on a large-scale general text corpus or a corpus related to the contract field, is used to map the content words in the contract clause to corresponding low-dimensional vector representations. Further, these word vectors are combined (such as average pooling, weighted average, etc.) to generate a text vector representation of the entire contract clause content.
[0101] Further preferably, in an embodiment of the present application, according to the knowledge in the contract field, some additional key feature information is extracted, such as the amount digital features involved in the terms (whether there are descriptions related to large amounts, the order of magnitude of the amount, etc.), time-related features (whether there is a clear performance period, deadline, etc.), and key entity features (involved important enterprises or matters, etc.), and the key feature information is numerically processed. For example, for the amount digital features and time feature constraints, binary variables can be used, where 1 represents the existence of a large amount or time constraint description, and 0 represents the non-existence. Also, for example, one-hot numerical encoding can be used for the key entity features for numerical encoding, such as the involved important matters or the strictness of the default clause.
[0102] S1224. Integrate the clause text vector, feature vector, and numerically processed feature information to obtain a comprehensive feature vector.
[0103] Specifically, in an embodiment, the numerically processed feature information is concatenated or fused in a suitable manner with the encoded contract type, clause label features, and text vector features to form a comprehensive feature vector for finally inputting into the model.
[0104] S1225. Divide the preprocessed feature vector data into a training set, a validation set, and a test set.
[0105] Specifically, the prepared data set is divided into a training set, a validation set, and a test set according to a suitable ratio (commonly such as 8:1:1 or 7:2:1, etc.). Ensure that in the division process, the clause samples of various contract types and different importance levels / score values are reasonably distributed in each subset to avoid data bias.
[0106] S1226. Input the training set data into the model in batches, perform multiple rounds of iterative training according to the preset loss function and optimization algorithm, calculate the loss value, and update the parameters of the model according to the optimization algorithm.
[0107] Specifically, in an embodiment of the present application, the training set data is input into the model in batches (Batch), and multiple rounds (Epoch) of iterative training are performed according to the set loss function and optimization algorithm. Among them, the optimization algorithms include stochastic gradient descent (SGD) and its variants Adagrad, Adadelta, Adam, etc.
[0108] In an embodiment of the present application, a suitable loss function is selected according to the nature of the output (score value or level). If the output is a specific score value, the mean squared error (MSE) or other forms of loss functions can be selected. It measures the average squared error between the predicted score and the true score and can effectively guide the model to learn accurate numerical mappings.
[0109] If the importance level is output, a preset cross-entropy loss function (Cross-Entropy Loss) can be used to regard the level as a classification category, so as to prompt the model to correctly classify and grade the terms of different levels. Preferably, in one embodiment, the cross-entropy loss function used in this application satisfies:
[0110] ,
[0111] Among them, Loss is the cross entropy loss function, The encoding vector corresponding to the true category of the jth sample is (dimension is C), The probability distribution vector of each category predicted by the model is (the dimension is also , and each element value is between 0 and 1, and the sum is 1 ). In other embodiments, those skilled in the art can also adjust the corresponding loss function according to actual needs.
[0112] S1227. Adjust the hyperparameters according to the performance of the validation set to obtain the final pre-trained second model.
[0113] Specifically, in each training round, the loss value is calculated and the parameters of the model are updated according to the optimization algorithm. At the same time, the performance of the model is regularly evaluated on the validation set (such as calculating the loss value, accuracy and other indicators on the validation set. The accuracy can be calculated according to the proportion of correctly classified samples based on the level output). The hyperparameters (such as learning rate, batch size, number of training rounds, etc.) are adjusted according to the performance of the validation set to prevent overfitting.
[0114] When determining the importance of each clause, the above scheme of this application not only considers the core text information of the clause content, but also incorporates additional dimensional information such as the contract type and clause label. The contract type can reflect the business background and general framework of the clause from a macro level. The criteria for judging the importance of key clauses in different types of contracts often vary; the clause label classifies and summarizes the clauses from a micro perspective, which helps to grasp the functional attributes of the clauses more accurately. By comprehensively utilizing these three aspects of input data, the model can more comprehensively understand the connotation of the importance of the clauses, and can mine richer semantics and related information than methods that rely solely on single text features.
[0115] Furthermore, we have also specially designed the extraction of specific key feature information based on contract domain knowledge, such as amount numerical features, time-related features, etc. These domain features play an intuitive and key role in judging the importance of contract terms. For example, terms involving large amounts of money are usually more important, and performance terms with clear time limits cannot be ignored for corresponding types of contract terms. In this way, the importance of the corresponding terms can be effectively strengthened or weakened through the above-mentioned specific key feature information, making the output results more accurate and more in line with expectations.
[0116] S13. Extract one or more target clauses according to the user attribute information and / or the current system attribute information and the importance of each clause, and form a target text according to a preset rule.
[0117] Specifically, in the solution of this application, the user's attribute information includes user role information (such as Party A, Party B, etc.), user identity (such as business, finance, legal affairs, administration, etc.), and the current system attribute includes the interface attribute information of the current system. For example, in the collaborative interface, approval interface, archiving interface, etc. in the system, the purposes of the user are different under different interfaces, and the generated abstracts are also different.
[0118] Preferably, in an embodiment of this application, extracting one or more relevant target clauses to form a target text includes:
[0119] S131. Obtain the user's role information, identity information and / or the interface information where the current system is located, and screen the set of labeled clauses associated with the user and / or the current system interface.
[0120] Preferably, in the solution of this application, first obtain the user's role information and identity information, and / or obtain the functional interface where the current system is located, and screen the labeled clauses related to the user and / or the current functional interface according to the user role, user identity and / or the current system interface to obtain a clause set including one or more labeled clauses, where the clause labels associated with the user role and user identity are recorded in the policy table and can be quickly queried. For example, the financial staff of Party B is concerned about the clauses corresponding to the labels such as "payment method" and "liability for breach of contract for overdue payment", and the business staff of Party A is concerned about the clauses corresponding to the labels such as "delivery time" and "acceptance method".
[0121] Furthermore, the solution of this application may further include: adjusting the importance of each clause through the user's attribute information and / or the current system attribute information, such as setting a corresponding adjustment coefficient to adjust the importance level and / or quantization value obtained above.
[0122] S132. Determine one or more target clauses from the set of labeled clauses according to the importance of each clause, and summarize them in a preset order to obtain a target text.
[0123] Specifically, in an embodiment of this application, after obtaining the user-related labeled clause set, screen the clauses obtained in the previous step from high to low according to the clause importance. For example, one or more target clauses with an importance greater than a certain threshold or a certain level can be screened, and further summarized, merged and / or sorted according to a preset rule according to the clause labels to obtain a target text as the input text for abstract generation.
[0124] S14. Generate the abstract text of the contract document based on the target text.
[0125] The solution of this application determines and extracts the explicit content and / or emphasized content information of each target clause as the target element information for generating the abstract of each clause, and then combines the target element information of each clause with the attributes of the contract (framework contract, non-framework contract, etc.) to generate the abstract according to the preset abstract text expression structure.
[0126] S141. Judge the attribute of the contract text. If it is a framework contract, determine the participating objects, term, signing purpose, and / or additional information of the signing purpose and term corresponding to each target clause to obtain the element information for generating the abstract of each clause.
[0127] In one embodiment, for a framework contract, which has no substantial transaction content, therefore, the solution of this application forms the target element information for generating the abstract by extracting or clarifying the information including one or more of the following in each target clause: participating objects, term, signing purpose, additional information of the signing purpose and term, etc.
[0128] S142. If it is a non-framework contract, determine the participating objects, business subject / transaction essence, and / or delivery / settlement information in each target clause to obtain the element information for generating the abstract of each clause.
[0129] For a non-framework contract, which is a formal contract with substantial transaction content, the solution of this application forms the target element information for generating the abstract by extracting or clarifying the information including one or more of the following in each clause: participating objects, business subject, transaction essence, and / or delivery / settlement information, etc.
[0130] S143. Generate the abstract according to the preset text expression structure for the target information for generating the abstract of each clause.
[0131] Specifically, a configuration file or database can be used to manage the preset text expression structure or template, such as setting one or more expression combinations or templates for framework and non-framework types. Then, based on the target information for generating the abstract of each clause extracted, generate the abstract according to the corresponding text expression structure or template.
[0132] The above solution of this application clarifies and determines the necessary content and key information for generating the corresponding abstract by judging the essence of the contract attribute. Through the preset text expression structure and target information extraction technology, the automatic generation of the contract clause abstract can be realized. This method combines rule and machine learning technologies, can efficiently and accurately generate a standardized abstract text, making the generated abstract concise and to the point.
[0133] Further, after obtaining the target text, the solution of the present application inputs the target text into a preset large language model in combination with a pre-constructed second target prompt word to generate abstracts corresponding to each clause, and generates corresponding abstract titles, where the titles can be generated based on the target information when generating the abstracts.
[0134] Further, considering the hallucination of the large model, the abstracts and the content of each contract clause are input into the large language model and the pre-constructed third target prompt word again for semantic verification to obtain the final abstract text. The third target prompt word involves constructing semantic content such as "verifying the results of each abstract and the content of the contract clause to determine whether it conforms to the original meaning or whether the abstract needs to be modified".
[0135] As Figure 3 shown, it is a schematic diagram of the intelligent generation of an abstract for a house lease contract by the above solution in an embodiment of the present application. It can be seen from the generation result diagram that the solution of the present application can accurately generate a concise and to-the-point contract abstract related to the user.
[0136] Through the above solution of the present application, the importance of each clause in the current contract text is intelligently identified in a targeted manner, and the corresponding key clause content is extracted in a targeted manner as the text basis for abstract generation. Further, the contract type judgment, clause label recognition, clause importance judgment, and clause importance screening links are incorporated into the contract abstract technical solution, achieving the purpose of making the contract abstract both concise and to-the-point, and greatly improving the accuracy of abstract generation. At the same time, the solution of the present application further combines the attribute information of the user and / or to generate contract abstract information related to the user and / or the current function interface, which helps different users quickly, efficiently, and accurately obtain and understand the contract key information associated with them, and improves the work efficiency of various links such as contract drafting, negotiation, and approval.
[0137] In one embodiment, as Figure 4 shown, it is a structural block diagram of a contract abstract generation system provided by the present application. The system includes:
[0138] A contract type judgment unit, configured to obtain a contract document to be processed and determine the contract type;
[0139] A clause label recognition unit, configured to parse the content of each clause of the contract document and assign associated labels to each clause;
[0140] A clause importance determination unit, configured to determine the importance of each clause according to the content of each clause, the contract type, and the associated labels of each clause;
[0141] A contract abstract generation unit is configured to extract one or more target clauses according to user attribute information and / or current system attribute information and the importance of each clause, form a target text according to a preset rule; and generate an abstract text of the contract document based on the target text.
[0142] In one embodiment, the present application further provides a computer-readable storage medium storing a computer program, which when executed by a processor, causes the processor to perform the following steps:
[0143] Obtain a contract document to be processed and determine the contract type of the contract document;
[0144] Parse the content of each clause of the contract document and assign an associated label to each clause;
[0145] Determine the importance of each clause according to the content of each clause, the contract type and the associated label of each clause;
[0146] Extract one or more target clauses related to the user according to the user attribute information and the importance of each clause to form a target text;
[0147] Generate an abstract text of the contract document based on the target text.
[0148] In one embodiment, as Figure 5 shown, the present application further provides a computer device including a memory and a processor, the memory stores a computer program, which when executed by the processor, causes the processor to perform the following steps:
[0149] Obtain a contract document to be processed and determine the contract type of the contract document;
[0150] Parse the content of each clause of the contract document and assign an associated label to each clause;
[0151] Determine the importance of each clause according to the content of each clause, the contract type and the associated label of each clause;
[0152] Extract one or more target clauses related to the user according to the user attribute information and the importance of each clause to form a target text;
[0153] Generate an abstract text of the contract document based on the target text.
[0154] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0155] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0156] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A method for generating a contract summary, characterized in that: The method comprises: Obtaining a contract document to be processed and determining a contract type of the contract document; Parsing the contents of each clause of the contract document and assigning associated tags to each clause; Determine the importance of each clause based on its content and the tags associated with it; Extract one or more target clauses according to user attribute information and / or current system attribute information and the importance of each clause, and form a target text according to preset rules; Generate a summary text of the contract document based on the target text; The determination of the importance of each clause includes: Training the historical contract text based on a preset large language model and / or NLP model to obtain a trained second model; Inputting the content of the clauses of the contract, the labels associated with each clause, and the contract type into the trained second model for importance grading to obtain the importance level and / or quantitative value of each clause; and / or, The initial importance level and / or quantitative value of each preset clause label in the contract type is queried, and the key feature information of the clause content is extracted to strengthen or weaken it, so as to obtain the importance level and / or quantitative value of each clause.
2. The method according to claim 1, characterized in that The method further comprises: Inputting the contents of each clause of the contract document into a preset large language model and combining it with a pre-built first target prompt word to assign one or more associated tags to each clause; and / or, Based on the first model trained with a large language model and / or an NLP model, one or more associated tags are assigned to each term.
3. The method according to claim 1, characterized in that The method further comprises: Determine a corresponding clause label subset according to the contract type; Each clause is assigned an associated label based on the content of each clause and the clause label subset.
4. The method according to claim 1, characterized in that: The training of the historical contract text based on the preset large language model and / or NLP model to obtain the second training model includes: Building a second initial model based on a preset large language model and / or NLP model; Encode the contract type and clause label to obtain a feature vector; Generate a clause text vector based on the clause content, and digitize the key feature information in the clause content; The clause text vector, feature vector and numerical feature information are integrated to obtain a comprehensive feature vector; Divide the preprocessed comprehensive feature vector data into training set, validation set and test set; Input the training set data into the second initial model in batches, perform multiple rounds of iterative training according to the preset loss function and optimization algorithm, and update the parameters of the model; The hyperparameters are adjusted according to the performance of the validation set to obtain the final trained second model.
5. The method according to claim 1, characterized in that The method further comprises: Obtaining the user's role information, identity information and / or the current system interface information, and filtering one or more tag terms associated with the user and / or the current system interface; One or more target clauses are determined according to the importance of each clause, and they are summarized in a preset order to obtain the target text.
6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: If the current contract is a framework contract, extract the additional information of the participants, term, contract purpose and / or contract purpose and term in each target clause to obtain the summary element information of each target clause; If the current contract is not a framework contract, extract the participating objects, business subject / transaction substance and / or delivery / settlement information in each target clause to obtain the summary element information of each target clause; The summary element information of each objective clause is summarized according to the preset text expression structure.
7. The method according to claim 6, characterized in that The method further comprises: Input the target text into the preset large language model and combine it with the pre-built second target prompt words to generate the corresponding summary of each clause and the corresponding summary title; The content of each summary and each contract clause is input into the large language model and the pre-built third target prompt word for semantic verification to obtain the final summary text.
8. A contract summary generation system, characterized in that: The system comprises: A contract type determination unit, used to obtain a contract document to be processed and determine the contract type of the contract document; A clause label recognition unit, used to parse the contents of each clause of the contract document and assign an associated label to each clause; A clause importance determination unit, used to determine the importance of each clause according to the content of each clause, the contract type and the associated tags of each clause; A contract summary generating unit, configured to extract one or more target clauses according to user attribute information and / or current system attribute information and the importance of each clause, and form a target text according to a preset rule; and generate a summary text of the contract document based on the target text; The determination of the importance of each clause includes: Training the historical contract text based on a preset large language model and / or NLP model to obtain a trained second model; Inputting the content of the clauses of the contract, the labels associated with each clause, and the contract type into the trained second model for importance grading to obtain the importance level and / or quantitative value of each clause; and / or, The initial importance level and / or quantitative value of each preset clause label in the contract type is queried, and the key feature information of the clause content is extracted to strengthen or weaken it, so as to obtain the importance level and / or quantitative value of each clause.
9. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 7.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Legal document key information extraction method and system
CN118520881A