Method, apparatus and device for converting natural language contracts into intelligent legal contracts

By inserting low-rank matrix layer and designing comprehensive loss functions in the pre-trained model, the modeling limitations in the field of intelligent legal contract code are solved, the modeling efficiency is improved and the cost is reduced, and efficient conversion of natural language contracts is achieved.

CN119902751BActive Publication Date: 2025-07-04CENTRAL UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510087627.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-07-04
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

The limitations of equivalent modeling, low modeling efficiency and high modeling cost in the field of smart legal contract code in the prior art.

Method used

By building a corpus of specific fields, inserting a low-rank matrix layer into the encoder of the basic pre-trained model, and designing a comprehensive loss function, fine-tuning the big model parameters based on the semantic and structural differences between natural language contracts and intelligent legal contracts, and iterative training is used using the Seq2Seq framework and Transformer architecture.

Benefits of technology

It improves the modeling efficiency of smart legal contract code, reduces modeling costs, enhances the understanding of legal terms and regulatory logic in natural language contracts, and promotes cross-domain collaboration between technicians and legal personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119902751B_ABST
    Figure CN119902751B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of intelligent legal contracts, and particularly to a method, device and equipment for converting natural language contracts into intelligent legal contracts. The method includes: generating training corpus for a target large model according to natural language contract data and business process data; generating an input sequence for a fine-tuning framework according to the training corpus and intelligent legal contract code, and iteratively training the target large model under the fine-tuning framework by using the input sequence, wherein the fine-tuning framework includes a decoder, an encoder and a comprehensive loss function, adding low-rank matrix layers to multiple parameter weight matrices of the encoder respectively, freezing some model parameters of the target large model, and adjusting the model parameters of the target large model based on the low-rank matrix layers and the comprehensive loss function; realizing the intelligent legal contract modeling function by using the trained target large model. Thereby, the problems of equivalent modeling limitations, low modeling efficiency and high modeling cost in the field of intelligent legal contract code in the related art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent legal contracts, and particularly to a method for converting natural language contracts into intelligent legal contracts. Background Art

[0002] Intelligent legal contracts are established between intelligent contracts and natural language contracts, having both the legal characteristics and comprehensibility of contracts and the normativity of computer program codes. Modeling natural language contracts as intelligent legal contracts not only helps legal personnel monitor the performance risks of contracts, but also enables computer personnel to understand legal logic and behaviors, promoting cross-domain collaboration between technical and legal personnel. However, the professionalism of the terms, legal rigor, and scenario complexity of natural language contracts themselves pose challenges such as difficult rule conversion, complex system construction, and high maintenance costs in modeling them as intelligent legal contracts.

[0003] Related technologies achieve equivalent modeling in the field of natural language contracts based on rule-based modeling methods and symbolic modeling and formalization methods. The rule-based modeling method defines a set of rules and conventions for automatically processing the terms of a contract and the relationships between terms, and uses technologies such as decision tables and state machines to represent the logic and conditions of contract terms. The method based on symbolic modeling and formalization converts the terms and logical relationships of a contract into a formal language or symbolic representation for precise verification and execution.

[0004] However, the rule-based modeling method has poor flexibility and adaptability. When the business involved in natural language contracts changes, the original rules often need to be extensively modified and redefined to be adapted; the method based on symbolic modeling and formalization requires professional technical knowledge, has a high learning threshold, and has high costs and low efficiency in processing natural language contracts with high complexity. Summary of the Invention

[0005] The present invention provides a method, device, and equipment for converting natural language contracts into intelligent legal contracts to solve problems such as limitations in equivalent modeling, low modeling efficiency, and high modeling costs in the field of intelligent legal contract codes in related technologies.

[0006] An embodiment of the first aspect of the present invention provides a method for converting a natural language contract into an intelligent legal contract, including the following steps: obtaining natural language contract data and business process data expressing the business process model of the natural language contract; generating training corpus for the target large model according to the natural language contract data and the business process data; generating an input sequence of the fine-tuning framework according to the training corpus and the intelligent legal contract code, and iteratively training the target large model under the fine-tuning framework by using the input sequence, where the fine-tuning framework includes a decoder, an encoder, and a comprehensive loss function, adding low-rank matrix layers to multiple parameter weight matrices of the encoder, freezing some model parameters of the target large model, and adjusting the model parameters of the target large model based on the low-rank matrix layers and the comprehensive loss function; implementing an intelligent legal contract modeling function by using the trained target large model, where the intelligent legal contract modeling function is to model a natural language contract into an intelligent legal contract.

[0007] Optionally, generating the training corpus for the target large model according to the natural language contract data and the business process data includes: performing format conversion and data cleaning on the natural language contract data and the business process data; identifying various clauses and elements in the natural language contract data, establishing a mapping relationship between various clauses and elements and various elements in the business process data, and forming training samples according to the natural language contract data, the business process data, and the mapping relationship; splitting the training samples into a training set, a validation set, and a test set, and generating the training corpus for the target large model according to the training set, the validation set, and the test set.

[0008] Optionally, generating the input sequence of the fine-tuning framework according to the training corpus and the intelligent legal contract code includes: converting the natural language contract data and the training corpus into the input format of the fine-tuning framework; performing word segmentation on the intelligent legal contract code corresponding to the natural language contract data and the training corpus; loading the target large model under the fine-tuning framework, marking the word tokens after word segmentation according to the tokenizer of the target large model, and assigning a unique index to the marked word tokens; encoding the marked word tokens into a vector matrix through the embedding layer of the fine-tuning framework, and using the vector matrix as the input sequence of the fine-tuning framework.

[0009] Optionally, the attention mechanism module and the fully connected feed-forward module of the encoder, and the multiple parameter weight matrices include the first to third parameter weight matrices, where the attention weight of the attention mechanism module is determined based on the interaction of the first parameter weight matrix and the second parameter weight matrix, the third attention weight is the integrated information weight of the attention mechanism module, and the features of the natural language contract data and the business process data are assisted in learning through the three parameter weight matrices.

[0010] Optionally, the comprehensive loss function includes a semantic consistency loss, a structural loss, and a semantic and structural alignment loss, where the formula of the comprehensive loss function is:

[0011] L = λ1L token + λ2L s + λ3L align

[0012] Wherein, L represents the value of the loss function, λ1, λ2, and λ3 represent weight parameters, and L token represents the semantic consistency loss, and L s represents the structural loss, and L align represents the semantic and structural alignment loss.

[0013] Optionally, iteratively training and fine-tuning the target large model in the fine-tuning framework using the input sequence includes: setting the hyperparameters of the fine-tuning framework and starting the fine-tuning of the target large model; iteratively training the target large model using the input sequence to calculate the value of the loss function during the training process according to the comprehensive loss function, and optimizing the model parameters of the target large model according to the loss value; after meeting the iterative stop condition, completing the iterative training of the target large model.

[0014] Optionally, the encoder includes a self-attention module and a fully connected feed-forward module, wherein a low-rank matrix layer is added after the self-attention module.

[0015] Optionally, the fine-tuning framework includes a Seq2Seq framework, and the decoder and encoder are based on the Transformer architecture.

[0016] An embodiment of the second aspect of the present invention provides a natural language contract to intelligent legal contract conversion device, including: an acquisition module for acquiring natural language contract data and business process data of a business process model expressing the natural language contract; a processing module for generating training corpus for the target large model according to the natural language contract data and the business process data; a fine-tuning module for generating an input sequence of the fine-tuning framework according to the training corpus and the intelligent legal contract code, and iteratively training the target large model in the fine-tuning framework using the input sequence, wherein the fine-tuning framework includes a decoder, an encoder, and a comprehensive loss function, adding low-rank matrix layers to multiple parameter weight matrices of the encoder respectively, freezing some model parameters of the target large model, and adjusting the model parameters of the target large model based on the low-rank matrix layers and the comprehensive loss function; a deployment module for implementing the intelligent legal contract modeling function using the trained target large model, wherein the intelligent legal contract modeling function is to model the natural language contract into an intelligent legal contract.

[0017] An embodiment of the third aspect of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the program to implement the natural language contract to intelligent legal contract conversion method as described in the above embodiment.

[0018] Thus, the present invention has the following beneficial effects:

[0019] In the embodiments of the present invention, a corpus in a specific field (natural language contract - intelligent legal contract) is constructed. LoRA layers are inserted near three weight matrices in the self-attention module of the basic pre-trained model encoder, and a loss function is designed based on the semantic loss and structural loss between natural language contracts and intelligent natural language contracts to fine-tune the parameters of the large model, enabling the large model to better understand legal terms and regulatory logics in natural language contracts, thereby optimizing the intelligent legal contract code modeling process. Thus, the technical advantages of the large model, such as its complex structure understanding and capture ability, cross-domain dynamic adaptation and generalization ability, and efficient modeling and processing ability, can be utilized to address the limitations of equivalent modeling in the field of intelligent legal contract codes, improve the modeling efficiency, and reduce the modeling cost.

[0020] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0022] Figure 1 is a flowchart of a method for converting a natural language contract into an intelligent legal contract according to an embodiment of the present invention;

[0023] Figure 2 is the expiration date and renewal clause according to an embodiment of the present invention

[0024] Figure 3 is an example diagram of the distributor liability clause according to an embodiment of the present invention;

[0025] Figure 4 is an example diagram of the service confirmation clause according to an embodiment of the present invention;

[0026] Figure 5 is a flowchart of a method for converting a natural language contract into an intelligent legal contract according to an embodiment of the present invention;

[0027] Figure 6 is a block diagram of a device for converting a natural language contract into an intelligent legal contract according to an embodiment of the present invention;

[0028] Figure 7 is a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.

[0030] For the technical terms used in the following embodiments, the explanations are as follows:

[0031] Intelligent legal contract: It is a computer program containing the elements of a contract and covering the performance agreement reached by the contracting parties according to offers and acceptances. It is a transitional means between real natural language contracts and smart contracts.

[0032] Rule-based modeling: It is a modeling method that describes the behavior, constraints, and logical relationships of a system by defining a series of rules. These rules usually take the form of "if... then..." and are used for automated decision-making, reasoning, or behavior selection. The rule-based modeling method is widely applied in multiple fields, including expert systems, natural language contract modeling, decision support systems, etc.

[0033] Symbolic modeling and formal method-based: It is a modeling method that describes complex systems, concepts, or problems in a symbolic way and uses a formal language to represent and analyze them precisely. This method relies on precise syntax, semantics, and reasoning rules and can help solve problems such as system verification, optimization, reasoning, and proof.

[0034] Large model fine-tuning: Fine-tuning usually adjusts the parameters of the model with a small amount of specific task data on the basis that the pre-trained model has learned general knowledge, so that it can better meet the specific task requirements.

[0035] Modeling natural language contracts as intelligent legal contracts helps legal personnel monitor contract performance risks, helps computer personnel understand legal logic and behavior, promotes cross-domain collaboration between technical and legal personnel, and makes the contract processing process more efficient, transparent, and intelligent.

[0036] Since large models have technologies such as good understanding and capture capabilities for complex structures, cross-domain dynamic adaptation and generalization capabilities, and efficient modeling and processing capabilities, the present invention proposes a method for converting natural language contracts into intelligent legal contracts, constructs a corpus for a specific domain to fine-tune the parameters of the large model, enables the large model to better understand legal terms and regulatory logic in natural language contracts, and further optimizes the code modeling process of intelligent legal contracts.

[0037] The following describes the method, device, and electronic device for converting natural language contracts into intelligent legal contracts according to the embodiments of the present invention with reference to the accompanying drawings. Specifically, Figure 1Schematic diagram of a method for converting a natural language contract into an intelligent legal contract according to an embodiment of the present invention.

[0038] As Figure 1 shown, the method for converting a natural language contract into an intelligent legal contract includes the following steps:

[0039] In step S101, natural language contract data and business process data representing the business process model of the natural language contract are obtained.

[0040] It can be understood that in the embodiments of the present invention, various types of natural language contracts and BPMN business process data representing natural language contracts can be obtained through various channels such as legal databases, open source projects, and GitHub. Among them, the natural language contract data includes various types of natural language contracts.

[0041] In step S102, training corpus for the target large model is generated based on the natural language contract data and the business process data.

[0042] In the embodiments of the invention, generating the training corpus for the target large model based on the natural language contract data and the business process data includes: performing format conversion and data cleaning on the natural language contract data and the business process data; identifying various types of clauses and elements in the natural language contract data, establishing a mapping relationship between various types of clauses and elements and various types of elements in the business process data, and forming training samples according to the natural language contract data, the business process data, and the mapping relationship; splitting the training samples into a training set, a validation set, and a test set, and generating the training corpus for the target large model according to the training set, the validation set, and the test set.

[0043] It can be understood that in the embodiments of the present invention, after collecting a large number of various types of natural language contracts, various types of clauses and elements in the contract are determined and mapped to various types of elements of BPMN. Through data processing methods such as data cleaning, format conversion, alignment, and splitting, high-quality input data is obtained.

[0044] Specifically, S11 preprocesses and cleans the natural language contract and the BPMN business process data to remove redundant data.

[0045] S12 data alignment, establishing a one-to-one correspondence between the natural language contract clauses and their corresponding BPMN models to form standardized training samples.

[0046] In the process of identifying and determining the key terms of a contract, natural language contracts of different types and covering a wide range of business fields are analyzed, and typical terms that are common to these contracts, highly representative, and can reflect the core essence of the contract are accurately extracted. For the classification of core terms, the classification of key terms summarized from multiple contracts in the CUAD dataset can be referred to. The corpus of mapping a single key term to BPMN elements and mapping a combination of multiple key terms to BPMN includes: such as Figure 2 the validity period and renewal clause shown as Figure 3 the distributor liability clause shown as Figure 4 the service confirmation clause shown as

[0047] Thus, in the data processing stage, the embodiments of the present invention can, based on the accurate analysis of the content of the contract text, combine the theories of law and data science to label the original text document with elements such as the contract subject, validity period, default clause, applicable law, etc.

[0048] In step S103, an input sequence of the fine-tuning framework is generated according to the training corpus and the intelligent legal contract code, and the target large model under the fine-tuning framework is iteratively trained using the input sequence. Among them, the fine-tuning framework includes a decoder, an encoder, and a comprehensive loss function. Low-rank matrix layers are respectively added to multiple parameter weight matrices of the encoder, some model parameters of the target large model are frozen, and the model parameters of the target large model are adjusted based on the low-rank matrix layers and the comprehensive loss function.

[0049] Among them, the fine-tuning framework may include a Seq2Seq framework, and the decoder and encoder are based on the Transformer architecture.

[0050] It can be understood that the embodiments of the present invention collect a large number of natural language contracts and intelligent legal contract element annotation corpora of multiple types, and design corresponding loss functions and insert LoRA layers at appropriate positions according to the structural differences between natural language contracts and intelligent natural language contracts, and continuously experiment and optimize through parameter-efficient fine-tuning.

[0051] Therefore, based on the principle of low-rank approximation, LoRA can learn only a small number of additional parameters without changing the structure of the original model too much. LoRA can utilize the language knowledge and semantic representations already learned in the pre-trained model and only fine-tune a small number of parameters related to the specific vocabulary and clause logic of natural language contracts, thereby improving the model's understanding and processing ability of natural language contract content; a loss function is designed for the semantic and structural differences between natural language contracts and intelligent natural language contracts to enable the large model to better correspond natural language contracts with intelligent legal contracts; the solution for fine-tuning the large language model for intelligent legal contract modeling avoids the complex grammar and strict rules of rule-based modeling and symbolic modeling languages and formal methods, and improves the efficiency of complex legal contract conversion.

[0052] In the embodiment of the present invention, as Figure 5 shown, the encoder includes a self-attention module and a fully-connected feed-forward module. When computing resources permit, in addition to adding LoRA layers near the three weight matrices W q , W k , W v of the self-attention module of the encoder, it is possible to consider adding LoRA layers near the cross-attention mechanism of the decoder to further learn the complex logic and structural differences between legal contracts and intelligent legal contracts.

[0053] In the embodiment of the present invention, an input sequence of the fine-tuning framework is generated according to the training corpus and the intelligent legal contract code, including: converting natural language contract data and the training corpus into the input format of the fine-tuning framework; performing word segmentation on the intelligent legal contract code and the training corpus corresponding to the natural language contract data; loading the target large model into the fine-tuning framework, marking the word tokens after word segmentation according to the tokenizer of the target large model, and assigning a unique index to the marked word tokens; encoding the marked word tokens into a vector matrix through the embedding layer of the fine-tuning framework, and using the vector matrix as the input sequence of the fine-tuning framework.

[0054] It can be understood that the target large model can be a large language model, such as LegalBert. During the process of fine-tuning the large language model, for example, it can be fine-tuned based on LegalBert, which can better understand legal language and optimize legal logic. Compared with the related technology where fine-tuning the basic language model may not correctly understand legal language and errors may occur during the process of extracting elements, the large language model fine-tuned in the embodiment of the present invention can accurately extract the key clauses in natural language contracts, perform intelligent legal contract modeling, and prevent potential legal risks.

[0055] Specifically, the corpus is converted into a format suitable for input under the Seq2Seq framework to construct a good input sequence, as follows:

[0056] S2.1: For natural language contracts and BPMN corpora, first convert them into a format suitable for input to the Seq2Seq framework.

[0057] S2.2 Load the basic pre-trained model into the fine-tuning framework, and use a tokenizer compatible with LegalBERT to tokenize the segmented Solidity and BPMN corpora. This process will convert the tokens into corresponding markers and assign a unique index to each marker.

[0058] Table 1 compares two target large models, RoBERTa and LegalBERT. Generally speaking, choosing LegalBERT as the target large model can well understand the logical relationship between legal terms and related complex contracts because this model is pre-trained with legal literature and legal-specific corpora.

[0059] Table 1

[0060]

[0061] S2.3: Encode these markers into vector representations through the embedding layer. These vectors will serve as the basic units for subsequent model input and contain the semantic and positional information of the tokens.

[0062] S2.4: Add a decoder based on the Transformer structure under the seq2seq framework.

[0063] In the embodiment of the present invention, a low-rank matrix is added at an appropriate position, and a loss function is designed. The comprehensive loss function includes semantic consistency loss, structural loss, and semantic and structural alignment loss. The attention mechanism module and the fully connected feed-forward module of the encoder, and multiple parameter weight matrices include the first to third parameter weight matrices. Among them, the attention weights of the attention mechanism module are determined based on the interaction of the first parameter weight matrix and the second parameter weight matrix, and the third attention weight is the integrated information weight of the attention mechanism module. The three parameter weight matrices assist in learning the characteristics of natural language contract data and business process data.

[0064] Specifically, a low-rank matrix is added at an appropriate position, and a loss function is designed to prepare for fine-tuning, as follows:

[0065] S31: Taking LegalBERT as an example, apply the LoRA layer to W in the self-attention module of LegalBERT q , W k , W v Compared with only one parameter weight matrix in the related technology, for W q , W k , W vAdd LoRA layers. When processing natural language contracts and BPMN data, this can perform more refined adjustments to the self-attention mechanism. The three parameter weight matrices in the self-attention mechanism are beneficial for better learning of the characteristics of natural language contract data and business process data. W q , W k interactions determine the attention weights and affect the degree of attention to information at different positions. W v is then used to integrate information. Adding LoRA layers separately allows the model to perform independent fine-tuning on these different functional dimensions for the complex semantic relationships in natural language contracts and the process structures of BPMN, in order to better learn the characteristics of the data.

[0066] S32: Based on the differences between natural language contracts and BPMN data, comprehensively consider from three aspects: semantic consistency loss, structural loss, and semantic-structure alignment loss, and design a comprehensive loss function.

[0067] Calculating the semantic consistency loss is mainly to ensure that the generated BPMN representation can correctly reflect the semantics of the input natural language contract. Using the cross-entropy loss function as the basic function to measure the difference between each token generated by the model and the tokens of the target BPMN, then the cross-entropy loss function is where, y t is the t-th token in the target BPMN sequence, and P(y t |y<t,x) represents the probability of the token y t appearing at time step t given the input natural language contract.

[0068] When considering the structural loss, it is necessary to ensure that the generated BPMN's json data has the correct hierarchy and structure. This solution measures the difference by comparing whether the two BPMN.JSON files are consistent from the root node to the key elements. Then the structural loss function is: where, match(p1,p2) represents the consistency between p1 and p2, and L s =0 means that the paths of all activity elements match and the structural similarity is high. The closer L s is to 1, the greater the path difference and the lower the structural similarity. Among them, p a1 is the set of paths of all activity elements from the middle to J1, and p a2 is the set of paths of all activity elements from the middle to J2.

[0069] When considering the semantic-structure loss, use the embedding representation of the generated BPMN (through the vector of the last layer of the model) and the embedding representation of the input contract (through the output of the encoder) to calculate the contrast loss: L align =1 - cos(Z contract , z BPMN ), where, zcontract The vector representation representing the contract embedding, z BPMN The vector representation representing the BPMN embedding, cos(z contract , Z BPMN ), uses cosine similarity to measure z contract , z BPMN The similarity between them.

[0070] Therefore, the comprehensive loss function is:

[0071] L = λ1L token + λ2L s + λ3L align

[0072] where L represents the loss function value, λ1, λ2, and λ3 represent weight parameters, and L token represents the semantic consistency loss, and L s represents the structural loss, and L align represents the semantic and structural alignment loss.

[0073] S33: Freeze most of the parameters of the base model (LegalBERT), initialize the parameters of the LoRA layer, and combine the initialized LoRA layer (A and B matrices) with the corresponding parameters of LegalBERT. The LoRA layer is a technique used in fine-tuning pre-trained models. By introducing low-rank matrices, it fine-tunes the weights of the model, enabling it to better adapt to specific tasks without changing the main structure of the pre-trained model. Matrix A: It is a matrix with a shape of (hidden layer dimension, r), where r is the rank of the low-rank decomposition, much smaller than the hidden layer dimension. It maps the input features to a low-dimensional space. Matrix B: It is a matrix with a shape of (r, hidden layer dimension), which maps the features in the low-dimensional space back to the original high-dimensional space.

[0074] It should be noted that here mainly the word embedding layer is frozen. When fine-tuning the large model of intelligent legal contracts, these general semantic and lexical information still plays an important fundamental role in specific tasks in the legal field and is not likely to change drastically during the fine-tuning process. Therefore, it is considered to freeze the parameters of the word embedding layer. At the same time, the underlying Transformer layers mainly learn relatively general language knowledge such as the basic grammar structure of the language and the basic semantic combinations of sentences. This knowledge has a certain degree of generality in different legal tasks, and the selected model has been pre-trained on a large-scale corpus, and the underlying parameters have been relatively stable and optimized. Therefore, the parameters of several underlying Transformer layers can be frozen to preserve the general language knowledge, while reducing the computational amount and time of training. At this time, the unfrozen parameters are only the LoRA layer, which is set according to the steps of S31.

[0075] In an embodiment of the present invention, iteratively training and fine-tuning a target large model using an input sequence includes: setting hyperparameters of the fine-tuning framework and starting the fine-tuning of the target large model; iteratively training the target large model using the input sequence, calculating the loss function value during the training process according to a comprehensive loss function, and optimizing the model parameters of the target large model according to the loss value; after meeting the iteration stop condition, completing the iterative training of the target large model.

[0076] Among them, the iteration stop condition may include that the performance of the target large model meets the target requirements, or the number of iteration training reaches the target number, etc.

[0077] It can be understood that in an embodiment of the present invention, after setting corresponding hyperparameters (such as learning rate, training length, and number of training epochs, etc.), model fine-tuning can be started, and through multiple iterative trainings, the model can be adjusted according to the results of the validation set, and the optimal model parameters can be saved.

[0078] Specifically, as Figure 5 shown, in an embodiment of the present invention, after setting corresponding hyperparameters (learning rate, training length, number of training epochs), model fine-tuning is started, specifically as follows:

[0079] S41: Input the preprocessed natural language contract - BPMN data into the target large model.

[0080] S42: Calculate the loss according to the determined comprehensive loss function, and record the change of the loss value for subsequent analysis of the training progress of the model.

[0081] S43: According to the obtained loss function value, calculate the gradient using the backpropagation algorithm and update the parameters using the optimization algorithm.

[0082] S44: Perform iterative training multiple times, adjust the model according to the results of the validation set, and save the optimal model parameters. Specifically: repeat the above forward propagation, loss calculation, backpropagation, and parameter update processes for multiple epochs of training. Regularly evaluate the performance of the model on the validation set using relevant evaluation metrics (F1 value). If it is found that the model is overfitting, some regularization measures are taken. If it is found that the model performance improves slowly, try to adjust the loss function weight, learning rate, or other hyperparameters. Among them, the F1 value is an index used to evaluate the balance of precision and recall in the model. It is the harmonic mean of precision and recall, providing a single score that considers both at the same time. The value range of the F1 value is between 0.0 and 1.0, and the closer the value is to 1.0, the better the model.

[0083] In step S104, the trained target large model is used to implement the intelligent legal contract modeling function, where the intelligent legal contract modeling function is to model a natural language contract into an intelligent legal contract.

[0084] It can be understood that after the training of the embodiments of the present invention is completed, the fine-tuned model parameters are saved and encapsulated, and the intelligent legal contract modeling function is embedded in the intelligent contract design platform. Specifically, as Figure 5 shown: The embodiments of the present invention can use an API (Application Programming Interface, program programming interface framework) (such as FastAPI) to build an intelligent contract modeling service. Embedding the intelligent legal contract modeling function into the intelligent contract design platform can improve usability.

[0085] In summary, the embodiments of the present invention collect various types of natural language contracts, determine various types of terms and elements in the contracts, and align the legal contract elements with the intelligent legal contract elements, so as to obtain high-quality input data and alleviate the problems of semantic association difficulties and information loss caused by different semantic fields and structures between natural language contracts and intelligent legal contracts; use the obtained corpus to fine-tune the pre-trained large language model in the legal field, so that the fine-tuned model can be aligned with the intelligent legal contract in BPMN form based on the analysis of natural language contract elements, and make the fine-tuned model better understand the concepts and logical relationships of legal terms; the fine-tuned large language model can be integrated with other platforms, for example, embedding it into the intelligent contract generation platform can improve the generation efficiency of natural language contracts to intelligent legal contract code. Therefore, the embodiments of the present invention can build a specific domain corpus of natural language contracts and intelligent legal contracts based on the large model fine-tuning technology, enabling the large model to better understand the legal terms and regulatory logic in natural language contracts, thereby optimizing the intelligent legal contract code modeling process. At the same time, completing the intelligent legal contract modeling of natural language contracts is not only beneficial to users' analysis of contract terms and legal risk review, but also beneficial to technicians' transformation of legal contracts into intelligent contracts, promoting the modernization of the legal society.

[0086] The intelligent contract modeling method based on large model fine-tuning will be further elaborated through a specific embodiment below, as Figure 5 shown, which mainly includes three stages, namely S1 data processing stage, S2 model fine-tuning stage, and S3 model evaluation and deployment stage, specifically as follows: (1) Collect various types and languages of natural language contracts, determine various types of terms and elements in the contracts, perform structured annotation, and align the data with the intelligent legal contract in BPMN business process modeling form to train the model; (2) Input the preprocessed dataset into the basic pre-training of LegalBert. The large model designs a loss function according to the differences between natural language contracts and intelligent legal contracts, and then inserts a LoRA layer at an appropriate position in the basic large model for fine-tuning; (3) Evaluate and encapsulate the fine-tuned large model, and embed the function into the intelligent contract generation platform.

[0087] In the S1 data processing stage, by collecting natural language contracts of multiple types and languages, aligning the data with BPMN business process data, the smoothness of the large language model and the accuracy of extracting complex natural language contract logic are improved; in the S2 model fine-tuning stage, the annotated corpus is preprocessed and then input into the basic pre-trained large model of LegalBert for parameter fine-tuning; in the S3 model evaluation and deployment stage, the trained model is experimentally optimized and encapsulated, deployed on a local server and embedded into the intelligent contract generation platform, and at the same time, feedback is collected to update the corpus and the model.

[0088] Therefore, the embodiment of the present invention proposes an intelligent legal contract modeling method based on large model fine-tuning, constructs a corpus for natural language contracts - intelligent legal contracts in a specific domain, inserts LoRA layers near three weight matrices in the self-attention module of the basic pre-trained model encoder, and designs a loss function based on the semantic loss and structural loss between natural language contracts and intelligent legal contracts to achieve fine-tuning of the large model parameters, enabling the large model to better understand legal terms and regulatory logics in natural language contracts, and further optimizing the intelligent legal contract code modeling process. The present invention utilizes the technical advantages of the large model's complex structure understanding and capture ability, cross-domain dynamic adaptation and generalization ability, efficient modeling and processing ability, etc., aiming to solve the limitations of equivalent modeling in the existing intelligent legal contract code field, improve the modeling efficiency and reduce the modeling cost.

[0089] Next, a device for converting natural language contracts into intelligent legal contracts according to an embodiment of the present invention is described with reference to the accompanying drawings.

[0090] Figure 6 It is a block diagram of the device for converting natural language contracts into intelligent legal contracts according to an embodiment of the present invention.

[0091] As Figure 6 shown, the device 10 for converting natural language contracts into intelligent legal contracts includes: an acquisition module 100, a processing module 200, a fine-tuning module 300, and a deployment module 400.

[0092] Among them, the acquisition module 100 is used to acquire natural language contract data and business process data of the business process model expressing the natural language contract; the processing module 200 is used to generate training corpus for the target large model according to the natural language contract data and the business process data; the fine-tuning module 300 is used to generate an input sequence of the fine-tuning framework according to the training corpus and the intelligent legal contract code, and iteratively train the target large model under the fine-tuning framework by using the input sequence. Among them, the fine-tuning framework includes a decoder, an encoder, and a comprehensive loss function. Low-rank matrix layers are respectively added to multiple parameter weight matrices of the encoder, and some model parameters of the target large model are frozen, and the model parameters of the target large model are adjusted based on the low-rank matrix layer and the comprehensive loss function; the deployment module 400 is used to use the trained target large model to implement the intelligent legal contract modeling function, where the intelligent legal contract modeling function is to model the natural language contract into an intelligent legal contract.

[0093] It should be noted that the foregoing explanation of the embodiment of the conversion of the natural language contract to the intelligent legal contract also applies to the device for converting the natural language contract to the intelligent legal contract in this embodiment, and will not be elaborated here.

[0094] The device for converting a natural language contract to an intelligent legal contract proposed according to an embodiment of the present invention constructs a corpus in the specific field of natural language contract-intelligent legal contract, inserts LoRA layers near three weight matrices in the self-attention module of the encoder of the basic pre-trained model, and designs a loss function based on the semantic loss and structural loss between the natural language contract and the intelligent legal contract, so as to realize the fine-tuning of the large model parameters, so that the large model can better understand the legal terms and regulatory logic in the natural language contract, and then optimize the intelligent legal contract code modeling process, so as to utilize the technical advantages of the large model's complex structure understanding and capture ability, cross-domain dynamic adaptation and generalization ability, efficient modeling and processing ability, etc., aiming to solve the equivalent modeling limitations in the existing intelligent legal contract code field, improve the modeling efficiency, and reduce the modeling cost.

[0095] Figure 7 The structural schematic diagram of the electronic device provided by the embodiment of the present invention. The electronic device may include:

[0096] A memory 701, a processor 702, and a computer program stored on the memory 701 and executable on the processor 702.

[0097] When the processor 702 executes the program, it implements the method for converting a natural language contract to an intelligent legal contract provided in the foregoing embodiment.

[0098] Furthermore, the electronic device further includes:

[0099] A communication interface 703 for communication between the memory 701 and the processor 702.

[0100] A memory 701 for storing a computer program that can run on a processor 702.

[0101] The memory 701 may include a high-speed RAM (Random Access Memory) memory, and may also include a non-volatile memory, such as at least one disk memory.

[0102] If the memory 701, the processor 702, and the communication interface 703 are implemented independently, the communication interface 703, the memory 701, and the processor 702 can be interconnected through a bus and communicate with each other. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0103] Optionally, in a specific implementation, if the memory 701, the processor 702, and the communication interface 703 are integrated on a single chip, the memory 701, the processor 702, and the communication interface 703 can communicate with each other through an internal interface.

[0104] The processor 702 may be a CPU (Central Processing Unit), or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention.

[0105] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0106] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0107] Any process or method description shown in a flowchart or described otherwise herein can be understood to represent a module, segment, or portion of code including one or N executable instructions for implementing a customized logical function or process. And the scope of the preferred embodiments of the present invention includes additional implementations, where functions may be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.

[0108] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays, field programmable gate arrays, etc.

[0109] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method for implementing the above embodiments can be completed by instructing relevant hardware through a program. The above program can be stored in a computer-readable storage medium, and when executed, it includes one or a combination of the steps of the method embodiments.

[0110] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for converting natural language contracts into intelligent legal contracts, characterized in that, Including the following steps: Obtain natural language contract data and business process data representing the business process model of the natural language contract. Among them, the natural language contract data includes various types of natural language contracts, and the natural language contract is a contract text written in natural language; Generate training corpus for the target large model based on the natural language contract data and the business process data; Generate an input sequence for the fine-tuning framework according to the training corpus and the intelligent legal contract code, and iteratively train the target large model under the fine-tuning framework using the input sequence. Among them, the fine-tuning framework includes a decoder, an encoder, and a comprehensive loss function. Add low-rank matrix layers to multiple parameter weight matrices of the encoder, freeze some model parameters of the target large model, initialize the parameters of the low-rank matrix layer, combine the initialized low-rank matrix layer with the corresponding parameters of the target large model, and adjust the model parameters of the target large model based on the low-rank matrix layer and the comprehensive loss function; The encoder includes an attention mechanism module and a fully connected feed-forward module. The multiple parameter weight matrices include the first to third parameter weight matrices. Among them, the attention weights of the attention mechanism module are determined based on the interaction of the first parameter weight matrix and the second parameter weight matrix, and the third parameter weight matrix is the integrated information weight of the attention mechanism module. The features of the natural language contract data and the business process data are assisted in learning through the three parameter weight matrices; The comprehensive loss function includes semantic consistency loss, structural loss, and semantic and structural alignment loss. Among them, the formula of the comprehensive loss function is: , where represents the loss function value, , and represent weight parameters, represents the semantic consistency loss, represents the structure loss, represents the semantic and structural alignment loss; Implement the intelligent legal contract modeling function using the trained target large model. Among them, the intelligent legal contract modeling function is to model the natural language contract into an intelligent legal contract.

2. The method for converting a natural language contract into an intelligent legal contract according to claim 1, wherein The generating the training corpus for the target large model according to the natural language contract data and the business process data includes: Perform format conversion and data cleaning on the natural language contract data and the business process data; Identify various clauses and elements in the natural language contract data, establish a mapping relationship between the various clauses and elements and various elements in the business process data, and form training samples according to the natural language contract data, the business process data, and the mapping relationship; Split the training samples into a training set, a validation set, and a test set, and generate the training corpus for the target large model according to the training set, the validation set, and the test set.

3. The method for converting a natural language contract into an intelligent legal contract according to claim 1, wherein The generating the input sequence for the fine-tuning framework according to the training corpus and the intelligent legal contract code includes: Convert the natural language contract data and the training corpus into the input format of the fine-tuning framework; Perform word segmentation on the intelligent legal contract code corresponding to the natural language contract data and the training corpus; Load the target large model into the fine-tuning framework, mark the word tokens after word segmentation according to the tokenizer of the target large model, and assign a unique index to the marked word tokens; Encode the marked word tokens into a vector matrix through the embedding layer of the fine-tuning framework, and use the vector matrix as the input sequence of the fine-tuning framework.

4. The method for converting a natural language contract into an intelligent legal contract according to claim 1, characterized in that Iteratively training the target large model under the fine-tuning framework using the input sequence includes: Setting hyperparameters of the fine-tuning framework and starting the fine-tuning of the target large model; Iteratively training the target large model using the input sequence, calculating the loss function value during the training process according to the comprehensive loss function, and optimizing the model parameters of the target large model according to the loss value; After meeting the iteration stop condition, completing the iterative training of the target large model.

5. The method for converting a natural language contract into an intelligent legal contract according to claim 1, wherein A low-rank matrix layer is added after the attention mechanism module.

6. The method for converting a natural language contract into an intelligent legal contract according to claim 1, characterized in that, The fine-tuning framework includes a Seq2Seq framework, and the decoder and the encoder are based on the Transformer architecture.

7. A device for converting natural language contracts into intelligent legal contracts, characterized in that, Including: An acquisition module for acquiring natural language contract data and business process data of a business process model expressing natural language contracts, wherein the natural language contract data includes various types of natural language contracts, and the natural language contract is a contract text written in natural language; A processing module for generating training corpus for the target large model according to the natural language contract data and the business process data; A fine-tuning module for generating an input sequence for the fine-tuning framework according to the training corpus and intelligent legal contract code, and iteratively training the target large model under the fine-tuning framework using the input sequence, wherein the fine-tuning framework includes a decoder, an encoder, and a comprehensive loss function. A low-rank matrix layer is respectively added to multiple parameter weight matrices of the encoder, some model parameters of the target large model are frozen, the parameters of the low-rank matrix layer are initialized, the initialized low-rank matrix layer is combined with the corresponding parameters of the target large model, and the model parameters of the target large model are adjusted based on the low-rank matrix layer and the comprehensive loss function; the encoder includes an attention mechanism module and a fully connected feed-forward module, and the multiple parameter weight matrices include first to third parameter weight matrices, wherein the attention weights of the attention mechanism module are determined based on the interaction of the first parameter weight matrix and the second parameter weight matrix, and the third parameter weight matrix is the integrated information weight of the attention mechanism module, and the characteristics of natural language contract data and business process data are assisted in learning through the three parameter weight matrices; The comprehensive loss function includes semantic consistency loss, structural loss, and semantic and structural alignment loss, wherein the formula of the comprehensive loss function is: , where represents the loss function value, , and represent the weight parameters, represents the semantic consistency loss, represents the structure loss, represents the semantic and structure alignment loss; A deployment module for implementing the intelligent legal contract modeling function using the trained target large model, wherein the intelligent legal contract modeling function is to model a natural language contract into an intelligent legal contract.

8. An electronic device, characterized in that, Including: A memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the program to implement the method for converting a natural language contract into an intelligent legal contract according to any one of claims 1-6.

Citation Information

Patent Citations

  • Intelligent legal contract generation method and device, electronic equipment and storage medium

    CN117311726A

  • Intelligent contract generation system, method and device based on large model and storage medium

    CN118504542A