Contract generation method based on interactive large language model

Through an interactive large language model, confirming the contract template with users and searching relevant legal terms, and generating contracts with RAG technology, the professionalism and accuracy problems of the general model when generating contracts in the legal field are solved, and intelligent and personalized contract generation is achieved, improving generation efficiency and accuracy.

CN120373311APending Publication Date: 2025-07-25GUIZHOU UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510441263.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing general language model has problems such as incomplete knowledge and irregular format in the generation of contracts in the legal field, which leads to the lack of professionalism and accuracy of the generated contracts, which affects its actual applicability.

Method used

The interactive large language model is used to confirm the contract template with users, and the fine-tuned embeding embedding model is used to search for relevant legal terms. It is transmitted to the legal model through RAG technology to assist in the generation of contracts. It is optimized with user feedback to ensure the accuracy and efficiency of contract generation.

Benefits of technology

It improves the accuracy and efficiency of contract generation, reduces the risk of model illusions, provides an intelligent and personalized contract generation experience, ensuring information integrity and data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373311A_ABST
    Figure CN120373311A_ABST
Patent Text Reader

Abstract

The invention discloses a contract generation method based on an interactive big language model, and belongs to the technical field of electronic contract generation, and the method comprises the steps: S1, carrying out the contract template confirmation of a law big model and a user interactive question and answer; s2, after the contract category is confirmed by the law big model, legal terms associated with the contract category are retrieved by semantic matching through the fine-tuned embedding model, and contract generation and contract risk review are assisted; s3, transmitting the retrieved legal clauses and contract templates to a large law model through an RAG technology, and assisting the model to generate a specific contract; s4, after the law big model obtains the contract template and the law terms related to the contract category, the law big model interactively communicates with the user to determine the final content of the contract; s5, according to the final content of the contract, the contract is output after being subjected to standardization and logic processing, and the required contract is generated; according to the invention, the integrity and accuracy of information are effectively ensured, and the quality and reliability of contract generation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic contracts, and particularly to a contract generation method based on an interactive large language model. Background Art

[0002] In modern commercial transactions, contracts play a crucial role. They clearly define the rights and obligations of all parties and provide legal protection for transactions. However, the process of drafting contracts is usually complex and cumbersome, requiring not only legal expertise but also rich practical experience.

[0003] In recent years, large language models, as a natural language processing technology based on artificial intelligence, have demonstrated powerful language understanding and generation capabilities and have performed excellently in various language tasks. With the popularization of this technology, more and more developers have started training large language models for the needs of vertical fields to improve their adaptability in specific fields. However, due to the professionalism and sensitivity of vertical field knowledge, some information cannot be publicly obtained. Against this background, Retrieval-Augmented Generation (RAG) has emerged.

[0004] However, general large language models (such as ChatGPT) are usually not specifically covered with professional knowledge in specific fields during the training process, especially in the legal field and contract generation, the research is still not deep enough. Therefore, although these models perform excellently in language understanding and generation, when dealing with legal texts and generating contracts, there may be problems such as incomplete knowledge and non-standard formats.

[0005] Even a dedicated model (such as ChatLaw) fine-tuned with a large amount of legal data may ignore the structural information of the contract when generating a specific contract text. This shortcoming may lead to the lack of professionalism and accuracy of the generated contract, affecting its practical applicability.

[0006] Based on this, the present invention provides a contract generation method based on an interactive large language model. Summary of the Invention

[0007] The object of the present invention is to provide a contract generation method based on an interactive large language model. By retrieving and matching the corresponding contract template types from the contract template library, the legal large language model can more accurately understand the user's intention and provide templates that meet the requirements, reducing the risk of the model having "hallucination" phenomena and significantly improving the accuracy and efficiency of contract generation. In addition, the present invention continuously optimizes the interactive dialogue process by collecting user feedback, thereby providing a more intelligent and personalized contract generation experience.

[0008] To achieve the above object, the present invention provides a contract generation method based on an interactive large language model, including the following steps:

[0009] S1. The legal large model interacts with the user through questions and answers to confirm the contract template;

[0010] S2. After the legal large model confirms the contract category, the fine-tuned embeding embedding model uses semantic matching to retrieve legal clauses associated with the contract category to assist in contract generation and contract risk review;

[0011] S3. Transmit the retrieved legal clauses and contract templates to the legal large model through the RAG technology to assist the model in generating a specific contract;

[0012] S4. After the legal large model obtains the contract template and legal clauses related to the contract category, it interacts with the user to determine the final content of the contract;

[0013] S5. According to the final content of the contract, after normalizing and logicalizing it, output it to generate the required contract.

[0014] Preferably, after the legal large model in S2 confirms the contract category, the process of using the fine-tuned embeding embedding model to retrieve legal clauses associated with the contract category through semantic matching is as follows:

[0015] S11. Extract corresponding keywords and key phrases from the contract generation requirements input by the user; and confirm the contract template through ES retrieval;

[0016] S12. Embed the legal clauses into a text vector space, and through vectorized text retrieval, retrieve legal clauses related to the contract;

[0017] S13. Train the embedding model with contract template data and legal clause data, and perform fine-tuning using the contrastive learning method;

[0018] S14. Use the trained embedding model to vectorize the contract template determined in S1, match the most relevant legal clauses using vector similarity, and input them into the legal large model.

[0019] Preferably, the specific process of fine-tuning the embedding model using the contrastive learning method in S13 is as follows:

[0020]

[0021] A = <e p, e q >;

[0022] A′ = <e p, eq′ >;

[0023] where p and q represent paired positive samples, q' ∈ Q' represents a negative sample, and e q and e p represent the embedding vectors of texts q and p, A represents the inner product of these two positive text embedding vectors, A' represents the inner product of the positive and negative sample embedding vectors, and e q′ represents the text embedding vector of the negative sample q', and τ represents the temperature parameter.

[0024] Preferably, in S14, the contract template determined in S1 is vectorized using the trained embedding model, the most relevant legal provisions are matched using vector similarity, and the specific process of inputting them into the legal large model is as follows:

[0025] S21. Embed the legal provision library into a text vector space using the embedding model;

[0026] S22. Calculate the similarity between the vectorized text information of the contract template obtained in S14 and the text vector space of the legal provisions obtained in S21, and confirm the legal provisions related to the contract category; among them, the cosine value of the angle between two space vectors is measured by the cosine similarity formula to evaluate the similarity degree of the two space vectors, which is specifically expressed as:

[0027]

[0028] where A and B respectively represent the vectors in the vectorized text information and the text vector space, <A, B> represents the dot product of A and B, ||A|| and ||B|| represent the norms of A and B, and the legal provisions represented by the vectors close to the vectorized text information of the contract template in the text vector space of the legal provisions are the legal provisions related to the contract;

[0029] S23. Perform semantic matching retrieval on the contract category and legal provisions to find the legal provisions most relevant to the contract category, and the specific process can be expressed as:

[0030]

[0031] where sim() represents the similarity between vectors, represents the space vector of the i-th contract template, V u represents the space vector of the legal provisions, and argmax i represents the maximum value of the similarity, that is, it represents as the legal regulations and provisions related to the contract category.

[0032] Preferably, the process of transmitting the retrieved legal provisions and contract templates to the legal large model through RAG in S3 is as follows:

[0033] S31. Perform template extraction to extract the user requirement contract template retrieved from the template library.

[0034] S32. Match the extracted contract template with the legal clause library to match the legal clauses related to this contract category; and transmit them to the legal large model through API calls.

[0035] S33. After receiving the contract template and legal clauses, the legal large model parses and understands their structures and contents.

[0036] Preferably, the process of determining the final content of the contract in S4 is as follows:

[0037] S41. Identify the type of the contract template transmitted in S3.

[0038] S42. Conduct an interactive communication Q&A with the user based on the type of the contract template and the information required by this type of template.

[0039] S43. Fill the final result obtained from the interaction on the contract template to determine the final content of the contract; and conduct a risk review on the contract content according to relevant laws and regulations, and conduct contract risk assessment item annotation.

[0040] Preferably, the legal clauses mentioned in S2 are stored as private information in a dedicated vector library.

[0041] Preferably, the contrastive learning method mentioned in S13 means that the embedding model learns to embed similar texts closer while pulling the embedding distances between unrelated negative samples farther apart, minimizing the loss of the model for paired positive samples, thereby improving its ability to distinguish different samples.

[0042] Therefore, the present invention adopts a contract generation method based on an interactive large language model with the above structure, constructs a contract template using the structured information in the contract, and combines the retrieval-augmented generation (RAG) technology and interactive technology to achieve a conversational interaction between the large language model and the user. First, confirm the contract template through an interactive Q&A with the user, then match the legal clauses related to the contract category with the help of the embedding model, then prompt the user to gradually fill in the necessary information through interaction, and conduct a contract risk review and annotation based on the filled information and legal clauses; finally, generate a complete contract text.

[0043] The technical solution of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings

[0044] Figure 1This is the implementation process diagram of a contract generation method based on an interactive large language model according to the present invention;

[0045] Figure 2 This is the flow chart of a contract generation method based on an interactive large language model according to the present invention. Detailed implementation manners

[0046] Embodiment

[0047] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Generally, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.

[0048] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0049] As Figure 1 and Figure 2 shown, a contract generation method based on an interactive large language model according to the present invention includes the following steps:

[0050] S1. The legal large language model interacts with the user through questions and answers to confirm the contract template;

[0051] S2. After the legal large language model confirms the contract category, the fine-tuned embeding embedding model uses semantic matching to retrieve legal clauses associated with the contract category to assist in contract generation and contract risk review; the legal clauses are stored as private information in a dedicated vector library; wherein, the process of the legal large language model using the fine-tuned embeding embedding model to retrieve legal clauses associated with the contract category through semantic matching is as follows:

[0052] S11. Extract corresponding keywords and key phrases from the contract generation requirements input by the user; and confirm the contract template through ES retrieval;

[0053] S12. Embed the legal clauses into the text vector space, and retrieve legal clauses related to the contract through vectorized text retrieval;

[0054] S13. Train an embedding model with contract template data and legal clause data, and fine-tune it using the contrastive learning method. The specific process of fine-tuning the embedding model using the contrastive learning method is as follows:

[0055]

[0056] A = <e p, e q >;

[0057] A' = <e p, e q′ >;

[0058] where p and q represent paired positive samples, q' ∈ Q' represents negative samples, e q and e p represent the embedding vectors of texts q and p, A represents the inner product of these two positive text embedding vectors, A' represents the inner product of positive and negative sample embedding vectors, τ represents the temperature parameter, and e q′ represents the text embedding vector of the negative sample q'. Additionally, the contrastive learning method is specifically: the embedding model learns to bring similar text embeddings closer while pulling the embedding distances between unrelated negative samples farther apart, minimizing the loss of the model for paired positive samples and improving its ability to distinguish different samples.

[0059] S14. Vectorize the contract template determined in S1 using the trained embedding model, match the most relevant legal clauses using vector similarity, and input them into the large legal model. The specific process is as follows:

[0060] S21. Embed the legal clause library into a text vector space using the embedding model;

[0061] S22. Calculate the similarity between the vectorized text information of the contract template obtained in S14 and the text vector space of the legal clauses obtained in S21 to find the legal clauses related to the contract category. Among them, the cosine similarity formula is used to measure the cosine value of the angle between two space vectors to evaluate the similarity degree of the two space vectors, which is specifically expressed as:

[0062]

[0063] where A and B respectively represent the vectors in the vectorized text information and the text vector space, <A, B> represents the dot product of A and B, ||A|| and ||B|| represent the norms of A and B, and the legal clause represented by the vector in the text vector space of the legal clauses that is close to the vectorized text information of the contract template is the legal clause related to the contract;

[0064] S23. Semantically match and retrieve the contract category and legal terms to find the legal terms most relevant to the contract category. The specific process can be expressed as follows:

[0065]

[0066] Among them, sim() represents the similarity between vectors. represents the spatial vector of the i-th contract template, V u represents the spatial vector of the legal terms, argmax i represents the maximum similarity value.

[0067] S3. Transmit the retrieved legal terms and contract templates to the legal large model through the RAG technology to assist the model in generating a specific contract. The specific process is as follows:

[0068] S31. Perform template extraction to extract the user-demand contract template retrieved from the template library.

[0069] S32. Match the extracted contract template with the legal terms library to find the legal terms related to this contract category; and transmit them to the legal large model by means of API call.

[0070] S33. After receiving the contract template and legal terms, the legal large model parses and understands their structures and contents.

[0071] S4. After the legal large model obtains the contract template and the legal terms related to the contract category, it interacts with the user to determine the final content of the contract. Among them, the process of determining the final content of the contract is as follows:

[0072] S41. Identify the type of the contract template transmitted in S3.

[0073] S42. Conduct an interactive communication Q&A with the user according to the type of the contract template and the information required by this type of template.

[0074] S43. Fill the final result obtained from the interaction into the contract template to determine the final content of the contract; and conduct a risk review of the contract content according to the relevant laws and regulations, and mark the contract risk items.

[0075] Therefore, the contract generation method based on an interactive large language model with the above structure is adopted in the present invention. Through an interactive dialogue method, the user is guided to gradually input based on the structured information of the contract template, effectively ensuring the integrity and accuracy of the information. Moreover, the context Q&A of each interaction is more concise and clear, which helps to reduce the risk of the large model generating hallucinations, thereby improving the quality and reliability of the contract generation process. At the same time, the corresponding contract template type is retrieved and matched from the contract template library, enabling the legal large language model to more accurately understand the user's intention and provide a template that meets the requirements. And the contract is audited for risks through legal terms to ensure the quality and accuracy of the contract generation. At the same time, the contract template library is stored as private information in a dedicated vector library, which not only ensures the confidentiality and integrity of the user data, but also provides flexible retrieval capabilities, enabling users to quickly obtain relevant templates when needed, improving the data security in the contract generation process. Finally, the present invention also continuously optimizes the interactive dialogue process by collecting user feedback, thereby providing users with a more intelligent and personalized contract generation experience.

[0076] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A contract generation method based on an interactive large language model, characterized in that, It includes the following steps: S1. The legal large model interacts with the user for contract template confirmation through Q&A; S2. After the legal large model confirms the contract category, it uses the fine-tuned embedding model to retrieve legal clauses associated with the contract category through semantic matching to assist in contract generation and contract risk review; S3. Transmit the retrieved legal clauses and contract templates to the legal large model through RAG technology to assist the model in generating a specific contract; S4. After the legal large model obtains the contract template and legal clauses related to the contract category, it communicates with the user to determine the final content of the contract; S5. According to the final content of the contract, perform normalization and logical processing on it and then output to generate the required contract.

2. The contract generation method based on an interactive large language model according to claim 1, wherein The process of the legal large model in S2 retrieving legal clauses associated with the contract category through the fine-tuned embedding model after confirming the contract category is as follows: S11. Extract corresponding keywords and key phrases from the contract generation requirements input by the user; and confirm the contract template through ES retrieval; S12. Embed the legal clauses into a text vector space, and retrieve legal clauses related to the contract through vectorized text retrieval; S13. Train the embedding model with contract template data and legal clause data, and perform fine-tuning using the contrastive learning method; S14. Vectorize the contract template determined in S1 using the trained embedding model, match the most relevant legal clauses using vector similarity, and input them into the legal large model.

3. The contract generation method based on an interactive large language model according to claim 2, wherein: The specific process of fine-tuning the embedding model using the contrastive learning method in S13 is as follows: A = <e p , e q >; A′ = <e p, e q′ >; where p and q represent a pair of positive samples, q′ ∈ Q′ represents a negative sample, e q and e p represent the embedding vectors of texts q and p, A represents the inner product of these two positive text embedding vectors, A′ represents the inner product of the positive and negative sample embedding vectors, τ represents the temperature parameter, e q′ represents the text embedding vector of the negative sample q′.

4. The contract generation method based on an interactive large language model according to claim 3, wherein The specific process of vectorizing the contract template determined in S1 using the trained embedding model, matching the most relevant legal clauses using vector similarity, and inputting them into the legal large model in S14 is as follows: S21. Embed the legal clause library into a text vector space using the embedding model; S22. Calculate the similarity between the vectorized text information of the contract template obtained in S14 and the text vector space of the legal clauses obtained in S21, and find the legal clauses related to the contract category; among them, the cosine similarity formula is used to measure the cosine value of the angle between two space vectors to evaluate the similarity degree of the two space vectors, which is specifically expressed as: where A and B respectively represent the vector in the vectorized text information and the text vector space, <A,B> represents the dot product of A and B, ||A|| and ||B|| represent the norms of A and B, and the legal clauses represented by the vectors close to the vectorized text information of the contract template in the legal clause text vector space are the legal clauses related to the contract; S23. Perform semantic matching retrieval on the contract category and legal clauses to find the legal clauses most relevant to the contract category. The specific process can be expressed as: where sim() represents the similarity between vectors, represents the spatial vector of the i-th contract template, V u represents the spatial vector of legal terms, argmax i represents the maximum value of similarity.

5. The contract generation method based on an interactive large language model according to claim 4, wherein The process of transmitting the retrieved legal clauses and contract templates to the legal large model through RAG in S3 is as follows: S31. Perform template extraction to extract the contract template of the user's requirements retrieved from the template library; S32. Match the extracted contract template with the legal clause library to find legal clauses related to this contract category; And transmit it to the legal large model through API calls; S33. After receiving the contract template and legal clauses, the legal large model analyzes and understands their structures and contents.

6. The contract generation method based on an interactive large language model according to claim 5, characterized in that, The process of determining the final content of the contract in S4 is as follows: S41. Identify the type of the contract template transmitted in S3; S42. Conduct an interactive communication Q&A with the user based on the type of the contract template and the information required by this type of template; S43. Fill the final result obtained from the interaction into the contract template to determine the final content of the contract; And conduct a risk review of the contract content according to relevant laws and regulations, and mark the contract risk assessment items.

7. A contract generation method based on an interactive large language model according to claim 6, characterized in that: The legal clauses mentioned in S2 are stored as private information in a dedicated vector library.

8. The contract generation method based on an interactive large language model according to claim 7, wherein: The specific contrastive learning method mentioned in S13 is as follows: The embedding model learns to embed similar texts closer while pulling the embedding distances between irrelevant negative samples farther apart, minimizing the loss of the model for paired positive samples and improving its ability to distinguish different samples.

Citation Information

Cited By

  • Contract generation system and method based on multiple agents

    CN121234894A

  • Electronic contract generating and sending method, device and equipment based on large language model

    CN121764691A

  • Electronic contract generation and sending method, device and equipment based on large language model

    CN121764691B