Intelligent generation method for contract in specific field
Through self-distillation and code pre-training small language models, combined with instruction adjustment and neuron pre-training, the problems of low credibility and high computational volume in contract generation are solved, and efficient and accurate generation of contracts in specific fields are achieved.
Patent Information
- Application Number
- CN202510456920.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-12
- Publication Date
- 2025-07-11
AI Technical Summary
The existing large language models have low credibility, lack targetedness and excessive calculation volume when generating contracts, which cannot effectively meet the contract generation needs of specific fields.
Small language models are built through self-distillation and code pre-training, combining instruction adjustment and neuron pre-training, using the terms standard library and contract template library for field-specific contract generation, and feedback adjustments are made through reinforcement learning to improve credibility and targeting.
有效降低了计算参数量,提高了合同生成的可信度和针对性,确保生成的合同符合特定领域的法律要求,提高了合同起草的效率和质量。
Smart Images

Figure CN120298167A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of processing legal documents and belongs to the field of information and communication technologies specifically applicable to administrative, commercial, financial, management, or supervision purposes. Background Art
[0002] There are various large language models in the prior art that can automatically generate contracts through question and answer. However, when actually used for the purpose of generating contracts, these large language models have the following problems in several aspects:
[0003] A. Low credibility of output results: Through actual tests, all existing large language models will fabricate the basis on which the contract depends (including legal and regulatory basis and other objective basis), and the generated contracts will quote content that actually does not exist. For example, a certain labor contract generated will quote Article 5 of the Labor Contract Law, but its content is created by the large language model. In fact, the provisions of Article 5 of the Labor Contract Law are about the tripartite mechanism for coordinating labor relations and have no direct relation to labor contracts;
[0004] B. Lack of pertinence: The content generated by existing large language models only generally talks about the general situation involved in the user input content. Through actual tests, even if the user input content is adjusted to be specific to a certain field, the content generated by the large language model obtained is still a general concept of popularization and lacks content with strong pertinence;
[0005] C. Excessive computational amount: Among general large language models, BLOOM has 17.6 billion parameters, GPT-NeoX has 2 billion parameters, DeepSeek-V3 has 67.1 billion parameters, and ChatGPT-3.5 has 2 billion parameters. The number of computational parameters is too large, far exceeding the upper limit of 10 million parameters that can be accepted by general projects. Summary of the Invention
[0006] To solve the above technical problems, the present invention provides a method for intelligently generating contracts in a specific field. This method for intelligently generating contracts in a specific field effectively avoids the problem of excessive computational parameters of large language models based on the method of self-distillation and code pre-training, and can effectively ensure the credibility and pertinence of the output content based on the method of instruction tuning and neuron pre-training.
[0007] The present invention is achieved through the following technical solutions.
[0008] A method for intelligently generating contracts in a specific field provided by the present invention includes the following steps:
[0009] S1. Prepare the corpus: Convert multiple contract templates in a specific field into contract terms, store the contract terms in a clause standard library, form at least one type of contract template based on the contract terms, and store the contract template in a contract template library;
[0010] S2. Build a language model: Migrate an open-source LLM model and self-distill the open-source LLM model into a small language model;
[0011] S3. Code pre-training: Use the CodeBERT model to perform code pre-training on the small language model;
[0012] S4. Instruction tuning: Use a preset instruction dataset to perform instruction tuning on the small language model;
[0013] S5. Neuron pre-training: Use the data in the clause standard library and contract template library in step S1 to perform neuron pre-training on the small language model;
[0014] S6. Generate a contract: Generate a contract using the small language model according to the user's instruction.
[0015] After step S5, there is also step S7 that runs alternatively with step S6, that is:
[0016] S7. Contract review: According to the user's instruction, use the small language model to review the contract text uploaded by the user and return the review result to the user.
[0017] After step S6, there are also step S8 and step S9, that is:
[0018] S8. Contract verification: Verify the generated contract by means of manual correction or review and correction by the small language model, and output the verified contract text;
[0019] S9. Reinforcement learning feedback adjustment: In the way of reinforcement learning, use the verified contract text to perform feedback adjustment on the small language model.
[0020] In step S8, after verifying by means of at least 100 times of manual correction, then verify by means of review and correction by the small language model.
[0021] Step S1 specifically includes the following steps:
[0022] S1.1. Source standardization: Perform standardization processing on the contract templates in a specific field to obtain standardized contract templates;
[0023] S1.2. Clause atomization: Unseal the standardized contract templates article by article, and each article of the contract text is a contract clause;
[0024] S1.3. Marking Tags: After marking type tags for contract terms, store them in the clause standard library;
[0025] S1.4. Composing Templates: Combine the contract terms to obtain multiple contract templates, and store the combined contract templates in the contract template library.
[0026] The steps S1.1 and S1.2 are carried out using CN-Text-Normalizer.
[0027] The open-source LLM model used is GPT-NeoX.
[0028] The number of parameters of the small language model is not greater than 0.4 times the number of parameters of the open-source LLM model.
[0029] In the step S4, an instruction dataset is constructed based on the clause standard library and the standard instructions preset by the user.
[0030] The beneficial effects of the present invention are as follows: Based on the methods of self-distillation and code pre-training, the problem of excessive computational parameters of the large language model is effectively avoided. Based on the methods of instruction tuning and neuron pre-training, the credibility and pertinence of the output content can be effectively ensured. Description of the Drawings
[0031] Figure 1 is a schematic flowchart of at least one implementation manner of the present invention;
[0032] Figure 2 is Figure 1 the specific flowchart of the step of preparing the corpus in Detailed Embodiments
[0033] The technical solutions of the present invention are further described below, but the scope of protection claimed is not limited thereto.
[0034] The first embodiment of the present invention relates to a method for intelligent generation of contracts in a specific field as shown in Figure 1 and includes the following steps:
[0035] S1. Preparing the Corpus: Convert multiple contract templates in a specific field into contract terms, store the contract terms in the clause standard library, and form at least one type of contract template according to the contract terms, and store the contract template in the contract template library;
[0036] S2. Building a Language Model: Migrate the open-source LLM model and self-distill the open-source LLM model into a small language model;
[0037] S3. Code Pre-training: Use the CodeBERT model to perform code pre-training on the small language model; thereby being able to handle common problems of NL-PL, such as searching for code in natural language, automatically generating code, etc.;
[0038] S4. Instruction Adjustment: Use a preset instruction dataset to perform instruction adjustment on the small language model; Instruction adjustment is a technique for fine-tuning a large language model (LLM) on a labeled instruction prompt and corresponding output dataset, which can not only improve the model's performance in specific tasks but also enhance the model's ability to follow instructions in general, thereby helping to adjust the pre-trained model to make it suitable for actual use.
[0039] S5. Neuron Pre-training: Use the data in the clause standard library and contract template library in step S1 to perform neuron pre-training on the small language model; This step is the model pre-training, mainly used to optimize the grammar rules of the language;
[0040] S6. Generate Contract: Generate a contract using the small language model according to the user's instructions.
[0041] Generally, the above step S4 mainly performs fine-tuning by providing the model with clear instructions or examples for specific tasks. Usually, it retains the knowledge of the pre-trained model, can focus on fine-tuning for specific tasks, has strong adaptability, and retains the basic capabilities of the model. Although it may not fully exploit the potential of the model in some highly complex tasks, the small language model in this application is mainly used for intelligent contract generation in a specific field and generally does not involve highly complex tasks. Therefore, it can just avoid the weaknesses of instruction fine-tuning.
[0042] It is easy to understand that the contract template is a contract that has been reviewed by legal affairs and is legal and compliant within a specific field, thereby ensuring the quality of the input content.
[0043] The second embodiment of the present invention is substantially the same as the first embodiment. The main difference is that after step S5, it further includes step S7 that runs alternatively with step S6, that is:
[0044] S7. Contract Review: According to the user's instructions, use the small language model to review the contract text uploaded by the user and return the review result to the user.
[0045] After the said step S6, it further includes step S8 and step S9, that is:
[0046] S8. Contract Verification: Verify the generated contract by means of manual correction or review and correction by the small language model, and output the verified contract text;
[0047] S9. Reinforcement Learning Feedback Adjustment: Use the verified contract text to perform feedback adjustment on the small language model in the way of reinforcement learning.
[0048] Furthermore, in step S8, first verify it by means of at least 100 times of manual correction, and then verify it by means of review and correction by the small language model.
[0049] The third embodiment of the present invention is substantially the same as the first embodiment. The main difference is that step S1 specifically includes the following steps:
[0050] S1.1, Source standardization: Standardize the contract templates in a specific field to obtain standardized contract templates;
[0051] S1.2, Clause atomization: Unseal the standardized contract templates by clause, and each contract text is a contract clause;
[0052] S1.3, Marking labels: Mark type labels for the contract clauses and store them in the clause standard library;
[0053] S1.4, Composition templates: Combine the contract clauses to obtain multiple contract templates, and store the combined contract templates in the contract template library.
[0054] Further, steps S1.1 and S1.2 are carried out using CN-Text-Normalizer.
[0055] Preferably, the open-source LLM model uses GPT-NeoX.
[0056] Further, the number of parameters of the small language model is not greater than 0.4 times the number of parameters of the open-source LLM model.
[0057] Preferably, in step S4, an instruction data set is constructed based on the clause standard library and the standard instructions preset by the user.
[0058] Thus, the above-mentioned embodiments mainly effectively avoid the problem of excessive calculation parameters of the large language model based on the methods of self-distillation and code pre-training. On the premise of ensuring the effective realization of the basic purpose (that is, the intelligent generation of contracts for specific fields), the unnecessary calculation amount is effectively reduced; based on the methods of instruction tuning and neuron pre-training, the generalized language model can be more targeted and effective for contract generation, effectively improving the efficiency and quality of contract drafting.
[0059] Generally speaking, the technical solution of this application is mainly implemented by deploying code on the SAAS platform, and the overall advantages can be achieved in the following aspects:
[0060] 1. Use artificial intelligence algorithms to automatically atomize (disassemble) contract samples, mark clause labels, and generate clause big data; make the drafting and review of contracts more convenient, faster, and more efficient;
[0061] 2. By adopting artificial intelligence search engine technology and big data knowledge graphs, the most suitable templates can be quickly found from a vast number of contract templates, helping enterprises avoid potential legal risks while improving the efficiency and quality of contract drafting;
[0062] 3. An advanced neuron computing technology is used to establish a training model, combined with protocol rule configuration, to intelligently generate integrated protocols and improve contract generation efficiency;
[0063] 4. The IPA robot can be used to implement the intermediate audit and review function of business contracts. The AI intelligent recognition technology is used to accurately identify the risk points in the contract (such as detecting tampered information by comparison), reducing contract risks;
[0064] 5. It is convenient to realize the full-cycle management from before, during, and after signing. Among them, before signing: contract clause atomization, contract template library, contract automatic generation; during signing: contract automatic generation; after signing: contract intelligent review.
Claims
1. A method for intelligent generation of contracts in a specific field, characterized in that: It includes the following steps: S1. Prepare corpus: Convert multiple contract templates in a specific field into contract terms, store the contract terms in a clause standard library, form at least one type of contract template based on the contract terms, and store the contract template in a contract template library; S2. Build a language model: Migrate an open-source LLM model and self-distill the open-source LLM model into a small language model; S3. Code pre-training: Use the CodeBERT model to perform code pre-training on the small language model; S4. Instruction tuning: Use a preset instruction dataset to perform instruction tuning on the small language model; S5. Neuron pre-training: Use the data in the clause standard library and contract template library in step S1 to perform neuron pre-training on the small language model; S6. Generate a contract: Generate a contract using the small language model according to the user's instruction.
2. The specific field contract intelligent generation method according to claim 1, wherein: After step S5, there is also step S7 that runs alternatively with step S6, that is: S7. Contract review: Review the contract text uploaded by the user using the small language model according to the user's instruction and return the review result to the user.
3. The specific domain contract intelligent generation method according to claim 1, wherein: After step S6, there are also step S8 and step S9, that is: S8. Contract verification: Verify the generated contract by means of manual correction or review and correction by the small language model, and output the verified contract text; S9. Reinforcement learning feedback adjustment: Use the verified contract text to perform feedback adjustment on the small language model in the way of reinforcement learning.
4. The specific field contract intelligent generation method according to claim 3, wherein: In step S8, after verifying by means of at least 100 times of manual correction, verify by means of review and correction by the small language model.
5. The specific domain contract intelligent generation method according to claim 1, characterized in that: Step S1 specifically includes the following steps: S1.
1. Source standardization: Perform standardization processing on the contract templates in a specific field to obtain standardized contract templates; S1.
2. Clause atomization: Unseal the standardized contract templates by clause, and each contract text is a contract clause; S1.
3. Mark tags: Mark type tags on the contract clauses and then store them in the clause standard library; S1.
4. Form templates: Combine the contract clauses to obtain multiple contract templates, and store the combined contract templates in the contract template library.
6. The specific field contract intelligent generation method according to claim 1, characterized in that: Steps S1.1 and S1.2 are performed using CN-Text-Normalizer.
7. The specific field contract intelligent generation method according to claim 1, characterized in that: The open-source LLM model uses GPT-NeoX.
8. The specific domain contract intelligent generation method according to claim 1, characterized in that: The number of parameters of the small language model is not greater than 0.4 times the number of parameters of the open-source LLM model.
9. The specific field contract intelligent generation method according to claim 1, characterized in that: In step S4, an instruction dataset is constructed based on the clause standard library and the standard instructions preset by the user.