Method and system for automatically generating legal document through large model
Through large-scale model training and multi-dimensional data processing, the automatic generation of legal documents is achieved, which solves the problems of time-consuming and labor-intensive manual writing and unstable quality in existing technologies, and realizes efficient and accurate generation of legal documents.
Patent Information
- Application Number
- CN202510866606.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-09-26
AI Technical Summary
Existing legal document writing relies on manual operations, which is time-consuming and labor-intensive. Existing software also lacks the ability to analyze in-depth legal knowledge, resulting in unstable document quality.
Using large-scale model training methods, high-quality legal documents are generated through data preparation and annotation, model training, information processing and document generation, combined with natural language processing technology and knowledge graphs.
It significantly shortens the time for writing legal documents, reduces human errors, improves the efficiency and quality of document generation, and can keep up with legal changes in a timely manner, meet personalized needs, and improve user satisfaction.
Smart Images

Figure CN120706379A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and legal technology, and in particular to a method and system for automatically generating legal documents through a large model. Background Art
[0002] In traditional legal practice, legal documents are primarily drafted manually by legal professionals. This process is not only time-consuming and labor-intensive, requiring a high level of professionalism and experience, but also prone to errors and omissions, resulting in inconsistent quality in legal documents.
[0003] While some existing legal document writing software offers templates and basic text editing capabilities, it lacks the in-depth analysis and intelligent processing capabilities of legal knowledge, making it difficult to generate high-quality legal documents tailored to specific case circumstances. The rise of big model technology offers a new approach to addressing these issues.
[0004] Therefore, a method and system for automatically generating legal documents through a large model is needed to address the above-mentioned problems. Summary of the Invention
[0005] The purpose of the present invention is to solve the above problems and to propose a method and system for automatically generating legal documents through a large model.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions: A method for automatically generating legal documents through a large model includes the following parts: Data preparation and annotation: Obtain open-source legal corpora, desensitized authentic documents, government announcements, and industry-standard contract templates; clean the data, unify the text encoding, and then annotate and build a knowledge base; Large model training: After the model is determined, the model structure and parameters are adjusted according to the characteristics and requirements of legal document generation. The pre-processed data is divided into training sets, validation sets, and test sets according to the preset ratios for training. Information processing: receiving case-related information input by users, and analyzing and processing it; Document generation: Generate legal documents based on processed user input information and the big model; Output and feedback: Output the generated legal documents to the user and receive feedback from the user.
[0007] Preferably, the data is cleaned, unified into text code, annotated, and a knowledge base is constructed, specifically including: Labeling via hierarchical annotation: The entity layer marks the party information, time, and amount; The relationship layer marks the logical relationship between legal elements; Event layer, marking the legal event type; Building a knowledge base: Establishing a legal concept system; Extract relationships from legal texts through remote supervision, connect to legal database APIs, and implement regular updates of the knowledge base.
[0008] Preferably, the model training specifically includes three parts: model selection, model evaluation and model optimization, wherein the specific content of model selection includes the following parts: After determining the selected model, calculate the accuracy of the model on various tasks using the language understanding benchmark test set; Using automatic evaluation metrics, perplexity measures the model's ability to predict test data; Consider the model's parameter size, computational complexity, and hardware resource requirements; after selecting the base model, fine-tune the model based on the specific requirements of legal document generation.
[0009] Preferably, the model evaluation includes the following parts: The preprocessed data is divided into a training set, a validation set, and a test set according to a preset ratio. After using the adaptive learning rate optimization algorithm, a preset number of samples are selected from the training set for training each time, and the average loss of the batch samples is calculated. The parameters of the model are updated according to the optimization algorithm. During the training process, the model's loss value and accuracy indicators are monitored in real time; the training curve is drawn to observe the model's training process.
[0010] Preferably, the model optimization includes the following parts: Randomly discard a preset number of neurons in the model's neural network layer to prevent excessive dependence between neurons and improve the model's generalization ability; Add a regularization term to the loss function to penalize the complexity of the model; Fuse multiple trained models and synthesize the output results of multiple models; As new legal data is generated, the model is continuously trained on a regular basis.
[0011] Preferably, the information processing specifically includes the following parts: Use natural language processing technology to conduct in-depth analysis of user input and extract key information; Based on the extracted key information, a knowledge graph of the case is constructed to visualize the various entities in the case and their relationships; Establish an information verification rule library to verify the extracted key information according to legal provisions and business logic; Use logical reasoning algorithms to verify the logical consistency of information and check whether there are any contradictions or unreasonableness between the information; if problems are found in the information, prompt users to modify and supplement it in a timely manner.
[0012] Preferably, the document generation specifically includes the following parts: Establish a legal document template library, which contains various types of legal document templates; each template has been reviewed and optimized by professional legal personnel; Based on the case type and legal claims entered by the user, a template is selected from the template library using rule matching and machine learning algorithms. The template is adjusted and modified based on the special circumstances of the case and the user's personalized needs. Organize and optimize the content generated by the large model, segment, sort, and add titles to the content according to the logical structure and format requirements of legal documents; Use the legal rules engine to check the content of generated legal documents; Use logical reasoning algorithms to check the logical rationality of the document content; The generated legal documents will be submitted to professional legal personnel for manual review, and the content, format, language expression and other aspects of the documents will be comprehensively checked and corrected.
[0013] Preferably, the output and feedback specifically include the following parts: Convert generated legal documents into common document formats; Provides various output setting options, allowing users to personalize the output documents according to their preferences; Support users to download the generated legal documents to local computer and provide printing function; Develop a feedback interface and provide multiple feedback channels; Based on user feedback, the large model is further trained and optimized, and the system parameters and algorithms are adjusted.
[0014] A system for automatically generating legal documents through a large model, including the following parts: Data preparation and annotation module: This module obtains legal-related data, cleans it, removes noise, duplication, and erroneous information, and then unifies the text encoding. It then performs hierarchical annotation on the data to build a knowledge base. Model training module: After determining the candidate model, the accuracy of the model on tasks such as text classification and natural language inference is calculated using the language understanding benchmark test set. The pre-processed data is divided according to a preset ratio to update the model parameters. Multiple trained models are fused and the output results of multiple models are synthesized. Information processing module: Extracts key information after in-depth analysis of user input, builds a knowledge graph of the case, establishes an information verification rule base, and verifies the extracted key information according to legal provisions and business logic; Text module: Establish a legal document template library, select appropriate templates from the template library, generate the specific content of the legal document based on the processed user input information and the large model, and convert the generated legal document into a document format for output.
[0015] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are: 1. The present invention realizes the automatic generation of legal documents by combining large-scale model training and multi-dimensional data processing methods. In the data preparation stage, the comprehensiveness and accuracy of legal knowledge are ensured through layered annotation and knowledge base construction. In the large-scale model training, a variety of evaluation indicators, optimization algorithms and regularization methods are used to enable the model to accurately learn legal language patterns and logical relationships. In the information processing link, with the help of natural language processing technology and knowledge graph construction, it can deeply understand the case information input by the user. This series of operations greatly shortens the time for writing legal documents, reduces errors and omissions in manual writing, and significantly improves the efficiency and quality of legal document generation.
[0016] 2. On the one hand, the present invention realizes regular updates of the knowledge base by connecting to the legal database API, and continuously trains the model as new legal data is generated, so as to keep pace with the dynamic changes in the legal field and ensure that the generated legal documents comply with the latest legal provisions and judicial practices. On the other hand, it develops a feedback interface to collect user opinions, and further trains the large model and optimizes the system parameter algorithm based on the feedback, which can continuously meet the personalized needs of users, improve user satisfaction, and enable the system to continuously improve and progress in long-term use, maintaining good performance and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Further details, features and advantages of the present application are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which: Figure 1 Flowchart of the present invention. DETAILED DESCRIPTION
[0018] Several embodiments of the present application will be described in more detail below with reference to the accompanying drawings so that those skilled in the art can implement the present application. The present application can be embodied in many different forms and for many different purposes and should not be limited to the embodiments described herein. These embodiments are provided to make the present application comprehensive and complete and to fully convey the scope of the present application to those skilled in the art. The embodiments do not limit the present application.
[0019] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. It will be further understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the relevant art and / or the context of this specification, and will not be interpreted in an idealized or overly formal sense unless expressly defined as such herein.
[0020] See also Figure 1 As shown, the present invention provides a technical solution: A method for automatically generating legal documents using a large model, comprising the following parts: Data preparation and annotation: Obtain open-source legal corpora, desensitized authentic documents, government announcements, and industry-standard contract templates; clean the data, unify the text encoding, and then annotate and build a knowledge base; After cleaning the data, the unified text encoding is performed and then annotated and a knowledge base is constructed, including: Labeling via hierarchical annotation: At the entity level, the party information (name / company, ID number / registration number), time, and amount are marked; The relationship layer marks the logical relationship between legal elements (e.g., “breach of contract → liability for breach of contract”); Event layer, marking the legal event type (such as "contract formation" and "tort occurrence"); Building a knowledge base: Establish a legal concept system (e.g., "contract" includes the sub-concepts of "subject", "object" and "content"); Through remote supervision, the relationships between "law → case" and "crime → constituent elements" are extracted from legal texts, and connected to legal database APIs (such as Peking University Law Library and Westlaw) to achieve regular updates of the knowledge base; Large model training: After the model is determined, the model structure and parameters are adjusted according to the characteristics and requirements of legal document generation. The pre-processed data is divided into training sets, validation sets, and test sets according to the preset ratios for training. Model training specifically includes three parts: model selection, model evaluation, and model optimization. The specific content of model selection includes the following parts: After deciding on the model of choice, calculate the model's accuracy on various tasks (such as text classification and natural language inference) using a language understanding benchmark test set; Assume that in a text classification task, the model predicts the number of correct samples to be A and the total number of samples to be B, then the accuracy is A divided by B; Using automatic evaluation metrics, perplexity measures the model's ability to predict test data; The calculation formula is: , where N is the number of words in the test data, is the conditional probability of the model predicting the i-th word given the previous i-1 words; Consider the model's parameter size, computational complexity, and hardware resource requirements. Parameter size is typically expressed as the number of model parameters, M, while computational complexity can be measured by the number of floating-point operations (FLOPs) required for the model's forward and backward propagation. After selecting the basic model, fine-tune the model according to the specific needs of legal document generation; The fine-tuning process mainly involves adjusting some parameters of the model to adapt it to the language characteristics and knowledge system of the legal field. The objective function of fine-tuning usually adopts the cross-entropy loss function. For classification tasks, its calculation formula is: ; Where N is the number of samples, C is the number of categories, is the true label 0 or 1 that the i-th sample belongs to the j-th class). is the probability that the model predicts that the i-th sample belongs to the j-th class; Model evaluation consists of the following parts: The preprocessed data is divided into training set, validation set, and test set according to the preset ratio; the common division ratio is 80%, 10%, and 10%; the training set is used to update the model parameters, the validation set is used to evaluate the model performance and adjust the hyperparameters during the training process, and the test set is used to finally evaluate the generalization ability of the model; After using the adaptive learning rate optimization algorithm, a preset number of samples are selected from the training set for training each time, and the average loss of the batch samples is calculated, and the parameters of the model are updated according to the optimization algorithm; use The optimizer updates the network parameters. The update formula of the optimizer is: Compute the gradient: ; Compute the first moment estimate: ; Compute the second moment estimate: ; Modified first moment estimate: ; Modified second moment estimate: ; Update parameters: ; in yes The gradient of time, is the learning rate, is a preset constant, and is the decay rate; During training, monitor the model's loss and accuracy in real time. Draw a training curve to observe the model's training progress. If the loss stops decreasing or the validation set accuracy starts to decrease, it may indicate that the model is overfitting and requires appropriate adjustments. Model optimization includes the following parts: Randomly discard a preset number of neurons in the model's neural network layer to prevent excessive dependence between neurons and improve the model's generalization ability; Assume that in a certain layer, the output of the neuron is x and the Dropout rate is p, then the output y after Dropout processing is: ; where r is a random binary vector of the same dimension as x, each element is 0 with probability p and 1 with probability 1-p, Represents element-wise multiplication; Add a regularization term to the loss function to penalize the complexity of the model; The calculation formula of the L1 regularization term is:
[0021] The calculation formula of L2 regularization term is:
[0022] in, is the regularization coefficient, are the parameters of the model; The loss function after adding the regularization term is: ; Among them, L is the original loss function, and are the weights of L1 and L2 regularization respectively; Fuse multiple trained models and synthesize the output results of multiple models; As new legal data is generated, the model is trained regularly; Information processing: receiving case-related information input by users, and analyzing and processing it; Information processing specifically includes the following parts: Using natural language processing technologies such as word segmentation, part-of-speech tagging, named entity recognition, and syntactic analysis, we conduct in-depth analysis of user input and extract key information, such as the core facts of the case, legal relationships, and the rights and obligations of the parties involved. Based on the extracted key information, a knowledge graph of the case is constructed to visualize the various entities in the case (such as parties, legal provisions, events, etc.) and the relationships between them; Establish an information verification rule library to verify the extracted key information according to legal provisions and business logic; For example, verifying whether the identity information of the parties is legal and whether the legal claims comply with legal provisions; Use logical reasoning algorithms to verify the logical consistency of information and check whether there are any contradictions or unreasonableness between the information; if any problems are found in the information, prompt the user to modify and supplement it in a timely manner; Document generation: Generate legal documents based on processed user input information and the big model; The document generation process specifically includes the following parts: Establish a legal document template library containing various types of legal document templates, such as complaint templates, defense templates, contract templates, etc. Each template has been reviewed and optimized by professional legal personnel to ensure its standard format and complete content; Based on user input of case type, legal claims, and other information, the system selects templates from a template library using rule matching and machine learning algorithms (such as decision trees and support vector machines). The system also adjusts and modifies templates based on the specific circumstances of the case and the user's personalized needs. Organize and optimize the content generated by the large model, segment, sort, and add titles to the content according to the logical structure and format requirements of the legal document to make it clear and well-organized; Use the legal rules engine to check the content of generated legal documents to ensure that they comply with laws and regulations and that the cited legal provisions are accurate; Use logical reasoning algorithms to check the logical rationality of the document content, check whether the argumentation in the document is reasonable and whether the evidence is sufficient; Submit the generated legal documents to professional legal personnel for manual review, and conduct comprehensive inspection and correction of the content, format, language expression, etc. to ensure that the quality of the documents reaches professional standards; Output and feedback: Output the generated legal documents to the user and receive feedback from the user; Output and feedback specifically include the following parts: Convert generated legal documents into common document formats, such as Word, PDF, TXT, etc., to meet different user needs; Provides various output setting options, such as font, font size, page layout, watermark, etc., allowing users to personalize the output documents according to their preferences; Support users to download the generated legal documents to local computer and provide printing function for user convenience; Develop a feedback interface and provide multiple feedback channels, such as online messages, email feedback, questionnaires, etc., to facilitate users to provide opinions and suggestions on the generated legal documents; Based on user feedback, the large model is further trained and optimized, the system parameters and algorithms are adjusted, and the system performance and user satisfaction are continuously improved.
[0023] A system for automatically generating legal documents through a large model, including the following parts: Data preparation and annotation module: This module obtains legal-related data, cleans it, removes noise, duplication, and erroneous information, and then unifies the text encoding. It then performs hierarchical annotation on the data to build a knowledge base. Model training module: After determining the candidate model, the accuracy of the model on tasks such as text classification and natural language inference is calculated using the language understanding benchmark test set. The pre-processed data is divided according to a preset ratio to update the model parameters. Multiple trained models are fused and the output results of multiple models are synthesized. Information processing module: Extracts key information after in-depth analysis of user input, builds a knowledge graph of the case, establishes an information verification rule base, and verifies the extracted key information according to legal provisions and business logic; Text module: Establish a legal document template library, select appropriate templates from the template library based on the case information input by the user, and adjust and modify the template according to the special circumstances of the case and the user's personalized needs. Generate the specific content of the legal document based on the processed user input information and the large model, check it, and convert the generated legal document into a common document format for output.
[0024] The above formulas are obtained by collecting a large amount of data and performing software simulation, and a formula close to the actual value is selected. The influencing weight factors and specific coefficient values in the formula are set by technical personnel in this field according to actual conditions, and can be adjusted and modified later.
[0025] The above description of the embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for automatically generating legal documents through a large model, characterized in that: Includes the following sections: Data preparation and annotation: Obtain open-source legal corpora, desensitized authentic documents, government announcements, and industry-standard contract templates; clean the data, unify the text encoding, and then annotate and build a knowledge base; Large model training: After the model is determined, the model structure and parameters are adjusted according to the characteristics and requirements of legal document generation. The pre-processed data is divided into training sets, validation sets, and test sets according to the preset ratios for training. Information processing: receiving case-related information input by users, and analyzing and processing it; Document generation: Generate legal documents based on processed user input information and the big model; Output and feedback: Output the generated legal documents to the user and receive feedback from the user.
2. The method for automatically generating legal documents through a large model according to claim 1, characterized in that: The data is cleaned, unified into text codes, annotated, and a knowledge base is constructed, specifically including: Labeling via hierarchical annotation: The entity layer marks the party information, time, and amount; The relationship layer marks the logical relationship between legal elements; Event layer, marking the legal event type; Building a knowledge base: Establishing a legal concept system; Extract relationships from legal texts through remote supervision, connect to legal database APIs, and implement regular updates of the knowledge base.
3. The method for automatically generating legal documents using a large model according to claim 1, characterized in that: The model training specifically includes three parts: model selection, model evaluation, and model optimization. The specific content of model selection includes the following parts: After determining the selected model, calculate the accuracy of the model on various tasks using the language understanding benchmark test set; Using automatic evaluation metrics, perplexity measures the model's ability to predict test data; Consider the model's parameter size, computational complexity, and hardware resource requirements; after selecting the base model, fine-tune the model based on the specific requirements of legal document generation.
4. The method for automatically generating legal documents using a large model according to claim 3, characterized in that: The model evaluation consists of the following parts: Divide the preprocessed data into training set, validation set and test set according to the preset ratio; After using the adaptive learning rate optimization algorithm, a preset number of samples are selected from the training set for training each time, and the average loss of the batch samples is calculated, and the parameters of the model are updated according to the optimization algorithm; During the training process, the model's loss value and accuracy indicators are monitored in real time; Draw a training curve to observe the training process of the model.
5. The method for automatically generating legal documents through a large model according to claim 3, characterized in that: The model optimization includes the following parts: Randomly discard a preset number of neurons in the model's neural network layer to prevent excessive dependence between neurons and improve the model's generalization ability; Add a regularization term to the loss function to penalize the complexity of the model; Fuse multiple trained models and synthesize the output results of multiple models; As new legal data is generated, the model is continuously trained on a regular basis.
6. The method for automatically generating legal documents through a large model according to claim 1, characterized in that: The information processing specifically includes the following parts: Use natural language processing technology to conduct in-depth analysis of user input and extract key information; Based on the extracted key information, a knowledge graph of the case is constructed to visualize the various entities in the case and their relationships; Establish an information verification rule library to verify the extracted key information according to legal provisions and business logic; Use logical reasoning algorithms to verify the logical consistency of information and check whether there are any contradictions or unreasonable points between the information; If any problems are found in the information, the user will be promptly prompted to modify and supplement it.
7. The method for automatically generating legal documents through a large model according to claim 1, characterized in that: The document generation specifically includes the following parts: Establish a legal document template library, which contains various types of legal document templates; each template has been reviewed and optimized by professional legal personnel; Based on the case type and legal claims entered by the user, a template is selected from the template library using rule matching and machine learning algorithms. The template is adjusted and modified based on the special circumstances of the case and the user's personalized needs. Organize and optimize the content generated by the large model, segment, sort, and add titles to the content according to the logical structure and format requirements of legal documents; Use the legal rules engine to check the content of generated legal documents; Use logical reasoning algorithms to check the logical rationality of the document content; The generated legal documents will be submitted to professional legal personnel for manual review, and the content, format, language expression and other aspects of the documents will be comprehensively checked and corrected.
8. The method for automatically generating legal documents through a large model according to claim 1, characterized in that: The output and feedback specifically include the following parts: Convert generated legal documents into common document formats; Provides various output setting options, allowing users to personalize the output documents according to their preferences; Support users to download the generated legal documents to local computer and provide printing function; Develop a feedback interface and provide multiple feedback channels; Based on user feedback, the large model is further trained and optimized, and the system parameters and algorithms are adjusted.
9. A system for automatically generating legal documents through a large model, using the method for automatically generating legal documents through a large model according to any one of claims 1 to 8, characterized in that: Includes the following sections: Data preparation and annotation module: This module obtains legal-related data, cleans it, removes noise, duplication, and erroneous information, and then unifies the text encoding. It then performs hierarchical annotation on the data to build a knowledge base. Model training module: After determining the candidate model, the accuracy of the model on tasks such as text classification and natural language inference is calculated using the language understanding benchmark test set. The pre-processed data is divided according to a preset ratio to update the model parameters. Multiple trained models are fused and the output results of multiple models are synthesized. Information processing module: Extracts key information after in-depth analysis of user input, builds a knowledge graph of the case, establishes an information verification rule base, and verifies the extracted key information according to legal provisions and business logic; Text module: Establish a legal document template library, select appropriate templates from the template library, generate the specific content of the legal document based on the processed user input information and the large model, and convert the generated legal document into a document format for output.
Citation Information
Cited By
Legal case abstract generation method and system based on knowledge guidance prompt fine tuning
CN121501992A
Large model-based legal text generation proofreading method and system
CN121920329A
Large model-based legal text generation proofreading method and system
CN121920329B