Method and system for generating mediation document
Through deep learning and reinforcement learning, the mediation document generation method is optimized, combined with the legal corpus and mediator preferences, the flexibility and personalization of mediation document generation in the existing technology is solved, efficient and personalized mediation document generation is achieved, and the quality and efficiency of the document is improved.
Patent Information
- Application Number
- CN202510197709.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing mediation document generation methods rely on personal experience and fixed templates, which are difficult to adapt to the complex situations of different cases, lack flexibility and personalization, and lack of user feedback mechanisms, resulting in inconsistent document quality and inefficiency.
By obtaining case text data and legal corpus information, text preprocessing and naming entity recognition are carried out, preliminary mediation documents are generated by combining deep learning and sequence to sequence models, and standardized review and logical review are carried out, and document structure is optimized using reinforcement learning and language style transfer, mediator preference templates are constructed, and personalized adjustments are achieved.
It improves the efficiency and quality of mediation documents, ensures legal compliance, readability and personalization, and improves the intelligence and standardization of mediation work.
Smart Images

Figure CN120336375A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and more particularly, to a mediation document generation method and system thereof. Background Art
[0002] During the mediation process, the writing of mediation documents is an important link to ensure the accuracy, legality, and enforceability of the mediation results of cases. Traditional mediation documents mainly rely on the personal experience of mediators and manual editing, with low efficiency and being easily affected by individual knowledge levels and expression habits, resulting in uneven document quality. In addition, existing document generation methods usually rely on fixed template filling or simple text splicing, making it difficult to adapt to the complex situations of different cases, resulting in the generated documents lacking pertinence and flexibility. Some studies have attempted to use natural language processing technology for automated generation of mediation documents, but existing technologies mainly rely on rule-based methods or static deep learning models, and are unable to effectively handle semantic matching of case texts, logical consistency review, and legal clause adaptation issues. In addition, these methods often lack a user feedback mechanism, making it difficult to dynamically optimize the readability and professionalism of documents according to actual application requirements, and unable to meet the requirements for efficient, accurate, and personalized mediation document generation.
[0003] Therefore, there is an urgent need for a mediation document generation method and system to solve the above problems. Summary of the Invention
[0004] The purpose of the present invention is to provide a mediation document generation method and system to improve the above problems. To achieve the above purpose, the technical solutions adopted by the present invention are as follows:
[0005] In a first aspect, the present application provides a mediation document generation method, including:
[0006] Obtaining case text data information and legal information of a legal corpus;
[0007] Performing text preprocessing and named entity recognition on the case text data information, and performing association processing on the processed text information and the legal information to obtain structured case data;
[0008] Sending the structured case data to a mediation document generation model for processing, where case classification, semantic matching, and sequence-to-sequence model processing are performed on the structured case data to obtain a preliminary mediation document with placeholders;
[0009] Performing specification review and logical review processing on the preliminary mediation document with placeholders, where the specification of the preliminary mediation document is checked through the legal information, and the logical consistency analysis of the preliminary mediation document is reviewed to obtain a review result report;
[0010] Send the audit result report and the preliminary mediation document with placeholders to an optimization model for processing, where the optimization model is a model for text optimization of the preliminary mediation document with placeholders based on the audit result report, and obtain an optimized mediation document;
[0011] Construct a preference template for the optimized mediation document and the historical mediation document information of the preset mediators, and then obtain the mediation document corresponding to each mediator.
[0012] In a second aspect, the present application also provides a mediation document generation system, including:
[0013] An acquisition unit for acquiring case text data information and legal information of a legal corpus;
[0014] An identification unit for performing text preprocessing and named entity recognition on the case text data information, and performing associated processing on the processed text information and the legal information to obtain structured case data;
[0015] A generation unit for sending the structured case data to a mediation document generation model for processing, where case classification, semantic matching, and sequence-to-sequence model processing are performed on the structured case data to obtain a preliminary mediation document with placeholders;
[0016] An analysis unit for performing specification review and logical review processing on the preliminary mediation document with placeholders, where the specification of the preliminary mediation document is checked through the legal information, and the logical consistency analysis of the preliminary mediation document is reviewed to obtain an audit result report;
[0017] An optimization unit for sending the audit result report and the preliminary mediation document with placeholders to an optimization model for processing, where the optimization model is a model for text optimization of the preliminary mediation document with placeholders based on the audit result report, and obtain an optimized mediation document;
[0018] A construction unit for constructing a preference template for the optimized mediation document and the historical mediation document information of the preset mediators, and then obtaining the mediation document corresponding to each mediator.
[0019] The beneficial effects of the present invention are:
[0020] It can be understood that the present invention introduces a reinforcement learning optimization strategy, adjusts the document structure through the policy gradient method, optimizes the legal language using the style transfer technology, and combines the text generation adversarial network for semantic enhancement to improve the professionalism and logic of the document. The present invention improves the generation efficiency of the mediation document, and also optimizes the legal compliance, readability, and personalization level of the document, and enhances the intelligence and standardization of the mediation work.
[0021] Other features and advantages of the present invention will be described in the following specification. Moreover, some of them will become apparent from the specification or be understood by implementing the embodiments of the present invention. The objectives and other advantages of the present invention can be achieved and obtained by the structures specifically pointed out in the written specification, claims, and drawings. Description of the Drawings
[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0023] Figure 1 It is a schematic flowchart of the mediation document generation method described in the embodiments of the present invention;
[0024] Figure 2 It is a schematic structural diagram of the mediation document generation system described in the embodiments of the present invention.
[0025] In the figure: 701, acquisition unit; 702, recognition unit; 703, generation unit; 704, analysis unit; 705, optimization unit; 706, construction unit. Detailed Embodiments
[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but only represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0027] It should be noted that: Similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present invention, terms such as "first", "second", etc. are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.
[0028] Embodiment 1:
[0029] This embodiment provides a method for generating mediation documents.
[0030] See Figure 1 , which shows that this method includes steps S1, S2, S3, S4, S5 and S6.
[0031] Step S1: Obtain the case text data information and the legal information of the legal corpus;
[0032] It can be understood that in this step, the case text data information usually comes from court mediation records, legal documents submitted by lawyers, historical case libraries, etc., and the data forms include structured and unstructured texts. At the same time, the legal information of the legal corpus mainly includes laws and regulations, judicial interpretations, typical cases, etc.
[0033] Step S2: Perform text preprocessing and named entity recognition on the case text data information, and perform association processing on the processed text information and the legal information to obtain structured case data;
[0034] It can be understood that this step converts the case text into structured data, including basic case information, key legal terms, involved amounts, time nodes, etc., providing high-quality input data for subsequent document generation. The technical effect of this step is to achieve precise structuring of case information through the combination of deep learning and the legal information of the legal corpus, so that subsequent document generation is not only based on case facts, but also complies with legal norms, improving the accuracy and applicability of the documents. In this step, step S2 includes steps S21, S22, S23 and S24.
[0035] Step S21: Perform word segmentation and legal term recognition processing on the case text data information. Among them, through the pre-set BERT model, perform word segmentation on the case text data information in turn and extract legal terms related to the case text data in the legal corpus to obtain the initial word segmentation result and the legal term set;
[0036] It can be understood that in this step, first, in the word segmentation stage, based on the pre-trained BERT model, perform sub-word level segmentation on the case text to avoid incorrect word segmentation caused by out-of-vocabulary problems. For example, traditional word segmentation methods may split "liquidated damages" into "breach of contract" and "gold", while BERT can keep the term intact according to the context information. In addition, BERT word segmentation uses the WordPiece algorithm, which can dynamically adjust sub-word segmentation, so that low-frequency legal terms can still be correctly recognized.
[0037] Secondly, in the legal term recognition stage, a sequence labeling model based on BERT is used to extract legal terms. This model is fine-tuned on a legal corpus to identify specific legal concepts in case texts, such as "contract rescission", "force majeure", "civil compensation", etc. This task can be regarded as a named entity recognition problem, and a conditional random field model is adopted to enable the model to consider the relationship between words before and after when identifying terms, improving the recognition accuracy. In this step, by combining deep learning with a knowledge base, the recognition accuracy and integrity of legal terms are improved, ensuring that the case text data can be correctly mapped to the corresponding legal provisions in subsequent processing, and improving the accuracy and legal applicability of document generation.
[0038] Step S22: Name the core entities in the case text data information based on the preliminary word segmentation results and the legal term set, where the core entities include parties, lawyers, judges, legal provisions, amounts, times, and locations, to obtain a set of named core entities.
[0039] It can be understood that this step uses the BERT-CRF model to perform named entity recognition on the case text. BERT can model context semantic information, while the conditional random field can capture the dependencies between entities. For example, in the case of "Case of Contract Dispute between Li Moumou (Plaintiff) and Zhang Moumou (Defendant)", BERT can understand that "Li Moumou" and "Zhang Moumou" are personal names, and the conditional random field can use context clues to correctly classify them as "plaintiff" and "defendant".
[0040] Secondly, for entities such as amounts, times, and locations, since they usually have fixed formats (such as "5,000 yuan in RMB", "January 1, 2023"), regular expressions and rule-based matching methods can be combined for auxiliary recognition. For example, the amount entity can be matched by regular expression "[RMB|¥][\d,]+ yuan", the time entity can be matched by the format "\d{4} year \d{1,2} month \d{1,2} day", and the semantic consistency with the time expression can be confirmed by combining context information. In addition, for the recognition of legal provisions, the system will use the legal term set extracted in the previous step and combine it with a pre-constructed legal provision knowledge base for matching to ensure that the legal provisions mentioned in the case text can be accurately corresponded to the standard expressions in the legal corpus.
[0041] Step S23: Extract the case background statements in the case text data information based on the TextRank algorithm, and calculate the semantic similarity between the legal information in the legal corpus and the case background statements to obtain the case background statements that match the legal information in the legal corpus, and use them as structured case background data.
[0042] It can be understood that first, TextRank is an unsupervised keyword and important sentence extraction algorithm that analyzes the co-occurrence relationship of words in the text based on a graph model. Specifically, in this step, after splitting the case text into sentences, an undirected weighted graph is constructed, where each node represents a sentence, and the weight of the edge is calculated based on the word similarity between sentences. Subsequently, PageRank is used to iteratively calculate the importance score of each node (sentence), and several sentences with the highest scores are selected as the background sentences of the case. For example, in the case description of a contract dispute case, TextRank will automatically extract key background information such as the contract signing time, contract amount, and breach of contract situation. This step effectively improves the legal rigor of the mediation document, avoids the problem of the disconnection between the case background description and the legal basis, and thus lays a solid foundation for the subsequent legal clause matching and mediation document generation.
[0043] Step S24: Perform legal clause matching based on the structured case background data. Specifically, the applicable legal clauses are automatically retrieved through a pre-set graph neural network to obtain the final structured case data.
[0044] It can be understood that in this step, first, for the problem of legal clause matching, this step uses a graph neural network to construct a legal knowledge graph, and through the feature learning of nodes and edges, efficient legal clause recommendation is achieved. Specifically, the nodes of the knowledge graph include the core entities in the case background (such as contracts, breaches of contract, statute of limitations, etc.) and the key concepts of legal clauses, and the edges represent the logical or semantic relationships between entities, such as "contract - applicable clause - Article XX of the XX Law".
[0045] In the specific matching process, the structured case background data is first converted into a graph representation, and a pre-trained graph neural network is used for feature learning. Here, the graph neural network is used to update the embedding vector of the case background so that it can better capture the association between the case facts and legal clauses. For example, for the case background of "the payment for goods has not been made after the contract expires", the graph neural network can identify two key concepts, "contract" and "payment obligation", and find highly relevant matching content regarding liability for breach through the graph propagation mechanism.
[0046] In the legal clause recommendation stage, the output layer of the graph neural network adopts a Top-K sorting mechanism. According to the similarity calculation between the case background embedding and the legal clause embedding (such as dot product similarity or cosine similarity), the most relevant K legal clauses are screened out, and further screening is carried out in combination with the rules formulated by legal experts to ensure that the recommended clauses conform to legal logic. For example, if the case involves labor disputes, the graph neural network will give priority to recommending clauses in the XX Law rather than general clauses.
[0047] Step S3: Send the structured case data to the mediation document generation model for processing, where the structured case data is subject to case classification, semantic matching, and sequence-to-sequence model processing to obtain a preliminary mediation document with placeholders;
[0048] It can be understood that in this step, by case classification, it is ensured that the document generation conforms to the case type; by semantic matching, the reuse rate of historical mediation experience is improved; and by the Seq2Seq model, the legal compliance and language fluency of the mediation document are guaranteed. Compared with the traditional template filling method, this solution combines the natural language processing ability of deep learning, making the mediation document more intelligent and personalized, while reducing the cost of manual modification and providing a solid foundation for subsequent document optimization and adaptive adjustment. In this step, step S3 includes step S31, step S32, and step S33.
[0049] Step S31: Convert the structured case data into BERT word vectors, and input the BERT word vectors into a preset case classification model for classification. Among them, extract keywords from the BERT word vectors according to the numerical features and text features therein, and use the extracted keywords as the category labels of the case to mark the structured case data, obtaining the structured case data marked with the case category labels, where the numerical features are data of amount values and time values, and the text features are text data describing the case;
[0050] It can be understood that in this step, through the preset BERT model, this text data is converted into word vectors. The BERT model learns the dependency relationships between various words in the text in the global context through the self-attention mechanism. This enables the model to deeply understand the context of each word in the case text, thereby generating a more accurate word vector representation.
[0051] Secondly, the processing of numerical features is usually ignored in traditional text models, but in legal cases, these numerical features (such as amount values and time values) play an important differentiating role. To effectively combine numerical features and text features, this step adopts a feature fusion method. Specifically, the model first converts the text data into BERT word vectors and inputs the numerical features in the case as independent numerical values. Through feature fusion, while encoding the text data, the model takes into account the numerical information in the case, such as "the compensation amount is 50,000 yuan" or "the case occurred on April 5, 2023", and these numerical information will be introduced into the BERT model as additional features, enhancing the comprehensiveness of the case context.
[0052] Next, the model extracts keywords from the BERT word vectors. Keyword extraction is usually achieved through the attention mechanism. By focusing on the high-weight parts of the text, the model can automatically identify the keywords related to case classification (e.g., keywords such as "labor contract", "contract breach", or "debt dispute"). These keywords will serve as the category labels for the cases, thus providing a clear basis for subsequent case classification.
[0053] Finally, by combining the extracted keywords with the numerical features in the cases, the model labels the structured case data, assigning one or more category labels to each case. These category labels determine the legal area to which the case should be assigned, such as "labor dispute", "civil litigation", or "commercial contract", etc., thus providing an accurate template matching basis for the subsequent generation of mediation documents.
[0054] Step S32: Match the structured case data marked with the case category labels with the preset historical mediation document templates to obtain the preliminarily matched mediation document templates.
[0055] It can be understood that in this step, first, the system takes the marked case category labels as input and classifies the cases into the corresponding mediation document template sets according to the case types (e.g., labor disputes or debt disputes, etc.). Then, the system compares the historical mediation document templates with the case data based on the matching method of similarity calculation. For example, the cosine similarity is used to evaluate the matching degree between the case data and the templates, and the template that most conforms to the case description is selected.
[0056] Step S33: Input the preliminarily matched mediation document templates into the Seq2Seq model for legal language optimization, and input placeholders at the preset positions in the text to obtain the preliminary mediation documents with placeholders.
[0057] It can be understood that in this step, the Seq2Seq model will adjust the vocabulary, sentence structure, grammar, etc. in the mediation document to ensure that the output text conforms to the rigorous norms of legal texts. For example, the model will convert colloquial or ambiguous expressions (such as "as soon as possible") into more legally binding expressions (such as "within a reasonable time limit"). In addition, legal texts often need to cite specific legal provisions and case precedents, and the Seq2Seq model can automatically convert these contents into professional legal terms and standardized legal sentence patterns according to the context. Placeholders are usually used to identify the parts in the document that need to be filled according to specific case information, such as the names of the parties, the amount of the case, the issues in dispute, etc. By presetting placeholders in the document, the system not only ensures the integrity of the document structure but also facilitates subsequent automatic filling. Here, when generating the legally optimized document, the Seq2Seq model inserts placeholders at the corresponding positions according to the template design to ensure the flexibility and personalized needs of the document. This step improves the quality, professionalism, and flexibility of the mediation document through intelligent language optimization and placeholder insertion, laying a foundation for subsequent automated processing and personalized adjustment.
[0058] Step S4: Conduct a normative review and logical review process on the preliminary mediation document with placeholders. Specifically, check the norms of the preliminary mediation document through the legal information, and review the logical consistency analysis of the preliminary mediation document to obtain a review result report.
[0059] It can be understood that in this step, through the verification of legal information in the normative review, the legal accuracy of the document is ensured, and situations of incorrect legal citation or non-compliance with current legal provisions are avoided. Secondly, through the semantic and structural analysis of the document in the logical review, the logical consistency within the document is ensured, and semantic confusion and logical contradictions are avoided, thereby improving the overall quality of the document. This step greatly improves the professionalism and reliability of the document, making it more in line with the requirements of legal norms and mediation practices, and providing an important basic step for subsequent document optimization. Step S4 includes steps S41 and S42.
[0060] Step S41: Match the preliminary mediation document with placeholders and the legal information in the legal corpus. Specifically, analyze by extracting the legal terms in the document and the legal information to determine whether the legal terms are consistent with the legal information, and determine whether the case category labels corresponding to the legal provisions cited in the preliminary mediation document with placeholders conform to the legal provisions of the corresponding category in the legal information. If they conform, a preliminary mediation document that meets the norms is obtained.
[0061] It can be understood that this step will extract all legal terms from the preliminary mediation document, including but not limited to legal provision names, legal concepts, related terms, etc. Among them, through named entity recognition for term extraction, the system can accurately identify the legal terms involved in the document. For example, terms such as "liability for breach of contract", "contract termination clause", or "damages" mentioned in the mediation document will be extracted as objects for further verification.
[0062] This will be based on a pre-established legal knowledge base or legal corpus. The system determines whether there are inconsistent or misquoted situations by comparing the legal terms in the document with legal provisions, case precedents, and legal definitions in the knowledge base. For example, if a specific legal provision is mentioned in the document but it does not apply to the category of the current case, the system will mark it as a mismatch and provide modification suggestions. Through this step, the system ensures that the legal terms in the document are logically and legally consistent with the current legal system, avoiding improper citation or ambiguous expression. This step also determines whether the legal provisions cited in the document match the case category label. Each case involves specific types of legal provisions. For example, in civil cases, relevant provisions of the Civil Code can be cited, while in criminal cases, provisions of the Criminal Law need to be cited.
[0063] This step can effectively verify the accuracy of legal terms and provisions in the mediation document, ensure the correctness of the legal basis and citation in the document, and thus improve the standardization of the document. Especially in the process of large-scale automated generation of mediation documents, it can greatly reduce manual intervention and improve efficiency and accuracy.
[0064] Step S42: Detect the semantic consistency between each paragraph in the preliminary mediation document that meets the standardization, and determine whether there are inconsistent contents before and after and whether the time sequence conforms to the preset logic, to obtain an audit result report of the preliminary mediation document that meets the standardization.
[0065] It can be understood that this step will conduct semantic analysis on each paragraph of the document to verify whether there are logical contradictions or content inconsistencies between paragraphs. For example, in the document, if the previous paragraph mentions that the parties agree to pay a certain amount of compensation, and the subsequent paragraph shows that the compensation is not mentioned or has been modified, such an inconsistency before and after will be marked as an error by the system. In addition, the system will also detect whether there is information omission or duplication. Through context reasoning and semantic similarity calculation, the system can intelligently identify potential inconsistencies between paragraphs and even determine whether the narration of case-related facts is consistent with the case background.
[0066] Next, the system will also detect the temporality in the document. The logical structure of a mediation document usually unfolds in a certain time sequence. For example, contents such as the case background, fact confirmation, liability determination, mediation suggestions, etc. should appear in a predetermined order. The system analyzes the time expressions and event descriptions in the document to determine whether the time sequence in each paragraph conforms to the development order of the actual situation. If there are time contradictions in the document or the description order of events is inappropriate (for example, the mediation result is mentioned before the case facts), the system will mark it as a potential problem and remind that adjustment is needed.
[0067] Step S5: Send the audit result report and the preliminary mediation document with placeholders to the optimization model for processing, where the optimization model is a model for optimizing the text of the preliminary mediation document with placeholders based on the audit result report, and obtain an optimized mediation document;
[0068] It can be understood that this step can significantly improve the quality and efficiency of the mediation document. Especially while ensuring the standardization and legal accuracy of the document, the finally generated mediation document not only meets the legal requirements, but also has a more concise, clear and logically rigorous language, thus improving the efficiency of the mediation work and reducing the burden of manual review and modification. In this step, step S5 includes step S51, step S52 and step S53.
[0069] Step S51: According to the audit result report and the preliminary mediation document with placeholders, adopt a reinforcement learning optimization method to adjust the text structure, and optimize the chapter division, paragraph organization and logical order of the document through the policy gradient method to obtain a mediation document with optimized structure;
[0070] It can be understood that in this step, through the policy gradient method, the optimization model can learn how to weigh different chapters and paragraphs in text adjustment to achieve the best structure adjustment. Specifically, the policy gradient method calculates the reward value of each action (i.e., text adjustment or paragraph rearrangement) according to the current structure and logical relationship of the document, and uses these reward signals to guide the optimization process.
[0071] In this step, the reinforcement learning model can be expressed by the following formula:
[0072] MDP=(S,A,P,R,γ)
[0073] Where MDP is the Markov decision process, S is the state space, representing the structure, paragraph order and logical relationship of the current mediation document, A is the action space, including paragraph exchange, chapter adjustment, logical connection optimization, P is the state transition probability, R is the reward function, and γ is the discount factor, which is used to balance short-term rewards and long-term rewards.
[0074] In specific implementation, the reinforcement learning model is trained with a series of training data and document samples, and gradually learns effective strategies for text structure optimization. For example, the model can identify which chapters or paragraphs in the mediation document have unreasonable arrangement orders, or which information should be mentioned earlier to enhance logical clarity, and then make appropriate adjustments. These adjustments not only include changes in language or format, but also content optimization, making the mediation document more compliant from a legal perspective and more well-organized.
[0075] Step S52: According to the mediation document with optimized structure, use the language style transfer technology to optimize the legal language. Through the StyleGAN and GYAFC models, convert the expression sentences into legal professional terms to obtain the mediation document with optimized language.
[0076] It can be understood that the StyleGAN model, as a generative adversarial network, is mainly used to generate text with a specific style. Here, it can convert the ordinary language in the mediation document into more formal and standardized legal language. Among them, the StyleGAN model learns the style characteristics (such as word usage, sentence structure, etc.) of a large number of legal documents, and converts ordinary expressions into more legal and rigorous expressions. By analyzing and learning the differences between the informal corpus and the formal corpus, the GYAFC model can automatically identify the informal expressions in the text and adjust the expressions according to the specifications of legal documents to meet the formality requirements of legal language. The main components of StyleGAN include: a mapping network that inputs a random noise and converts it into a style vector through multiple fully connected layers, which is used to control the semantic style of text generation. An adaptive normalization layer that affects the normalization operation of each layer of the generator through the style vector, enabling the model to learn the language styles of different mediators and control the sentence structure. The core structure of GYAFC includes a Seq2Seq structure and an attention mechanism. The Seq2Seq structure includes: an encoder that encodes the input informal text into a context vector. A decoder that generates formalized text based on the context vector to ensure the rigor of legal language. The attention mechanism: Through the attention mechanism, the model can focus on key information in the input text, such as legal terms, core entities (person names, amounts, time), etc., to ensure that the converted text still meets legal requirements.
[0077] Step S53: According to the mediation document with optimized language, use the text generation adversarial network to enhance the semantics. Through generative adversarial training, supplement the citation basis and cited cases in the legal information to obtain the mediation document with optimized language.
[0078] It can be understood that this step of the text generation adversarial network mainly consists of two parts: a generator and a discriminator. The generator is responsible for generating new text content, while the discriminator is used to evaluate whether the generated content meets the actual legal language standards and logical consistency. Based on the content of the input legal document and the existing legal corpus, the generator generates supplementary legal bases and case precedents related to the document content. These cited bases and case precedents usually include judgment citations, relevant legal provisions, past cases, etc. By learning the expression patterns and structures of legal language in the document, the generator automatically supplements appropriate legal citations to ensure the completeness of the legal information in the document. The task of the discriminator is to judge whether the generated cited bases and case precedents are reasonable, accurate, and in line with the context and legal background of the document. Through adversarial training with the generator, the discriminator gradually learns how to identify legal content that meets the requirements of the document, thus ensuring that the generated legal information not only has grammatical and linguistic correctness but also ensures the relevance and authority of legal provisions or case precedents.
[0079] Step S6: Construct a preference template from the optimized mediation document and the historical mediation document information of the preset mediator, and then obtain the mediation document corresponding to each mediator.
[0080] It can be understood that through constructing a preference template in this step, the generated mediation document can be automatically adjusted according to the working style, language preference, and historical cases of the mediator, making it more in line with the personalized requirements of the mediator. This technology not only improves the efficiency of document generation but also greatly enhances the job satisfaction of the mediator and reduces the need for manual intervention. In this step, step S6 includes step S61, step S62, and step S63.
[0081] Step S61: Use the Few-shot Learning algorithm to learn the preference information of the historical mediation document information based on the historical mediation document information of the mediator. Among them, by extracting the preference information of the historical mediation document information, the preference information includes lexical information, sentence patterns, and frequently occurring words;
[0082] It is understandable that in this step, first, through a comprehensive analysis of the mediator's historical mediation documents, lexical information (such as common words, professional terms), syntactic structures (such as common grammatical structures, sentence lengths, collocation methods of sentence patterns), and frequently-occurring words (such as words that appear frequently in multiple documents) are systematically extracted. This information will help the algorithm identify the expression patterns and language structures that mediators tend to use more in their actual work. The Few-shot Learning technique is used to identify and record the words commonly used by mediators, legal terms, and words with a higher frequency of occurrence in specific cases. These words may include specific legal terms, case keywords, and sentence modifiers commonly used by mediators. Extract the syntactic structures commonly used by mediators, such as common attributive clauses, adverbial clauses, parallel structures, etc., as well as the length and complexity of sentences. Extract the words and phrases frequently used by mediators in multiple documents to understand their tendencies. The extracted lexical information, syntactic structures, and frequently-occurring words are encoded through natural language processing techniques (such as word vector models, syntactic analysis) to form a structured preference information vector.
[0083] Step S62: Among them, the weight of each piece of lexical information in the preference template is adjusted through a meta-learning algorithm, and a template generation strategy is adjusted according to the previously learned syntactic structure to obtain a mediation document preference template corresponding to each mediator.
[0084] It is understandable that in this step, first, a meta-learning algorithm is used to analyze the mediator's historical mediation documents to extract high-frequency words, specific expression patterns, and syntactic structures. The importance of each word is calculated through TF-IDF and word embedding models. In terms of syntactic structure optimization, dependency syntactic analysis is used to extract common expression patterns and classify them through a clustering algorithm to form a personalized syntactic database. Subsequently, reinforcement learning is combined to optimize the selection strategy of the document template to ensure that the generated mediation document conforms to the mediator's expression style. Finally, based on these learning results, a personalized document preference template for the mediator is dynamically constructed and adaptively adjusted during the subsequent document generation process, and continuous optimization is supported as the usage data increases.
[0085] Step S63: The optimized mediation documents are corrected according to the mediation document preference template corresponding to each mediator to obtain the mediation documents corresponding to each mediator.
[0086] It is understandable that in this step, a document that conforms to both legal norms and the mediator's personal style is generated, thereby improving the acceptability and usage efficiency of the mediation document. At the same time, as the system continuously learns the mediator's revision habits, its personalized adjustment ability will also be continuously optimized, making the documents generated in the future more accurately meet the mediator's needs.
[0087] Example 2:
[0088] As shown Figure 2 in the figure, this embodiment provides a mediation document generation system. Refer to Figure 2 the system includes an acquisition unit 701, an identification unit 702, a generation unit 703, an analysis unit 704, an optimization unit 705, and a construction unit 706.
[0089] The acquisition unit 701 is configured to acquire case text data information and legal information of a legal corpus;
[0090] The identification unit 702 is configured to perform text preprocessing and named entity recognition on the case text data information, and perform association processing on the processed text information and the legal information to obtain structured case data;
[0091] The generation unit 703 is configured to send the structured case data to a mediation document generation model for processing, where case classification, semantic matching, and sequence-to-sequence model processing are performed on the structured case data to obtain a preliminary mediation document with placeholders;
[0092] The analysis unit 704 is configured to perform specification review and logic review processing on the preliminary mediation document with placeholders. Specifically, the specification of the preliminary mediation document is checked through the legal information, and the logical consistency analysis of the preliminary mediation document is reviewed to obtain a review result report;
[0093] The optimization unit 705 is configured to send the review result report and the preliminary mediation document with placeholders to an optimization model for processing, where the optimization model is a model for optimizing the text of the preliminary mediation document with placeholders based on the review result report to obtain an optimized mediation document;
[0094] The construction unit 706 is configured to construct a preference template from the optimized mediation document and the historical mediation document information of a preset mediator, and then obtain a mediation document corresponding to each mediator.
[0095] It should be noted that regarding the system in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment related to the method, and will not be elaborated here.
[0096] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
[0097] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for generating mediation documents, characterized in that, Including: Obtaining the case text data information and the legal information of the legal corpus; Performing text preprocessing and named entity recognition on the case text data information, and performing correlation processing on the processed text information and the legal information to obtain structured case data; Sending the structured case data to a mediation document generation model for processing, where case classification, semantic matching, and sequence-to-sequence model processing are performed on the structured case data to obtain a preliminary mediation document with placeholders; Performing specification review and logical review processing on the preliminary mediation document with placeholders, where the specification of the preliminary mediation document is checked through the legal information, and the logical consistency analysis of the preliminary mediation document is reviewed to obtain a review result report; Sending the review result report and the preliminary mediation document with placeholders to an optimization model for processing, where the optimization model is a model for optimizing the text of the preliminary mediation document with placeholders based on the review result report to obtain an optimized mediation document; Constructing a preference template for the optimized mediation document and the historical mediation document information of the preset mediator, and then obtaining the mediation document corresponding to each mediator.
2. The mediation document generation method according to claim 1, wherein , The performing text preprocessing and named entity recognition on the case text data information, and performing correlation processing on the processed text information and the legal information to obtain structured case data, including: Performing word segmentation and legal term recognition processing on the case text data information, where word segmentation is sequentially performed on the case text data information through a preset BERT model, and legal terms related to the case text data in the legal corpus are extracted to obtain a preliminary word segmentation result and a legal term set; Naming the core entities in the case text data information based on the preliminary word segmentation result and the legal term set, where the core entities include parties, lawyers, judges, legal provisions, amounts, times, and locations, to obtain a set of named core entities; Extracting the case background statements in the case text data information based on the TextRank algorithm, and calculating the semantic similarity between the legal information of the legal corpus and the case background statements to obtain case background statements matching the legal information of the legal corpus, and using them as structured case background data; Performing legal provision matching according to the structured case background data, where the applicable legal provisions are automatically retrieved through a preset graph neural network to obtain the final structured case data.
3. The mediation document generation method according to claim 1, wherein , Sending the structured case data to a mediation document generation model for processing, where case classification, semantic matching, and sequence-to-sequence model processing are performed on the structured case data to obtain a preliminary mediation document with placeholders, including: Convert the structured case data into BERT word vectors, and input the BERT word vectors into a preset case classification model for classification. Among them, extract keywords from the BERT word vectors according to the numerical features and text features therein, and use the extracted keywords as the category labels of the cases to label the structured case data, obtaining structured case data labeled with case category labels. Among them, the numerical features are data of amount values and time values, and the text features are text data describing the cases; Match the structured case data labeled with case category labels with a preset historical mediation document template to obtain a preliminarily matched mediation document template; Input the preliminarily matched mediation document template into a Seq2Seq model for legal language optimization, and input placeholders at preset positions in the text to obtain a preliminary mediation document with placeholders; 4. The mediation document generation method according to claim 3, characterized in that ,Perform a specification review and a logical review process on the preliminary mediation document with placeholders. Among them, check the specification of the preliminary mediation document through the legal information, and review the logical consistency analysis of the preliminary mediation document to obtain a review result report, including: Match the preliminary mediation document with placeholders with the legal information in the legal corpus. Among them, analyze by extracting legal terms in the document and the legal information to determine whether the legal terms are consistent with the legal information, and determine whether the case category label corresponding to the legal clause cited in the preliminary mediation document with placeholders conforms to the legal clause of the corresponding category in the legal information. If it conforms, obtain a preliminary mediation document that meets the specification; Detect the semantic consistency between each paragraph in the preliminary mediation document that meets the specification, and judge whether there are inconsistent contents before and after and whether the time sequence conforms to the preset logic, obtaining a review result report of the preliminary mediation document that meets the specification; 5. The mediation document generation method according to claim 1, characterized in that ,Send the review result report and the preliminary mediation document with placeholders to an optimization model for processing. Among them, the optimization model is a model for optimizing the text of the preliminary mediation document with placeholders based on the review result report, obtaining an optimized mediation document, including: According to the review result report and the preliminary mediation document with placeholders, adopt a reinforcement learning optimization method to adjust the text structure, and optimize the chapter division, paragraph organization and logical order of the document through the policy gradient method, obtaining a mediation document with an optimized structure; According to the mediation document with an optimized structure, adopt a language style transfer technology to optimize the legal language. Through the StyleGAN and GYAFC models, convert the expression sentences into legal professional terms, obtaining a mediation document with optimized language; According to the mediation document with optimized language, adopt a text generation adversarial network to enhance the semantics. Through generative adversarial training, supplement the citation basis and citation cases in the legal information, obtaining a mediation document with optimized language; 6. A mediation document generation system, characterized in that, Including: An acquisition unit for acquiring case text data information and legal information in a legal corpus; An identification unit, configured to perform text preprocessing and named entity recognition on the case text data information, and perform association processing on the processed text information and the legal information to obtain structured case data; A generation unit, configured to send the structured case data to a mediation document generation model for processing, where case classification, semantic matching, and sequence-to-sequence model processing are performed on the structured case data to obtain a preliminary mediation document with placeholders; An analysis unit, configured to perform specification review and logical review processing on the preliminary mediation document with placeholders, where the specification of the preliminary mediation document is checked through the legal information, and the logical consistency analysis of the preliminary mediation document is reviewed to obtain a review result report; An optimization unit, configured to send the review result report and the preliminary mediation document with placeholders to an optimization model for processing, where the optimization model is a model for optimizing the text of the preliminary mediation document with placeholders based on the review result report to obtain an optimized mediation document; A construction unit, configured to construct a preference template from the optimized mediation document and the historical mediation document information of a preset mediator, so as to obtain a mediation document corresponding to each mediator.
7. The mediation document generation system according to claim 6, wherein The identification unit includes: A first identification subunit, configured to perform word segmentation and legal term recognition processing on the case text data information, where word segmentation is sequentially performed on the case text data information through a preset BERT model, and legal terms related to the case text data in a legal corpus are extracted to obtain a preliminary word segmentation result and a legal term set; A second identification subunit, configured to name the core entities in the case text data information based on the preliminary word segmentation result and the legal term set, where the core entities include parties, lawyers, judges, legal provisions, amounts, times, and locations, to obtain a set of named core entities; A third identification subunit, configured to extract case background statements in the case text data information based on the TextRank algorithm, and calculate the semantic similarity between the legal information in the legal corpus and the case background statements to obtain case background statements matching the legal information in the legal corpus, and use them as structured case background data; A fourth identification subunit, configured to perform legal provision matching according to the structured case background data, where applicable legal provisions are automatically retrieved through a preset graph neural network to obtain the final structured case data.
8. The mediation document generation system according to claim 6, wherein The generation unit includes: A first generation subunit, configured to convert the structured case data into BERT word vectors, and input the BERT word vectors into a preset case classification model for classification, where the BERT word vectors are used to extract keywords according to the numerical features and text features therein, and the structured case data is marked with the extracted keywords as case category labels to obtain structured case data marked with case category labels, where the numerical features are data of amount values and time values, and the text features are text data describing the case; A second generation subunit, configured to match the structured case data marked with case category labels with a preset historical mediation document template to obtain a preliminarily matched mediation document template; A third generation subunit, configured to input the preliminarily matched mediation document template into a Seq2Seq model for legal language optimization, and input placeholders at preset positions in the text to obtain a preliminary mediation document with placeholders; 9. The mediation document generation system according to claim 8, wherein The analysis unit includes: A first analysis subunit, configured to match the preliminary mediation document with placeholders with legal information in a legal corpus. Specifically, by extracting legal terms in the document and analyzing them with legal information, it is determined whether the legal terms are consistent with the legal information, and whether the case category labels corresponding to the legal clauses cited in the preliminary mediation document with placeholders conform to the legal clauses of the corresponding categories in the legal information. If they conform, a preliminary mediation document that meets the norms is obtained; A second analysis subunit, configured to detect the semantic consistency between each paragraph in the preliminary mediation document that meets the norms, determine whether there are inconsistent contents before and after and whether the time sequence conforms to the preset logic, and obtain an audit result report of the preliminary mediation document that meets the norms; 10. The mediation document generation system according to claim 6, characterized in that, The optimization unit includes: A first optimization subunit, configured to perform text structure adjustment according to the audit result report and the preliminary mediation document with placeholders, and optimize the chapter division, paragraph organization and logical order of the document through the policy gradient method to obtain a mediation document with optimized structure; A second optimization subunit, configured to perform legal language optimization on the mediation document with optimized structure by using a language style transfer technique. Through the StyleGAN and GYAFC models, the expression statements are converted into legal professional terms to obtain a mediation document with optimized language; A third optimization subunit, configured to perform semantic enhancement on the mediation document with optimized language by using a text generation adversarial network. Through generative adversarial training, the citation basis and citation cases in the legal information are supplemented to obtain a mediation document with optimized language.
Citation Information
Cited By
Method and system for automatically generating personalized mediation protocol based on mediation case characteristics
CN119963374A
Method and system for automatically generating individualized mediation agreement based on mediation case characteristics
CN119963374B
Formatted document intelligent generation method and system
CN121009858A
Text specification marking method and system based on rule corpus
CN121072515A
Multi-modal large model-based law enforcement case handling document intelligent auditing system and method
CN121212989A