Judgment document summary generation method based on three-stage GRPO reinforcement learning
Through the three-stage GRPO reinforcement learning method, hierarchical training of large language models and combined with multiple reward function optimization, the problems of insufficient efficiency and accuracy in the generation of judicial document summaries are solved. The generated summaries have clear structure, accurate content and strong adaptability.
Patent Information
- Application Number
- CN202510758056.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-09
AI Technical Summary
Existing technologies have poor efficiency, accuracy and adaptability in generating summaries of judicial documents, making it difficult to effectively include key information and maintain the interpretability of the summaries.
A three-stage GRPO reinforcement learning method is adopted. By modeling a three-stage thinking chain, a large language model is trained with a layered data set. Multi-stage GRPO reinforcement learning and supervised fine-tuning are combined, and format rewards, language fluency rewards, content accuracy rewards, and context similarity rewards are designed to optimize the model generation process.
The efficiency, accuracy and adaptability of judicial document summary generation have been improved. The generated summaries have clear structure, accurate content and rigorous logic, and can adapt to inputs of various complexities, thereby improving the robustness and interpretability of the model.
Smart Images

Figure CN120278126B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of data processing technology, and in particular to a method for generating judicial document summaries based on three-stage GRPO reinforcement learning. Background Art
[0002] At present, the rapid development of large language models (LLMs) has brought new opportunities for text generation technology. A large number of NLP tasks are facing the dilemma of being completely conquered. For text summarization tasks that lack standard answers, especially the generation of judicial document summaries, how to reasonably include key information in the summary remains a challenge that needs to be solved urgently.
[0003] It can be seen that there is an urgent need for a judicial document summary generation method based on three-stage GRPO reinforcement learning with high generation efficiency, accuracy and adaptability. Summary of the Invention
[0004] In view of this, an embodiment of the present invention provides a method for generating judicial document summaries based on three-stage GRPO reinforcement learning, which at least partially solves the problems of poor generation efficiency, accuracy and adaptability in the existing technology.
[0005] The embodiment of the present invention provides a method for generating judicial document summaries based on three-stage GRPO reinforcement learning, comprising:
[0006] Step 1: Model a three-stage thinking chain;
[0007] Step 2: Distill and stratify the original judgment document dataset according to the three-stage thinking chain to obtain different types of datasets, where the types include high correlation, medium correlation, and low correlation;
[0008] Step 3: Use a highly relevant dataset to perform SFT supervised fine-tuning training on the large language model;
[0009] Step 4: Use the entire dataset to perform multi-stage GRPO reinforcement learning training on the trained large language model to obtain the target model;
[0010] Step 5: Input the target judgment document into the target model to generate a target summary.
[0011] According to a specific implementation of the embodiment of the present invention, step 1 specifically includes:
[0012] Step 1.1, defining a three-paragraph summary format, wherein the three-paragraph summary format includes entity extraction, analytical reasoning, and summary generation of the case;
[0013] Step 1.2, set the process of guiding the model to generate output content in a three-paragraph summary format through a predefined Prompt template, forming a three-paragraph thinking chain.
[0014] According to a specific implementation of the embodiment of the present invention, step 2 specifically includes:
[0015] Step 2.1: Use the large model as the teacher model and regenerate a three-stage reasoning chain summary for each judgment document in the original judgment document dataset based on the three-stage thinking chain to form a training sample;
[0016] Step 2.2: Use the AI model to perform a relevance score between the training sample and the original summary of each judgment document in the original judgment document dataset. The expression of the relevance score is:
[0017] ;
[0018] in, It is an AI model The relevance score of summary i is directly calculated, ranging from [0,1]. is the generated three-part reasoning chain summary, is the original abstract;
[0019] In step 2.3, the three-segment reasoning chain summaries of the training samples are sorted by relevance score, and then the sorted summaries are stratified according to different proportions to form a high relevance dataset, a medium relevance dataset, and a low relevance dataset.
[0020] According to a specific implementation of the embodiment of the present invention, step 3 specifically includes:
[0021] A large language model is trained using a preset amount of data from a high-correlation dataset. The training goal is to accurately map the input judicial documents to a standard three-segment structure output, and the cross-entropy loss function is used to fine-tune all parameters of the large language model.
[0022] According to a specific implementation of the embodiment of the present invention, step 4 specifically includes:
[0023] Step 4.1, set the data introduction strategy for multi-stage GRPO reinforcement learning training;
[0024] Step 4.2: Randomly select a three-segment reasoning chain summary from the high-relevance dataset as a context learning template;
[0025] Step 4.3: Set the formatting reward, language fluency reward, content accuracy reward, and context similarity reward to form the total reward;
[0026] Step 4.4, using the context-learning template to generate multiple candidate summaries corresponding to each judgment document in the original judgment document dataset;
[0027] Step 4.5, calculate the relative reward of each candidate summary based on the total reward;
[0028] In step 4.6, based on the relative reward and data, the strategy is introduced and the large language model is optimized through the policy gradient to obtain the target model.
[0029] According to a specific implementation of the embodiment of the present invention, step 4.1 specifically includes:
[0030] Step 4.1.1: Set the initial training phase to use 100% high-relevance dataset for multi-stage GRPO reinforcement learning training;
[0031] Step 4.1.2: Set the mid-term training phase to randomly select 70% high-correlation datasets and 50% medium-correlation datasets for multi-stage GRPO reinforcement learning training;
[0032] In step 4.1.3, in the later training stage, 40% of the high-correlation datasets, the remaining medium-correlation datasets, and all the low-correlation datasets are randomly selected for multi-stage GRPO reinforcement learning training.
[0033] According to a specific implementation of an embodiment of the present invention, step 4.3 specifically includes:
[0034] Step 4.3.1, set the format reward to
[0035] ;
[0036] in, Indicates the correctness of entity extraction labels, Indicates the correctness of the analytical reasoning label, Indicates the correctness of summary generated labels;
[0037] Step 4.3.2, set the language fluency reward to
[0038] ;
[0039] in, Indicates the summary generated by the pre-trained language model bert model The degree of confusion;
[0040] Step 4.3.3, set the content accuracy reward based on the similarity score reward, key entity coverage reward and source consistency reward as
[0041] ;
[0042] ;
[0043] ;
[0044] ;
[0045] in, Represents the similarity score reward, represents the key entity coverage reward, represents the source consistency reward, , , The similarity between the generated summary and the standard summary , , The corresponding weight, The entity collection in entity extraction, For the original text, is the key entity in analytical reasoning, is its importance weight, if the entity Appears in abstract middle, , otherwise 0;
[0046] Step 4.3.4, calculate the cosine similarity between the generated summary and the contextual learning template as the contextual similarity reward
[0047] ;
[0048] in, Learning templates for context, Learn the cosine similarity between templates for the generated summary and the context;
[0049] Step 4.3.5: Combine the format reward, language fluency reward, content accuracy reward, and context similarity reward to form the total reward.
[0050] .
[0051] According to a specific implementation of an embodiment of the present invention, the expression of the relative reward is:
[0052] ;
[0053] ;
[0054] in, Relative reward, Summary of the current candidate The total reward, is the average reward of all candidate summaries, Indicates the total number of current candidate summaries.
[0055] The judicial document summary generation scheme based on three-stage GRPO reinforcement learning in the embodiment of the present invention includes: step 1, modeling a three-stage thinking chain; step 2, performing data distillation and stratification on the original judicial document data set according to the three-stage thinking chain to obtain different types of data sets, wherein the types include high correlation, medium correlation and low correlation; step 3, using the high-correlation data set to perform SFT supervised fine-tuning training on the large language model; step 4, using the entire data set to perform multi-stage GRPO reinforcement learning training on the trained large language model to obtain a target model; step 5, inputting the target judicial document into the target model to generate a target summary.
[0056] The beneficial effects of the embodiments of the present invention are as follows: through the scheme of the present invention, a three-stage thinking chain structure is designed to imitate the logic of legal experts, and the summary is divided into three stages: entity extraction, analytical reasoning, and summary generation, thereby improving the transparency and interpretability of the model reasoning path; a high-quality data distillation and stratification mechanism is designed to generate reasoning chains and summary results through a large model, and the data is divided into three layers of high, medium, and low based on the relevance score to build a robust training process; a supervised fine-tuning and multi-stage GRPO training mechanism is constructed, first mastering the structured output format through the SFT training model, and then gradually introducing complex samples through multi-stage reinforcement learning, and combining multiple reward functions to optimize the model performance, thereby improving the efficiency, accuracy and adaptability of summary generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0058] Figure 1 A flowchart of a method for generating judicial document summaries based on three-stage GRPO reinforcement learning provided by an embodiment of the present invention;
[0059] Figure 2 This is a diagram showing the overall framework of a method for generating judicial document summaries based on three-stage GRPO reinforcement learning, provided by an embodiment of the present invention;
[0060] Figure 3 A schematic diagram of a data distillation and data stratification process provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0061] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0062] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0063] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on the present invention, those skilled in the art will appreciate that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement an apparatus and / or practice a method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this apparatus and / or practice this method.
[0064] It should also be noted that the illustrations provided in the following embodiments are merely schematic illustrations of the basic concept of the present invention. The illustrations only show components related to the present invention and are not drawn according to the number, shape, and size of components in actual implementation. In actual implementation, the type, quantity, and proportion of each component may be changed arbitrarily, and the component layout may also be more complex.
[0065] Additionally, in the following description, specific details are provided to provide a thorough understanding of the examples. However, one skilled in the art will appreciate that the aspects described can be practiced without these specific details.
[0066] A good summary should be detailed and entity-centric. However, traditional summary generation methods struggle to adapt to different types of legal texts, especially when dealing with complex judicial documents. They often fail to accurately extract key information, leading to information omissions or redundancy, and the quality of the generated summaries varies. Furthermore, most end-to-end deep learning-based models generate summaries directly, lacking a clear reasoning path. This makes it difficult to assess the credibility of the summary content and often fails to meet practical needs due to a lack of interpretability, low training efficiency, and slow convergence. In the field of natural language processing, and especially in judicial summary generation tasks, model interpretability is a key challenge. Existing research primarily uses Transformer-based pre-trained models (such as BERT and T5) to generate legal document summaries, but still faces challenges such as opaque reasoning logic, difficult-to-control output formats, and unstable summary quality.
[0067] An embodiment of the present invention provides a method for generating judicial document summaries based on three-stage GRPO reinforcement learning, which can be applied to the process of generating judicial document summaries in text data processing scenarios.
[0068] See also Figure 1 , is a flow chart of a method for generating judicial document summaries based on three-stage GRPO reinforcement learning provided by an embodiment of the present invention. Figure 1 and Figure 2 As shown, the method mainly includes the following steps:
[0069] Step 1: Model a three-stage thinking chain;
[0070] In specific implementation, the three-stage thinking chain modeling process mainly includes the following steps:
[0071] Step 1.1: Define a three-part summary format, corresponding to case entity extraction, logical reasoning process, and final summary content.
[0072] <entities>Extracting key legal entities from documents, such as people, time points, contract information, and legal clauses. The model first identifies and extracts key information elements from the input judgment document, including the main entities in the case (such as the parties, time, and cause of action) and key facts. This step is equivalent to the reading comprehension stage. The model scans the entire text and captures the names, places, events, and legal points that are most relevant to the case outcome and summary. This helps better grasp the context of the text and provides a foundation for subsequent reasoning.
[0073] <analysis>: Perform logical reasoning on the extracted entities, clarifying causal relationships and the basis for the judgment, and assigning weights to the entities. After obtaining the entities and facts, the model analyzes and infers these elements, constructing the logical context of the case, analyzing the importance of these entities in the text, and assigning weights to each entity. The model combines the extracted entities with the causal relationships, legal basis, and other content in the original text to infer the development of the case and the reasons for the judgment. At this stage, the model determines which entities are crucial to the generation of the summary and the relationship between them. This step is equivalent to the thinking that humans do before writing a summary: clarifying the context of the case and the core arguments;
[0074] <summary>: Based on the aforementioned analysis, the model generates the overall summary content, ensuring smooth language and clear key points. The model generates the final summary text based on the reasoning results of the previous stage, and generates a concise and accurate summary based on the weight of each entity in the text. Thanks to the support of the aforementioned analysis, the generated summary can specifically include the core information and conclusions of the judgment. At the same time, the model will organize the summary content according to the predetermined format requirements to ensure smooth language and clear expression. After this stage, the originally lengthy and complex judgment documents are compressed into a short and concise summary for users to read.
[0075] Step 1.2: Use the predefined prompt template to guide the model to generate output content according to the steps of "entity extraction → analysis and reasoning → summary generation". The prompt template specifically includes:
[0076] Please generate a summary of the judgment document according to the following format:
[0077] <entities> ...< / entities> <analysis> ...< / analysis> <summary>...< / summary> ;
[0078] Through the three-stage structure of entity extraction, analytical reasoning, and summary generation, the model is gradually guided from understanding to generalization. During the generation process, the model explicitly outputs intermediate results such as "entity extraction" and "analytical reasoning", making each step traceable. The model can generate more precise, clear summaries that contain rich entity relationships, significantly improving the model's interpretability and summary accuracy.
[0079] Step 2: Distill and stratify the original judgment document dataset according to the three-stage thinking chain to obtain different types of datasets, where the types include high correlation, medium correlation, and low correlation;
[0080] In practice, to ensure the quality of model training data, we introduced a high-quality data distillation strategy. Data distillation in this context refers to using a powerful teacher model to automatically generate intermediate reasoning processes and answers, passing them as a knowledge source to the student model, thereby extracting high-quality training samples. This approach provides the model with "expert-level" thinking examples without the need for manual annotation of reasoning processes, providing a reliable data foundation for subsequent model fine-tuning and reinforcement learning, improving the model's reasoning ability and the accuracy of the results. The specific process is as follows: Figure 3 As shown in the figure, the specific process of performing data distillation and stratification on the original judgment document dataset according to the three-stage thinking chain to obtain different types of datasets may include:
[0081] Step 2.1: Use the large model DeepSeek-R1 as the teacher model to regenerate the three-segment reasoning chain summary for the original judicial document dataset to form training samples.
[0082] Step 2.2: Use DeepSeek-V3 to score the relevance of the above samples and the original data summary.
[0083] ;
[0084] in, is the correlation score of sample i directly calculated by the model, which usually ranges from [0,1]. is the generated summary, is a standard reference summary.
[0085] Step 2.3: Sort and stratify by score into: high-relevance data (Top 50%); medium-relevance data (30%); low-relevance data (20%), providing a basis for subsequent phased training.
[0086] ;
[0087] in, Represents the relevance score of the jth sample after sorting (from high to low).
[0088] Step 2.3.1: Highly relevant dataset (top 50%): top score ranking data, of which:
[0089] ;
[0090] ;
[0091] For the size of the high-relevance dataset, the top 50% of the data with the highest scores are marked as high-relevance datasets. , this part of the data is closest to the standard answer and represents the distillation sample with the best quality.
[0092] Step 2.3.2: Medium relevance dataset (30%): score ranking between arrive Data between, where:
[0093] ;
[0094] ;
[0095] The size of the medium correlation data set, the remaining first 30% of the data is divided into medium correlation data ,These data have a high correlation with the standard answer, but slightly lower than the high correlation data, and are used for further generalization training of the model.
[0096] Step 2.3.3: Low relevance data (20%): Score ranking arrive Data between, where:
[0097] ;
[0098] ;
[0099] The size of the low-correlation data set, the last 20% of the data is divided into low-correlation data , these data have the lowest correlation with the standard answer and contain a certain amount of noise. They can be used as "challenging" samples to further enhance the robustness and generalization ability of the model.
[0100] Step 3: Use a highly relevant dataset to perform SFT supervised fine-tuning training on the large language model;
[0101] In practice, after constructing high-quality distillation data and designing a three-stage thought chain format, we first performed supervised fine-tuning (SFT) on the model to help it initially master the task format and basic capabilities. Through SFT, the model learns basic chain reasoning and summary generation patterns before reinforcement learning, laying a solid foundation and improving the efficiency of the reinforcement learning phase. In particular, in our task, the SFT phase teaches the model to adhere to the "three-stage" output format and style, which serves as a precursor to subsequent reinforcement learning format alignment.
[0102] We used the 30% highly relevant distilled data from the aforementioned screening for SFT, performing supervised fine-tuning of all model parameters to avoid interference from low-relevance samples. Using cross-entropy loss, the model learned to map the judgment document input to a three-step output (i.e., the results of entity extraction, analytical reasoning, and summary generation). SFT supervised fine-tuning provided a stable starting point for the model: it already possessed a certain level of task knowledge and format awareness, significantly reducing the difficulty and instability of the subsequent reinforcement learning phase. Because these training samples were inherently highly relevant and reliable, the model quickly learned the correct chain reasoning and the output format of the three-step thought chain during the SFT phase.
[0103] For example, the process of performing SFT supervised fine-tuning training is as follows:
[0104] Step 3.1: Select the first 30% of the highly correlated distilled data and use the Qwen2.5-3B model to fine-tune all parameters.
[0105] Step 3.2: The training goal is to accurately map the input document to a standard three-part output structure, optimized using a cross-entropy loss function. This phase aims to establish the model's format perception and basic reasoning capabilities, providing a robust starting point for reinforcement learning.
[0106] Step 4: Use the entire dataset to perform multi-stage GRPO reinforcement learning training on the trained large language model to obtain the target model;
[0107] In specific implementation, after completing SFT fine-tuning, we further optimized the model using multi-stage GRPO reinforcement learning to maximize summary quality, format consistency, and generalization capabilities. GRPO (Group Relative Policy Optimization) is an efficient and innovative reinforcement learning algorithm proposed by the DeepSeek-AI team in 2024. It is derived from the improvement of the policy optimization algorithm PPO. It can reduce reliance on independent value networks and use the relative ranking of grouped policy outputs to guide optimization. Compared with traditional reinforcement learning methods, GRPO is more stable in large-scale language model optimization and can effectively reduce computing resource consumption. The specific process of multi-stage GRPO reinforcement learning training is as follows:
[0108] Step 4.1: Multi-stage GRPO data ingestion strategy, including:
[0109] Step 4.1.1: Initial stage: Use 100% highly relevant data to stabilize the reinforcement learning process;
[0110] In this initial stage, only previously labeled high-relevance data is used for GRPO training. The model only needs to be optimized on high-quality samples. Its feedback signal is reliable and clear, which helps the model converge quickly and learn the core features of high-quality data.
[0111] Step 4.1.2: Mid-term: Randomly select 70% of the highly relevant data and introduce some moderately relevant data (accounting for 50%) to improve model generalization;
[0112] While continuing to train on highly relevant data, we also introduce moderately relevant data for optimization training. As the model gradually adapts to high-quality examples, we gradually increase the diversity of training samples, allowing the model to begin to handle more complex or less relevant cases, improving its generalization ability.
[0113] Step 4.1.3: Late stage: Randomly select 40% of the high-correlation data, the remaining medium-correlation data, and all low-correlation data to enhance robustness and adaptability to extreme cases.
[0114] The remaining low- or even extremely low-relevance data is incorporated into the training. At this point, the model has already learned strong summarization capabilities and format constraints in the first two stages. By training on difficult samples, its robustness and generalization capabilities can be further improved.
[0115] Step 4.2: Construct a contextual learning prompt and randomly select a high-quality example as a contextual learning example to add to the prompt. This helps the model understand the task requirements more quickly, improves the efficiency of reinforcement learning, and makes it easier for the model to follow the format requirements and master the three-stage thinking chain reasoning mode.
[0116] Step 4.3: Introduce four types of reward functions: format reward, language fluency reward, content accuracy reward, and context similarity reward.
[0117] Step 4.3.1: Format Reward: Encourages the model to produce output that conforms to the three-part thought chain structure. Rewards are awarded based on whether the output correctly distinguishes between the three components of "entity extraction," "analytic reasoning," and "summary generation." This ensures the model consistently delivers answers in the required format. This mechanism improves the controllability and interpretability of the results. The specific calculation method is as follows:
[0118] ;
[0119] When the model fully outputs the three-segment format, , otherwise the score is reduced according to the missing part.
[0120] in:
[0121] when <entities> ...< / entities> When the tag is closed correctly =1, otherwise 0;
[0122] when <analysis> ...< / analysis> When the tag is closed correctly =1, otherwise 0;
[0123] when <summary>...< / summary> When the tag is closed correctly =1, otherwise 0.
[0124] Step 4.3.2: Fluency Reward: This measures the language quality of the text generated by the model, including whether the wording is appropriate and whether the sentences are smooth and coherent. Using the pre-trained language model BERT, we calculate the perplexity (PPL) to measure the natural fluency of the summary. We give higher rewards to smooth and readable output. The specific calculation method is as follows:
[0125] ;
[0126] in, Represents the summary generated by the BERT model Lower perplexity indicates more fluent language expression, so its negative logarithm is taken as a reward, so that summaries with low PPL receive higher scores.
[0127] Step 4.3.3: Accuracy Reward: This focuses on evaluating the degree to which the summary content aligns with the facts in the original document and the reference summary. By calculating the degree of match between the model summary and the reference summary on key facts and legal points (e.g., using the ROUGE metric or training a discriminative model), we reward outputs that include key information and are semantically faithful, while penalizing outputs that deviate or omit information.
[0128] Accuracy rewards have three aspects:
[0129] Source consistency rewards, check <entities>Whether the entity in the tag comes from the original text:
[0130] ;
[0131] in, In the abstract <entities>Part of the entity collection, For the original text.
[0132] ROUGE score reward, which measures the similarity between the generated summary and the standard summary:
[0133] ;
[0134] in, , , They are , , The corresponding weight.
[0135] Key entity coverage reward, checks whether the summary contains high-weight key entities:
[0136] ;
[0137] in, yes <analysis>Key entities in the label, is its importance weight. Appears in abstract middle, =1 if the value is set to true, otherwise 0.
[0138] Final accuracy reward calculation method:
[0139] ;
[0140] Step 4.3.4: Contextual Similarity Reward: To encourage the model to imitate the reasoning style of contextual examples, we calculate the cosine similarity between the generated summary and the example summary:
[0141] ;
[0142] in, is a summary of contextual examples selected from a highly relevant dataset.
[0143] Step 4.3.5: The final total reward function combines all reward terms:
[0144] ;
[0145] in, This is the total reward applied to the model during each round of strategy optimization. Compared to a single reward signal, a multi-dimensional reward design comprehensively measures all aspects of summary quality, guiding the model to balance accurate content and fluent expression while maintaining correct formatting. The format reward accounts for 30%, ensuring that the model output conforms to the three-stage thought chain structure; the fluency reward accounts for 20%, encouraging smooth and natural language; the accuracy reward accounts for 40%, ensuring that the summary contains the correct key information and avoids false content; and the contextual similarity reward accounts for 10%, leveraging high-quality examples to accelerate learning. The final reward ranges from 0 to 10, with higher values indicating better summary quality.
[0146] Step 4.4: Candidate summary generation. For each training sample, multiple candidate summaries are generated based on the context template in step 4.2 for subsequent reward comparison.
[0147] Step 4.5: Relative reward calculation. The core idea of GRPO is to use the relative comparison of grouped policy outputs to guide optimization, that is, let the model generate multiple candidates for the same input, and then update the policy based on the relative merits of the candidates. In GRPO training, we do not directly use As a reward, instead use relative rewards to stabilize training:
[0148] ;
[0149] ;
[0150] in, Relative reward, is the total reward of the current candidate sample, is the average reward of all candidate summaries.
[0151] Step 4.6: Finally, optimize the objective through policy gradient to ensure that summaries with high rewards are more likely to be generated, while the probability of low-quality summaries is reduced.
[0152] ;
[0153] in, Maximize the expected reward E for the objective function and optimize the model parameters θ so that the probability of high-quality summaries increases. Relative reward, Representation model generates summary The logarithmic probability of .
[0154] Step 5: Input the target judgment document into the target model to generate a target summary.
[0155] In specific implementation, after the model training is completed, it can be used in online or offline summary service systems to generate three-paragraph summaries for massive judicial documents. The output results have the characteristics of standardized format, accurate content, and clear reasoning, and are suitable for scenarios such as legal assistants, smart courts, and public legal services. The final model shows obvious advantages in summary quality: First, thanks to the format reward and three-paragraph training, the summary output by the model has a clear structure, rigorous logic, and natural connection between each step; secondly, the guidance of fluency and accuracy rewards makes the summary not only fluent in language, but also highly faithful to the original information, covering the key information points in the judgment. More importantly, because it has been tested with samples of different relevance in reinforcement learning, the model has a strong adaptability to inputs of various complexities and its robustness has been significantly improved.
[0156] The method for generating judicial document summaries based on three-stage GRPO reinforcement learning provided in this embodiment designs a three-stage thinking chain structure, imitates the logic of legal experts, splits the summary into three stages: entity extraction, analytical reasoning, and summary generation, thereby improving the transparency and interpretability of the model reasoning path; designs a high-quality data distillation and stratification mechanism, generates reasoning chains and summary results through a large model, and divides the data into three layers: high, medium, and low based on relevance scores to build a robust training process; constructs a supervised fine-tuning and multi-stage GRPO training mechanism, first masters the structured output format through the SFT training model, then gradually introduces complex samples through multi-stage reinforcement learning, and combines multiple reward functions to optimize model performance, thereby improving the efficiency, accuracy, and adaptability of summary generation.
[0157] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.< / analysis> < / entities> < / entities> < / summary> < / analysis> < / entities>
Claims
1. A method for generating judicial document summaries based on three-stage GRPO reinforcement learning, characterized in that: include: Step 1: Model a three-stage thinking chain; The step 1 specifically includes: Step 1.1, defining a three-paragraph summary format, wherein the three-paragraph summary format includes entity extraction, analytical reasoning, and summary generation of the case; Step 1.2: Set the process of guiding the model to generate output content in a three-paragraph summary format using a predefined prompt template, forming a three-paragraph thought chain; Step 2: Distill and stratify the original judgment document dataset according to the three-stage thinking chain to obtain different types of datasets, where the types include high correlation, medium correlation, and low correlation; Step 3: Use a highly relevant dataset to perform SFT supervised fine-tuning training on the large language model; Step 4: Use the entire dataset to perform multi-stage GRPO reinforcement learning training on the trained large language model to obtain the target model; The step 4 specifically includes: Step 4.1, set the data introduction strategy for multi-stage GRPO reinforcement learning training; Step 4.2: Randomly select a three-part reasoning chain summary from the high-relevance dataset as a context learning template; Step 4.3: Set the formatting reward, language fluency reward, content accuracy reward, and context similarity reward to form the total reward; Step 4.4, using the context-learning template to generate multiple candidate summaries corresponding to each judgment document in the original judgment document dataset; Step 4.5, calculate the relative reward of each candidate summary based on the total reward; Step 4.6: Based on the relative rewards and data, introduce the strategy and optimize the large language model through policy gradient to obtain the target model; Step 5: Input the target judgment document into the target model to generate a target summary.
2. The method according to claim 1, characterized in that The step 2 specifically includes: Step 2.1: Use the large model as the teacher model and regenerate a three-stage reasoning chain summary for each judgment document in the original judgment document dataset based on the three-stage thinking chain to form a training sample; Step 2.2: Use the AI model to perform a relevance score between the training sample and the original summary of each judgment document in the original judgment document dataset. The expression of the relevance score is: ; in, It is an AI model The relevance score of summary i is directly calculated, ranging from [0,1]. is the generated three-part reasoning chain summary, is the original abstract; In step 2.3, the three-segment reasoning chain summaries of the training samples are sorted by relevance score, and then the sorted summaries are stratified according to different proportions to form a high relevance dataset, a medium relevance dataset, and a low relevance dataset.
3. The method according to claim 2, characterized in that The step 3 specifically includes: A large language model is trained using a preset amount of data from a high-correlation dataset. The training goal is to accurately map the input judicial documents to a standard three-segment structure output, and the cross-entropy loss function is used to fine-tune all parameters of the large language model.
4. The method according to claim 3, characterized in that The step 4.1 specifically includes: Step 4.1.1: Set the initial training phase to use 100% high-relevance dataset for multi-stage GRPO reinforcement learning training; Step 4.1.2: Set the mid-term training phase to randomly select 70% high-correlation datasets and 50% medium-correlation datasets for multi-stage GRPO reinforcement learning training; In step 4.1.3, in the later training stage, 40% of the high-correlation datasets, the remaining medium-correlation datasets, and all the low-correlation datasets are randomly selected for multi-stage GRPO reinforcement learning training.
5. The method according to claim 4, characterized in that Described step 4.3 specifically comprises: Step 4.3.1, set the format reward to ; in, Indicates the correctness of entity extraction labels, Indicates the correctness of the analytical reasoning label, Indicates the correctness of summary generated labels; Step 4.3.2, set the language fluency reward to ; in, Indicates the summary generated by the pre-trained language model bert model The degree of confusion; Step 4.3.3, set the content accuracy reward based on the similarity score reward, key entity coverage reward and source consistency reward as ; ; ; ; in, Represents the similarity score reward, represents the key entity coverage reward, represents the source consistency reward, , , The similarity between the generated summary and the standard summary , , The corresponding weight, The entity collection in entity extraction, For the original text, is the key entity in analytical reasoning, is its importance weight, if the entity Appears in abstract middle, =1, otherwise 0; Step 4.3.4, calculate the cosine similarity between the generated summary and the contextual learning template as the contextual similarity reward ; in, Learning templates for context, Learn the cosine similarity between templates for the generated summary and the context; Step 4.3.5: Combine the format reward, language fluency reward, content accuracy reward, and context similarity reward to form the total reward. 。 6. The method according to claim 5, characterized in that The expression of the relative reward is ; ; in, Relative reward, Summary of the current candidate The total reward, is the average reward of all candidate summaries, Indicates the total number of current candidate summaries.