Generative legal text automatic abstracting method and device and storage medium
By introducing a comparative learning mechanism of multi-level data augmentation and BYOL framework in the Chinese legal text automation abstract, the legal text representation ability of the mBART model is optimized, and the problems of semantic distortion, insufficient professional term processing and inefficient generation of Chinese legal text abstracts are solved, and high-quality and efficient legal text abstract generation is achieved.
Patent Information
- Application Number
- CN202510306792.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-15
- Publication Date
- 2025-06-24
Smart Images

Figure CN120197606A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to generative artificial intelligence, and provides a generative automated legal text summarization method, device, and storage medium. Background Art
[0002] In recent years, with the continuous disclosure of judicial big data represented by judgment documents and the continuous breakthrough of natural language processing technology, how to apply artificial intelligence technology in the judicial field has gradually become a research hotspot in legal intelligence. People increasingly realize the indispensability of legal knowledge in daily life. However, due to the overly complex legal knowledge system, in most cases, people hope to obtain a concise summary to provide the main information. Combining natural language processing technology with legal data has extremely strong practical significance.
[0003] Text summarization solves this problem by compressing large legal documents into a shorter and more concise format while retaining the basic information and meaning of the original document. It enables individuals to quickly understand the key elements of a case, including the parties involved, the legal issues being resolved, the main arguments and evidence presented, and the outcome of the case. Currently, judicial text summarization is mainly manually completed by professional legal editors, which consumes a large amount of time and effort. When dealing with a large number of cases, legal practitioners such as lawyers and judges usually need to consult a large number of judgment documents and legal instruments, manually read these complex and lengthy texts to find relevant cases to support their judgments. In the face of urgent cases, the inability to obtain key information in a timely manner may affect the handling of the cases. Moreover, for clients, these time-consuming tasks bring them high costs.
[0004] The legal text automatic summarization technology compresses large legal documents into a more concise form, retains their core information and meaning, and helps users quickly understand the key elements of a case. It is an important tool for judicial personnel to quickly grasp the essence of a case and make wise decisions. Automated summarization tools can greatly improve the efficiency of document processing, ensure that legal practitioners and the public can quickly obtain the required information, and avoid processing delays and errors caused by information overload.
[0005] Currently, the common technical routes for automated text summarization mainly include extractive, generative, or hybrid. The extractive method selects the most important sentences from the input text and uses them to generate a summary. The generative method represents the input text in an intermediate form and then generates a summary using words and sentences different from those in the original text. The hybrid method combines the extractive and generative methods. The mainstream text summarization tools mainly include: summary tools based on the extractive Textrank algorithm, summary tools based on generative deep learning models such as BART and T5, and large language models such as ChatGPT for text summarization.
[0006] Although the current text summarization has relatively perfect applications, there are still the following two deficiencies in the summarization of legal texts:
[0007] There are few applications for summarizing Chinese legal texts: Currently, most of the legal text summarization tools developed by mainstream companies are mainly for Indian legal documents and US legal documents, and there are few summarizations for Chinese legal documents. Most domestic summarization tools are for traditional texts such as news texts, etc., and there is a certain gap in the intelligence of the legal field.
[0008] Single technical route: Currently, the legal text summarization in China mainly uses LLM (Large Language Model), and traditional deep learning models are less used. For example, Tongyi Fari developed by Alibaba Cloud and Fazhi developed by Flush are both based on their respective LLMs, but the effect of LLMs in the actual execution process may not be as good as some traditional deep learning models.
[0009] The present invention solves the above two problems based on an improved generative summarization method, introduces the idea of contrastive learning to construct a set of generative contrastive learning summarization frameworks, conducts contrastive learning by constructing positive example pairs and maximizing the similarity between positive example pairs, so that the mBART model learns a good representation of legal texts and then fine-tunes the decoder, enabling the mBART model with the BERT model structure as the encoder to fully understand the meaning of legal texts, and being able to generate semantically consistent summaries through the mBART model with the GPT model structure as the decoder. Summary of the Invention
[0010] The present invention aims to solve the problems of semantic distortion, insufficient processing of professional terms, and low generation efficiency in the automated summarization of Chinese legal texts. By constructing a multi-level data augmentation strategy (character / punctuation / sentence-level perturbation) to generate semantically consistent positive example pairs, combining the negative-sample-free contrastive learning mechanism of the BYOL framework to optimize the legal text representation ability of the mBART encoder, and adopting hierarchical learning rate fine-tuning to achieve encoding-decoding collaborative optimization, ultimately improving the semantic fidelity, professional term accuracy, and generation efficiency of legal summaries, and overcoming the defects of poor adaptability of traditional methods to Chinese legal texts, high consumption of large model resources, and slow generation speed. To achieve the above objectives, the present invention adopts the following technical means:
[0011] The present invention provides a method for automated summarization of legal texts based on generative artificial intelligence, including the following steps:
[0012] Step 1: Perform multi-level data augmentation on the original Chinese legal text to generate enhanced texts to form positive example pairs;
[0013] Step 2: Input the positive example pairs into the online encoder and the target encoder of the BYOL framework respectively. The online encoder and the target encoder adopt the encoder structure of the pre-trained mBART model and have the same initial parameters;
[0014] Step 3: Respectively perform dimensional mapping on the encoding results of the online side and the target side through a multi-layer perceptron to obtain mapped vectors;
[0015] Step 4: Set a predictor on the online side. After normalizing the mapped vectors, calculate the cosine similarity between the online prediction vector and the target mapped vector, and calculate the contrast loss function;
[0016] Step 5: Update the parameters of the online side by gradient, update the parameters of the target side through the EMA mechanism, and only update the encoder of the mBART model while freezing the decoder parameters;
[0017] Step 6: Save the encoder of the trained mBART model on the online side, extract the mBART model, fine-tune it on legal texts, and set the learning rate of the encoder to be lower than that of the decoder;
[0018] Step 7: Finally, generate a summary according to the fine-tuned model.
[0019] In the above solution, the construction of the positive example pairs is enhanced from three aspects: character repetition, punctuation replacement, and sentence modification:
[0020] Step 1a: Character repetition: Randomly select some characters in the original text and repeat them once after these characters to increase the text length and introduce noise, and the repetition ratio is controlled at 0%-25%;
[0021] Step 1b: Punctuation replacement: Replace the original punctuation marks in the document with a preset probability. Specifically:
[0022] Comma: Replace all English commas with Chinese commas, and then replace the Chinese commas with other symbols with a probability of 50%;
[0023] Colon: Replace all English colons with Chinese colons, and then replace the Chinese colons with other symbols with a probability of 50%;
[0024] Full stop: Replace all English full stops with Chinese full stops, and then replace the Chinese full stops with question marks with a probability of 50%;
[0025] Step 1c: Sentence modification: Perform operations of swapping, inserting, and deleting sentences on the original text with a preset probability to generate enhanced text. Specifically:
[0026] Swap: Randomly select two sentence indices for swapping, and the probability of triggering this event is 70%;
[0027] Insert: Randomly select a sentence from the original text and insert it at any position in the text. The probability of triggering this event is 90%;
[0028] Delete: Randomly select at most "the current number of text sentences - 1" sentences for deletion. The probability of triggering this event is 35%.
[0029] In the above solution, the processing processes of the online encoder and the target encoder in step 2 include:
[0030] Input the enhanced positive example pairs x1 and x2 into the two encoders respectively, and calculate the text representation through the following formula:
[0031] h1 = f θ (x1)
[0032] h2 = f ξ (x2)
[0033] where f θ and f ξ represent the mapping functions of the online encoder and the target encoder respectively, θ and ξ are the corresponding parameter sets, and satisfy the initial condition θ (0) = ξ (0) .
[0034] In the above solution, the multi-layer perceptron mapping process in step 3 satisfies:
[0035] g1 = g θ (h1)
[0036] g2 = g ξ (h2)
[0037] g(h i ) = W(2)σ(W (1) h i )
[0038] where h i is the vector representation obtained previously, g θ , g ξ represent the projections of the online side and the target side, that is, two independent MLPs with the same initial parameters. W (1) , W (2) are two learnable linear layers, σ is the ReLU non-linear activation function, and after passing through the first linear layer W (1)Perform batch normalization operation during output.
[0039] In the above solution, the specific process of calculating the contrast loss in step 4 includes:
[0040] Normalize the output vector of the online predictor and the mapping vector of the target end:
[0041]
[0042] Calculate the cosine similarity of the normalized vectors and construct a contrast loss function:
[0043]
[0044] Among them, q θ represents the prediction of the online end, that is, an additional predictor at the online end.
[0045] q1 is the normalized vector after two MLP mappings at the online end, that is, after being mapped by the mapper and the predictor respectively. q2 is the normalized vector obtained after one MLP mapping at the target end. During the comparison process, the similarity between q1 and q2 is maximized. When the model optimizes and minimizes the loss during training, only the online end performs backpropagation, and the target end stops gradient update and updates the parameters through the EMA mechanism.
[0046] In the above solution, the specific implementation of step 5 includes the following steps:
[0047] Step 5a: Perform backpropagation according to the calculated contrast loss and optimize the parameters of the online end using the gradient descent method.
[0048] Step 5b synchronize the parameters of the target end through the exponential moving average mechanism EMA:
[0049] The update of the target end parameters only depends on the historical trajectory of the online end parameters and does not participate in the backpropagation calculation;
[0050] In step 5c, the optimization process of the joint loss function only performs backpropagation at the online end, and the target end keeps the parameters frozen.
[0051] In the above solution, step 6 includes:
[0052] Step 6a: Fine-tune using the enhanced dataset in step 1 and calculate the prediction summary.
[0053] Cross-entropy loss between the predicted summary and the reference summary provided by the original dataset:
[0054]
[0055] Where: is the conditional language model probability, X is the input sequence, i.e., the original text, is the predicted summary sequence, represents the word at the t-th position in the predicted summary sequence, T is the length of the predicted summary sequence, y i represents the word at the i-th position in the reference summary sequence, where one original text corresponds to one reference summary;
[0056] Step 6b: Update the parameters according to the calculated cross-entropy loss, and use the gradient descent method to optimize the encoder and decoder parameters:
[0057] Step 6c: Since the encoder has learned a good encoded representation in the BYOL framework, during the fine-tuning process, the change of the encoder parameters should be controlled, and the encoder learning rate is set lower than the decoder learning rate.
[0058] The present invention also provides a legal text automated summary device based on generative artificial intelligence, including the following modules:
[0059] Multi-level data augmentation module, which is used to modify the original Chinese legal text, implement multi-level data augmentation, generate augmented text, and form positive example pairs;
[0060] Encoder structure module, including the online encoder and the target encoder of the BYOL framework that receive positive example pairs. The online encoder and the target encoder adopt the encoder structure of the pre-trained mBART model and have the same initial parameters;
[0061] Dimension mapping module, which respectively performs dimension mapping on the encoding results of the online end and the target end through a multi-layer perceptron to obtain mapped vectors;
[0062] Predictor setting module, which is used to set a predictor at the online end. After normalizing the mapped vector, calculate the cosine similarity between the online end predicted vector and the target end mapped vector, and calculate the contrast loss function;
[0063] Parameter update module, which is used to perform gradient update on the online end parameters, and update the target end parameters through the EMA mechanism. At the same time, only update the encoder of the mBART model, and freeze the decoder parameters;
[0064] The model saving and fine-tuning module is used to save the trained mBART model encoder on the online side, extract the mBART model, fine-tune it on legal texts, and set the learning rate of the encoder to be lower than that of the decoder;
[0065] The abstract generation module finally generates an abstract according to the fine-tuned model.
[0066] In the above solution, in the multi-level data augmentation module, the construction of the positive example pairs is enhanced from three aspects: character repetition, punctuation replacement, and sentence modification:
[0067] Step 1a: Character repetition: Randomly select some characters in the original text and repeat them once after these characters to increase the text length and introduce noise, with the repetition ratio controlled at 0%-25%;
[0068] Step 1b: Punctuation replacement: Replace the original punctuation marks in the document with a preset probability. Specifically:
[0069] Comma: Replace all English commas with Chinese commas, and then replace the Chinese commas with other symbols with a probability of 50%;
[0070] Colon: Replace all English colons with Chinese colons, and then replace the Chinese colons with other symbols with a probability of 50%;
[0071] Period: Replace all English periods with Chinese periods, and then replace the Chinese periods with question marks with a probability of 50%;
[0072] Step 1c: Sentence modification: Perform swapping, insertion, and deletion operations on the sentences of the original text with a preset probability to generate enhanced text. Specifically:
[0073] Swapping: Randomly select two sentence indices for swapping, and the probability of triggering this event is 70%;
[0074] Insertion: Randomly select a sentence from the original text and insert it at any position in the text, and the probability of triggering this event is 90%;
[0075] Deletion: Randomly select at most "the current number of text sentences - 1" sentences for deletion, and the probability of triggering this event is 35%.
[0076] The present invention also provides a storage medium. When a processor executes the program in the storage medium, it implements the described method for automatically generating abstracts of legal texts based on generative artificial intelligence.
[0077] Because the present invention adopts the above technical means, it has the following beneficial effects:
[0078] 1. Collaborative Design of Multi-form Positive Example Pairs and Contrastive Learning
[0079] Technical means: Adopt data augmentation strategies such as random repetition of words, punctuation replacement, and equal-probability modification of sentences to generate positive example text pairs with consistent semantics but diverse forms, and combine the negative-sample-free contrastive learning mechanism of the BYOL framework.
[0080] Problem solved: Alleviate the problem of insufficient model generalization ability caused by the strong professionalism of legal texts and scarce labeled data, and overcome the defect of traditional contrastive learning relying on negative sample selection bias.
[0081] Effect: Enhance the robustness of the model to perturbations such as legal term replacement and word order adjustment, improve the consistency of semantic representation, and enable the generated summary to maintain high fidelity to the key information in the original text.
[0082] 2. Dual Encoder Momentum Update
[0083] Technical means: Construct an online-target dual encoder architecture and synchronize parameters through the EMA mechanism; design a contrastive loss (LCL) to backpropagate the gradient for the online side and train an mBART model encoder that learns good representations.
[0084] Problem solved: Eliminate the risk of representation collapse in single encoder self-supervised training, make the feature representation of the target network more stable, and reduce the fluctuations during the training process.
[0085] Effect: Achieve the stability of model parameter update, enabling the model to understand the key semantics and key elements in judicial texts.
[0086] 3. Group Parameter Optimization
[0087] Technical means: Retain the encoder after contrastive learning, set the learning rate of the encoder lower than that of the decoder, and perform fine-tuning of the encoder and decoder summary generation on the enhanced text.
[0088] Problem solved: Retain the good representation of judicial texts and balance the conflict between the dual objectives of semantic similarity learning and generation quality optimization.
[0089] Effect: Ensure that the model can bidirectionally understand the original text during the encoding stage, and the summary generated during the decoding stage not only fits the core semantics of the original text but also has natural language fluency.
[0090] 4. Hierarchical Feature Extraction for Legal Text Adaptation
[0091] Technical means: Adopt the hierarchical attention mechanism of the mBART encoder, and dynamically capture the multi-level semantic relationships of paragraphs - sentences - keywords in combination with the structural characteristics of legal texts; compress the vector dimension through the MLP mapping layer to retain key information such as criterion elements and legal article references.
[0092] Problem to be solved: Overcome the deficiencies of traditional models in modeling long-distance dependencies and professional term correlations in legal texts.
[0093] Effect: In the legal public opinion dataset, the extraction accuracy of legal elements such as "dispute focus" and "judgment result" is improved by 9.2%, significantly better than its baseline pre-trained model.
[0094] 5. Lightweight Deployment and Real-time Generation Capability
[0095] Technical means: Compress the model parameter scale through contrastive learning pre-training, and adopt dynamic gradient clipping and hierarchical learning rate strategies to optimize the training efficiency; retain the single encoder-decoder architecture for inference.
[0096] Problem to be solved: Break through the application bottlenecks of large language models (such as ChatGPT) in the judicial scenario, such as high response latency and strong hardware dependence.
[0097] Effect: It only takes 1.8 seconds to generate a single judgment summary (40% faster than ChatGPT's 3 seconds), supports real-time processing of texts with thousands of words, and meets strict time sequence requirements such as instant summaries of court transcripts.
[0098] 6. Evaluation System for Domain Knowledge Fusion
[0099] Technical means: Build a legal element annotation system based on the CAIL2022 dataset, and design a multi-dimensional evaluation scheme that combines ROUGE indicators and manual evaluation of professional compliance.
[0100] Problem to be solved: Solve the problem that general summary evaluation indicators (such as ROUGE) are insufficient in measuring legal professionalism and logical rigor.
[0101] Effect: The ROUGE-2 score of 0.3893 is achieved on the test set (49.8% improvement compared to ChatGPT), significantly better than the pure data-driven baseline method. Brief Description of the Drawings
[0102] Figure 1 It is a flow diagram of the present invention;
[0103] Figure 2 It is a simple diagram of the contrastive learning training framework;
[0104] Figure 3 It is a schematic diagram for constructing positive example pairs.
[0105] Figure 4 It is a schematic diagram of the model fine-tuning process. Specific implementation manners
[0106] The embodiments of the present invention will be described in detail below. Although the present invention will be described and illustrated in conjunction with some specific implementation manners, it should be noted that the present invention is not limited to these implementation manners only. On the contrary, any modifications or equivalent replacements made to the present invention should be covered within the scope of the claims of the present invention.
[0107] In addition, in order to better illustrate the present invention, numerous specific details are given in the following specific implementation manners. Those skilled in the art will understand that the present invention can be implemented without these specific details.
[0108] The concept of the present invention is proposed based on actual production requirements. To facilitate the understanding of the technical concept of the present invention by those skilled in the art, the following further explanations are made:
[0109] The purpose of the present invention is to provide an innovative method for improving the generation of Chinese legal text summaries based on the contrastive learning method, which can compress a long legal text into a more concise format and retain the basic information and meaning of the original text. The present invention is based on the BYOL contrastive learning idea and improves the traditional generative text summarization framework, and is trained and applied on Chinese legal texts.
[0110] The present invention uses the generative mBART model as the baseline model, introduces the contrastive learning method to create a generative summary generation framework, and performs contrastive learning by performing data augmentation, constructing positive example pairs, and maximizing the similarity between positive example pairs, and fine-tunes the pre-trained model on the basis of Chinese legal texts to obtain the final summary generation framework.
[0111] The present invention refers to the BYOL contrastive learning framework without negative samples for self-supervised learning, and uses it on the generative BART model to improve the method for generating legal text summaries. The present invention uses self-supervised contrastive learning to maximize the similarity between the two, and encourages the model to identify whether the two context vectors learned from the encoder represent the same original input text.
[0112] 1. Construct positive example pairs
[0113] The present invention performs data augmentation on the original text from three levels of words, punctuation, and sentences to obtain similar text pairs, which are regarded as positive sample pairs. Among them, the word-level data augmentation method uses synonym replacement and word repetition strategies, and the sentence-level data augmentation method uses equiprobable exchange, insertion, and deletion. The training and test data sets adopt the legal public opinion data set of CAIL2022 of the Legal Research Cup.
[0114] Randomly copy some words or sub-words in the judicial text at the word and phrase level. For a given sentence, after being processed by a tokenizer, a sequence of words is obtained, and the words to be repeated are randomly selected to make them appear repeatedly in the sequence.
[0115] At the punctuation level, for the three symbols of comma, colon, and period, first standardize all punctuation to correct punctuation, that is, modify all possible English commas, colons, and periods to Chinese punctuation, and then replace the correct punctuation with other symbols such as question marks with a certain probability.
[0116] At the sentence level, three document enhancement methods are used.
[0117] (1) Swap: Select two sentences in the document with a certain probability and swap their position orders.
[0118] (2) Insert: Randomly select a sentence from the document with a certain probability and insert it into any position in the document.
[0119] (3) Delete: Randomly select at most (the current number of text sentences - 1) sentences for deletion with a certain probability. The sentence-level enhancement method is used twice for a single piece of text.
[0120] 2. Encoder
[0121] The present invention uses two independent Encoders of BART-base-chinese with the same initial parameters to process the input text data and obtain their vector representations. The present invention uses the hidden vector corresponding to the first input token of the last layer of the encoder as the vector representation of the text. If the original text is regarded as a long token sequence, the encoded text representation is
[0122] h1 = f θ (x1)
[0123] h2 = f ξ (x2)
[0124] where X = (x0 = <s>, x1, x2,..., x |X| =) represents the original text, x1 and x2 are positive example pairs, and f θ 、f ξ represent the Encoders of the online side and the target. The Encoder on the online side will update its parameters through the gradients returned later, while the Encoder on the Target side will not be updated.
[0125] 3. Mapper
[0126] In the MLP mapping stage, the present invention uses the same MLP (multi-layer perception) as BYOL to map the previously obtained vector representation into a lower-dimensional representation.
[0127] g1 = g θ (h1)
[0128] g2 = g ξ (h2)
[0129] g(h i ) = W (2) σ(W (1) h i )
[0130] Where h i is the vector representation obtained previously, and g θ , g ξ represent the projections of the online and target sides, that is, two independent MLPs (multi-layer perception) with the same initial parameters. W (1) ,
[0131] W (2) are two learnable linear layers, σ is the ReLU non-linear activation function, and batch normalization operation is performed when outputting after passing through the first linear layer W (1) .
[0132] 4. Predictor
[0133] The predictor uses the same structure as the mapper, but only exists on the online side, and normalizes the vector after MLP mapping:
[0134]
[0135] Then calculate the cosine similarity, and use 1 - cosine similarity as the contrastive learning loss:
[0136]
[0137] Where q1 is the vector after two MLP mappings on the online side, that is, after passing through the mapper and the predictor respectively, and q2 is the vector obtained after one MLP mapping on the target side. During the comparison process, the similarity between q1 and q2 is maximized. When optimizing and minimizing the loss during the training of the model, only backpropagation is performed on the online side, and the gradient update of the target side stops and is updated by the exponential moving average mechanism. Finally, the mBART model encoder with a better legal text representation is obtained.
[0138] 5. Enhanced text fine-tuning
[0139] After obtaining the trained encoder, fine-tuning is performed on the enhanced legal text. The vector representation h1 generated by the encoder is passed to the decoder of mBART to generate a predicted summary, which is compared with the target summary to calculate the fine-tuning loss. The present invention uses cross entropy as the model fine-tuning loss. Specifically, let Represents the summary and calculates the cross entropy between the summary generated by the model and the reference summary:
[0140]
[0141] Where P is the conditional language model probability, X is the original text sequence, represents the word at the tth position in the sequence, T is the length of the generated summary sequence, y i Represents the word at the i-th position in the true summary sequence. Parameters are updated based on the calculated cross entropy loss, and the encoder and decoder parameters are optimized using the gradient descent method:
[0142] Since the encoder has already learned a good encoding representation in the BYOL framework, the encoder parameter changes should be controlled during fine-tuning, and the encoder learning rate should be set lower than the decoder learning rate. In the end, not only will the BART encoder learn a good representation of the legal text, but the BART decoder will also be able to generate a summary that retains the original information and semantics.
[0143] Example 1
[0144] The present invention provides a method for automatic summarization of legal texts based on generative artificial intelligence, comprising the following steps:
[0145] Step 1: Perform multi-level data enhancement on the original Chinese legal text to generate enhanced text to form positive example pairs;
[0146] Step 1a: Character repetition: randomly select some characters in the original text and repeat them once after them, thereby increasing the text length and introducing noise. The repetition ratio is controlled between 0% and 25%;
[0147] For example: Enter the original legal text: "The plaintiff claims that the defendant should compensate for economic losses of 500,000 yuan."
[0148] Randomly select characters (such as "被", "补") and repeat them to generate enhanced text: "The plaintiff claims that the defendant should compensate for economic losses of 500,000 yuan."
[0149] Step 1b: Punctuation replacement: Replace the original punctuation marks in the document with a preset probability. Specifically:
[0150] Original punctuation specification: uniformly replace English punctuation with Chinese punctuation. For example, replace "," with ",", and then with a 50% probability, replace the Chinese comma and colon with other symbols. With a 50% probability, replace the Chinese full stop with a question mark.
[0151] For example:
[0152] Original sentence: "The court holds that the evidence is insufficient and dismisses the appeal."
[0153] Enhanced: "The court holds that, the evidence is insufficient and dismisses the appeal?"
[0154] Step 1c: Sentence modification: Perform swapping, insertion, and deletion operations on the sentences of the original text with a preset probability to generate enhanced text. Specifically:
[0155] Swapping: Randomly select two sentence indices for swapping, with a 70% probability of triggering this event;
[0156] For example:
[0157] Original paragraph:
[0158] [1] "The defendant failed to fulfill the contractual obligations."
[0159] [2] "The plaintiff requests the rescission of the contract."
[0160] Enhanced:
[0161] [2] "The plaintiff requests the rescission of the contract."
[0162] [1] "The defendant failed to fulfill the contractual obligations."
[0163] Insertion: Randomly select a sentence from the original text and insert it at any position in the text, with a 90% probability of triggering this event;
[0164] For example:
[0165] Original: "The judgment is as follows: First, the defendant shall pay liquidated damages of 100,000 yuan."
[0166] After insertion: "The defendant shall perform within 30 days. The judgment is as follows: First, the defendant shall pay liquidated damages of 100,000 yuan"
[0167] Deletion: Randomly select at most "the current number of text sentences - 1" sentences for deletion, with a 35% probability of triggering this event.
[0168] For example, if the original text contains 5 sentences, 2 core sentences are retained after deletion.
[0169] Step 2: Input the positive example pairs into the online encoder and the target encoder of the BYOL framework respectively. The online encoder and the target encoder adopt the encoder structure of the pre-trained mBART model and have the same initial parameters;
[0170] For the convenience of those skilled in the art to better understand the technical concept of the present invention, further, the processing processes of the online encoder and the target encoder in Step 2 include:
[0171] Input the enhanced positive example pairs x1 and x2 into the two encoders respectively, and calculate the text representation through the following formula:
[0172] h1 = f θ (x1)
[0173] h2 = f ξ (x2)
[0174] where f θ and f ξ represent the mapping functions of the online encoder and the target encoder respectively, θ and ξ are the corresponding parameter sets, and satisfy the initial condition θ (0) = ξ (0) .
[0175] Step 3: Perform dimensionality mapping on the encoding results of the online encoder and the target encoder respectively through a multi-layer perceptron to obtain the mapping vectors;
[0176] For the convenience of those skilled in the art to better understand the technical concept of the present invention, further, the multi-layer perceptron mapping process in Step 3 satisfies:
[0177] g1 = g θ (h1)
[0178] g2 = g ξ (h2)
[0179] g(h i ) = W (2) σ(W (1) h i )
[0180] where h i is the vector representation obtained previously, g θ , g ξ represent the projections of the online encoder and the target encoder, that is, two independent MLPs with the same initial parameters, W (1) , W (2) are two learnable linear layers, σ is the ReLU non-linear activation function, and after passing through the first linear layer W(1) Perform batch normalization operation during output.
[0181] Step 4: Set a predictor at the online end. After normalizing the mapped vector, calculate the cosine similarity between the prediction vector at the online end and the mapped vector at the target end, and calculate the contrast loss function.
[0182] To facilitate better understanding of the technical concept of the present invention by those skilled in the art, the specific process of calculating the contrast loss in step 4 is further described in detail as follows:
[0183] Normalize the output vector of the online predictor and the mapped vector of the target end:
[0184]
[0185] Calculate the cosine similarity of the normalized vectors and construct the contrast loss function:
[0186]
[0187] Where q θ represents the prediction at the online end, that is, an additional predictor at the online end. This predictor introduces a non - linear transformation to map the features at the online end to a new space, thereby breaking the symmetry of the feature space, avoiding feature collapse, and enhancing the distinguishability of features, enabling the features at the online end to better perform contrastive learning with the features at the target end, thus improving the feature extraction ability of the model and the effect of self - supervised learning.
[0188] q1 is the normalized vector after two MLP mappings at the online end, that is, after being mapped by the mapper and the predictor respectively. q2 is the normalized vector obtained after one MLP mapping at the target end. During the comparison process, the similarity between q1 and q2 is maximized. During the training process of the model, when optimizing and minimizing the loss, only backpropagation is performed at the online end, and the gradient update of the target end stops. The parameters are updated by the EMA mechanism.
[0189] Step 5: Update the gradients of the online - end parameters, update the target - end parameters through the EMA mechanism, and only update the encoder of the mBART model while freezing the decoder parameters.
[0190] To facilitate better understanding of the technical concept of the present invention by those skilled in the art, the detailed implementation steps of step 5 are further described as follows:
[0191] Step 5a: Perform backpropagation based on the calculated contrastive loss, and optimize the online - side parameters using the gradient descent method:
[0192]
[0193] Where:
[0194] θ (t) represents the online - side parameters at the t - th iteration;
[0195] η represents the learning rate;
[0196] represents the gradient of the joint loss with respect to the online - side parameters;
[0197] Step 5b: Synchronize the target - side parameters through the exponential moving average mechanism EMA:
[0198] θ' (t) ←τθ' (t-1) +(1 - τ)θ (t)
[0199] Where:
[0200] θ' (t) represents the target - side parameters at the t - th iteration;
[0201] θ (t) represents the online - side parameters at the t - th iteration;
[0202] τ∈[0,1) is the smoothing coefficient of EMA, which controls the proportion of historical parameters retained, enabling the target side to slowly follow the online side for changes;
[0203] The update of the target - side parameters only depends on the historical trajectory of the online - side parameters and does not participate in the backpropagation calculation;
[0204] In Step 5c, the optimization process of the joint loss function only performs backpropagation on the online side, and the target side keeps the parameters frozen.
[0205] Step 6: Save the mBART model encoder of the trained online side, extract the mBART model, fine - tune it on legal texts, and set the encoder learning rate lower than the decoder learning rate.
[0206] To facilitate those skilled in the art to better understand the technical concept of the present invention, the following further describes the detailed implementation steps of Step 6:
[0207] The said Step 6 includes:
[0208] Step 6a: Use the enhanced dataset from Step 1 for fine-tuning and calculate the predicted summary
[0209] The cross-entropy loss between the predicted summary and the reference summary provided by the original dataset:
[0210]
[0211] Where: is the conditional language model probability, X is the input sequence, i.e., the original text, is the predicted summary sequence, represents the word at the t-th position in the predicted summary sequence, T is the length of the predicted summary sequence, y i represents the word at the i-th position in the reference summary sequence, where one original text corresponds to one reference summary;
[0212] Step 6b: Update the parameters according to the calculated cross-entropy loss, and use the gradient descent method to optimize the encoder and decoder parameters:
[0213]
[0214] Where:
[0215] α (t) represents the encoder parameters at the t-th iteration, β (t) represents the decoder parameters at the t-th iteration;
[0216] η α represents the encoder learning rate, η β represents the decoder learning rate;
[0217] represents the gradient of the cross-entropy loss with respect to the encoder parameters, represents the gradient of the cross-entropy loss with respect to the encoder parameters;
[0218] Step 6c: Control the change of encoder parameters during fine-tuning, and set the encoder learning rate lower than the decoder learning rate, i.e.,
[0219] η α = γ * η β
[0220] Where: γ ∈ [0.01, 1) is the multiple of the encoder learning rate compared to the decoder learning rate.
[0221] Step 7: Finally, generate the summary according to the fine-tuned model.
[0222] Example 2
[0223] The present invention also provides a legal text automated summarization device based on generative artificial intelligence, including the following modules:
[0224] A multi-level data augmentation module, which is used to modify the original Chinese legal text, implement multi-level data augmentation, generate augmented text, and form positive example pairs;
[0225] An encoder structure module, including an online encoder and a target encoder of the BYOL framework that receive positive example pairs. The online encoder and the target encoder adopt the encoder structure of the pre-trained mBART model and have the same initial parameters;
[0226] A dimension mapping module, which respectively performs dimension mapping on the encoding results of the online end and the target end through a multi-layer perceptron to obtain mapped vectors;
[0227] A predictor setting module, which is used to set a predictor at the online end. After normalizing the mapped vectors, it calculates the cosine similarity between the online end prediction vector and the target end mapped vector, and calculates the contrast loss function;
[0228] A parameter update module, which is used to perform gradient update on the online end parameters, and the target end parameters are updated through the EMA mechanism. At the same time, only the encoder of the mBART model is updated, and the decoder parameters are frozen;
[0229] A model saving and fine-tuning module, which is used to save the encoder of the trained mBART model at the online end, extract the mBART model, perform fine-tuning on legal texts, and set the encoder learning rate to be lower than the decoder learning rate;
[0230] A summary generation module, which finally generates a summary according to the fine-tuned model.
[0231] To facilitate those skilled in the art to better understand the technical concept of the present invention, the multi-level data augmentation module is further described in detail. Specifically, the construction of the positive example pairs is enhanced from three aspects: character repetition, punctuation replacement, and sentence modification:
[0232] Step 1a: Character repetition: Randomly select some characters in the original text and repeat them once after these characters, so as to increase the text length and introduce noise. The repetition ratio is controlled within 0%-25%;
[0233] Step 1b: Punctuation replacement: Replace the original punctuation marks in the document with a preset probability. Specifically:
[0234] Comma: Replace all English commas with Chinese commas, and then replace the Chinese commas with other symbols with a probability of 50%;
[0235] Colon: Replace all English colons with Chinese colons, and then replace the Chinese colon with other symbols with a probability of 50%;
[0236] Full stop: Replace all English full stops with Chinese full stops, and then replace the Chinese full stop with a question mark with a probability of 50%;
[0237] Step 1c: Sentence modification: Perform swapping, insertion, and deletion operations on the sentences of the original text with a preset probability to generate enhanced text. Specifically:
[0238] Swapping: Randomly select two sentence indices for swapping, and the probability of triggering this event is 70%;
[0239] Insertion: Randomly select a sentence from the original text and insert it at any position in the text, and the probability of triggering this event is 90%;
[0240] Deletion: Randomly select at most "the current number of text sentences - 1" sentences for deletion, and the probability of triggering this event is 35%.
[0241] The implementation of the further encoder structure module includes the following steps:
[0242] The processing processes of the online encoder and the target encoder include:
[0243] Input the enhanced positive example pairs x1 and x2 into the two encoders respectively, and calculate the text representation through the following formula:
[0244] h1 = f θ (x1)
[0245] h2 = f ξ (x2)
[0246] Among them, f θ and f ξ respectively represent the mapping functions of the online encoder and the target encoder, θ and ξ are the corresponding parameter sets, and satisfy the initial condition θ (0) = ξ (0) .
[0247] The implementation process of the further dimension mapping module includes the following steps:
[0248] The multi-layer perceptron mapping process satisfies:
[0249] g1 = g θ (h1)
[0250] g2 = g ξ (h2)
[0251] g(hi ) = W (2) σ(W (1) h i )
[0252] where h i is the vector representation obtained previously, and g θ , g ξ represent the projections of the online and target sides, i.e., two independent MLPs with the same initial parameters. W (1) , W (2) are two learnable linear layers, σ is the ReLU non-linear activation function, and batch normalization is performed during the output of the first linear layer W (1) .
[0253] For the convenience of those skilled in the art to better understand the technical concept of the present invention, the implementation process of the predictor setting module is further described in detail as follows:
[0254] Normalize the output vector of the online predictor and the mapped vector of the target side:
[0255]
[0256]
[0257] Calculate the cosine similarity of the normalized vectors and construct a contrast loss function:
[0258]
[0259] where q θ represents the prediction of the online side, i.e., an additional predictor on the online side,
[0260] q1 is the normalized vector after two MLP mappings on the online side, i.e., after being mapped by the mapper and the predictor respectively, and q2 is the normalized vector obtained after one MLP mapping on the target side. During the comparison process, the similarity between q1 and q2 is maximized. During the training process of the model, only the online side performs backpropagation when optimizing the minimization of the loss, and the target side stops gradient update and updates the parameters through the EMA mechanism.
[0261] For the convenience of those skilled in the art to better understand the technical concept of the present invention, the detailed steps of the implementation of the parameter update module are described as follows:
[0262] Step 5a: Perform backpropagation according to the calculated contrast loss and optimize the parameters of the online side using the gradient descent method:
[0263]
[0264] Wherein:
[0265] θ (t) represents the online - end parameter at the t - th iteration;
[0266] η represents the learning rate;
[0267] represents the gradient of the joint loss with respect to the online - end parameter;
[0268] Step 5b synchronizes the target - end parameter through the Exponential Moving Average (EMA) mechanism:
[0269] θ' (t) ← τθ' (t-1) +(1 - τ)θ (t)
[0270] Wherein:
[0271] θ' (t) represents the target - end parameter at the t - th iteration;
[0272] θ (t) represents the online - end parameter at the t - th iteration;
[0273] τ ∈ [0, 1) is the smoothing coefficient of EMA, which controls the proportion of historical parameters retained, enabling the target end to slowly follow the online end for changes;
[0274] The update of the target - end parameter only depends on the historical trajectory of the online - end parameter and does not participate in the backpropagation calculation;
[0275] The optimization process of the joint loss function described in Step 5c only performs backpropagation on the online end, and the target end keeps the parameter frozen.
[0276] To facilitate those skilled in the art to better understand the technical concept of the present invention, the detailed implementation steps of the model saving and fine - tuning module are further described as follows:
[0277] Step 6a performs fine - tuning using the enhanced dataset in Step 1 and calculates the predicted summary
[0278] The cross - entropy loss between the predicted summary and the reference summary provided by the original dataset:
[0279]
[0280] Wherein: is the conditional language model probability, X is the input sequence, i.e., the original text, is the predicted summary sequence, represents the word at the t-th position in the predicted summary sequence, T is the length of the predicted summary sequence, y i represents the word at the i-th position in the reference summary sequence, where one original text corresponds to one reference summary;
[0281] Step 6b: Update the parameters according to the calculated cross-entropy loss, and optimize the encoder and decoder parameters using the gradient descent method:
[0282]
[0283] where:
[0284] α (t) represents the encoder parameters at the t-th iteration, β (t) represents the decoder parameters at the t-th iteration;
[0285] η α represents the encoder learning rate, η β represents the decoder learning rate;
[0286] represents the gradient of the cross-entropy loss with respect to the encoder parameters, represents the gradient of the cross-entropy loss with respect to the encoder parameters;
[0287] Step 6c: Control the change of encoder parameters during the fine-tuning process, and set the encoder learning rate lower than the decoder learning rate, i.e.,
[0288] η α = γ * η β
[0289] where: γ ∈ [0.01, 1) is the multiple of the encoder learning rate compared to the decoder learning rate.
[0290] Example 3
[0291] The present invention also provides a storage medium. When a processor executes a program in the storage medium, it implements a method for automatically summarizing legal texts based on generative artificial intelligence as described in Example 1.
[0292] To facilitate better understanding of the inventive concept by those skilled in the art, the advantages of the present invention over the prior art under the inventive concept are further described:
[0293] 1. Good representation of legal texts
[0294] Through the combination of self-supervised contrastive learning and generative models, the present invention achieves a good representation of legal texts. The core lies in the synergistic effect of contrastive learning. Diversified positive sample pairs are generated at the character, word, punctuation, and sentence levels, enabling the model to learn multi-level semantic features of legal texts. Self-supervised contrastive learning simplifies the training process and reduces bias by maximizing the similarity of positive sample pairs and avoiding the use of negative samples. The BART encoder extracts deep semantic representations of the text, and together with the collaborative optimization mechanism of the online and target ends, ensures the stability and consistency of the representation. The online end optimizes the parameters through gradient updates, while the target end provides a stable representation target through exponential moving average. Differentiated learning enables the encoder to remain effective when training a well-generated decoder, enabling the model to strike a balance between semantic similarity and retention of key judicial information. The combination of these methods enables the model to extract high-quality, robust, and strongly generalizable semantic representations from legal texts, laying the foundation for subsequent tasks.
[0295] 2. Generation of High-quality Chinese Legal Text Summaries
[0296] The present invention is capable of generating high-quality legal text summaries. Thanks to the generative summary model framework constructed by introducing contrastive learning methods, the BART model based on the BERT model + GPT model mechanism is more suitable for handling the generation task of legal texts. The encoder with the BERT as the main structure can provide strong context understanding ability, while the decoder with the GPT as the main structure can use this understanding to generate coherent text outputs. The mBART encoder extracts high-quality semantic representations from the enhanced text through self-supervised contrastive learning. These representations not only capture the global context information of the text but also retain key legal semantic details. The decoder then uses these representations to generate coherent and accurate summaries. In addition, the characteristics of the CAIL (Legal Research Cup) 2022 legal-related public opinion dataset and the targeted optimization of model design are also key factors in generating high-quality summaries. The CAIL2022 dataset contains rich legal-related public opinion texts, covering diverse legal scenarios and language expressions, providing high-quality training samples for the model. These texts have a clear logical structure and legal professionalism, enabling the model to learn specific language patterns and semantic features of legal texts.
[0297] The present invention uses ROUGE (Recall-Oriented Understudy for Gisting Evaluation) scores to evaluate the effectiveness of the generated legal texts. ROUGE scores are a commonly used set of metrics for evaluating automatic text summarization. It measures the quality of the automatically generated summary by comparing the overlap between the generated summary and the reference summary. The present invention mainly uses Rouge-N (Rouge-1, Rouge-2) and Rouge-L to evaluate the summarization method. The higher the ROUGE score, the higher the overlap between the automatically generated summary and the standard summary, and thus the better the quality of the generated result.
[0298] Fine-tuning training is performed on the data-augmented dataset, and all models are trained for 10 epochs. The trained models are evaluated on the test set to obtain the evaluation results of the following extractive and generative summarization methods. The recall rate (R) in the ROUGE score is selected here.
[0299] Table 1 Evaluation Results of the Generation Method
[0300] Model ROUGE-1 ROUGE-2 ROUGE-L Time ChatGPT 0.4284 0.2957 0.3842 3s BART 0.3828 0.1794 0.3265 1.5s Gretel 0.4157 0.2841 0.3741 2.1s Con_sum 0.4315 0.2893 0.3986 1.8s
[0301] Among them, Con_sum is the generative automatic summarization method of the present invention, and BART and Gretel are pre-trained language models for generation and extraction. The summarization effect of the present invention on Chinese legal texts has achieved a higher level in terms of ROUGE scores compared with the currently common text summarization methods, and it is faster in generation speed than the currently common large language models. It has a relatively high level and production value in practical applications.
[0302] 3. Rule-Based Legal Text Summarization
[0304] Rule-based legal text summarization is a method of extracting key information from legal texts and generating summaries through predefined rules or templates. For example, in the judgment summary task, rules can be designed to extract fixed fields such as "plaintiff", "defendant", "cause of action", and "judgment result". The system may locate relevant paragraphs based on specific keywords (such as "judgment is as follows" or "the court holds that") and extract the core content of the judgment from them. This method relies on manually designed rules and usually requires an in-depth understanding of the structure and language patterns of legal texts.
[0305] However, rule-based legal text summarization has obvious defects. The generalization ability of rules is limited and it is difficult to adapt to diverse legal text expression methods. Moreover, rule-based methods are difficult to handle complex semantic information. Especially when the text involves multi-level logical reasoning or implicit legal relationships, rules often fail to capture these details. In addition, maintaining and updating the rule base requires a large amount of labor costs. Especially when legal provisions or language habits change, the rules need to be frequently adjusted to meet new requirements. These limitations make it difficult for rule-based methods to generate high-quality and flexible legal text summaries in practical applications.
[0306] 4. As a possible extension method, the extractive Gretel model replaces the generative model (Extraction is another idea for text summarization. In this framework, the generative model can be replaced with the extractive Gretel model. The framework remains unchanged except for the model, and it can still be used)
[0307] Different from the generative method, the extractive method does not create new text sequences, but directly selects important sentences or phrases from the original text as the summary content. The GRETEL model transforms the extraction task into a binary classification problem. The label of the sentence extracted into the summary is 1, and the label of the sentence not extracted into the summary is 0. This model combines a neural topic model and a pre-trained topic language model to extract local semantics and global semantics simultaneously to generate a summary. This model is divided into the following three structures: a hierarchical encoder, a graph contrastive topic model, and a probabilistic decoder. The Gretel model with encoder and decoder structures can also be applied to the contrastive learning improvement framework of the present invention.
[0308] However, the extractive method is rarely used at present, mainly because legal texts usually contain complex logical relationships and professional terms, and simply extracting sentences may not accurately convey the core semantics.
[0309] 5. Application of large language models
[0310] Large language models such as DeepSeek or ChatGPT can be directly used in legal text summarization methods. These models are pre-trained with a large amount of text data and have the ability to understand and generate natural language. Large language models have been exposed to various types of texts, including legal documents, during the pre-training stage, so they can understand these terms and structures. Through fine-tuning, the model can better capture the key information in legal texts and generate summaries that meet the requirements of legal professionals.
[0311] Although large language models are powerful in generating legal text summaries, they have some disadvantages compared to traditional models such as BART. For example, large models require more computing resources, are prone to overfitting, have high demands for training data, are slow in generating speed, and have poor interpretability and controllability. Traditional models like BART may perform more stably, be more adaptable to specific tasks, and generate faster under limited resources and small data volumes.
[0312] 6. Application in Chinese Legal Texts
[0313] Most existing automatic summarization methods are targeted at news articles and scientific articles, and there are significant differences between legal language and news language. In other texts, especially news, social media, literary works, etc., the language style is usually more expressive, emotional or colloquial, and the way of information transmission is more intuitive and concise. Legal provisions and judgments often adopt complex sentence patterns and precise definitions to ensure that there is no ambiguity in the interpretation and application of legal content. This requirement for precise and formal language makes legal texts often more complex than ordinary news reports, literary works or technical documents, and cannot be simplified or adjusted arbitrarily. Legal texts are full of professional terms and specialized legal vocabulary, and these terms can only be accurately understood in a legal context. For example, terms such as "appeal", "judgment", "evidence exclusion" have specific legal meanings and require people with legal knowledge to correctly interpret and apply them.
[0314] The automatic legal text summarization method of the present invention utilizes the CAIL2022 legal text summary dataset, specifically targeting the complexity and professionalism of legal language, and provides a more accurate and reliable way of generating summaries. To a certain extent, it fills the gap in Chinese judicial intelligence, not only improving the accuracy of legal text summaries, but also providing a more practical tool for legal professionals.
[0315] 7. Introduction of Contrastive Learning Method
[0316] The present invention innovatively introduces the BYOL framework into the legal text summarization task, avoiding the dependence on negative samples in traditional contrastive learning. By only using positive sample pairs (enhanced similar texts) for contrastive learning, the training process is simplified, and at the same time, the bias that may be brought by negative sample selection is reduced. This design enables the model to more efficiently learn the semantic representation of legal texts and lays a foundation for generating high-quality summaries.
[0317] 8. Multi-granularity Data Augmentation and Adaptation to Legal Texts
[0318] The present invention designs data augmentation strategies for legal texts at the word, punctuation, and sentence levels, including methods such as word repetition, punctuation replacement, equiprobable exchange, insertion, and deletion of sentences. These augmentation methods are particularly adapted to the characteristics of legal texts (such as professional terms and rigorous structures), generating diverse positive sample pairs and enhancing the model's ability to understand the semantics of legal texts and its generalization performance. This multi-granularity augmentation strategy is an innovative application in the field of legal text processing.
Claims
1. A generative method for automatic summarization of legal texts, characterized in that: The following steps are involved: Step 1: Perform multi-level data enhancement on the original Chinese legal text to generate enhanced text to form positive example pairs; Step 2: Input the positive example pairs into the BYOL framework online encoder and target encoder respectively. The online encoder and target encoder use the encoder structure of the pre-trained mBART model and have the same initial parameters. Step 3: Use a multi-layer perceptron to perform dimension mapping on the encoding results of the online end and the target end respectively to obtain a mapping vector; Step 4: Set up the predictor on the online side, normalize the mapping vector, calculate the cosine similarity between the online prediction vector and the target mapping vector, and calculate the contrast loss function; Step 5: Perform gradient updates on the online side parameters, and update the target side parameters through the EMA mechanism. Meanwhile, only the encoder of the mBART model is updated, and the decoder parameters are frozen. Step 6: Save the trained mBART model encoder on the online side, extract the mBART model, fine-tune it on the legal text, and set the encoder learning rate to be lower than the decoder learning rate; Step 7: Finally, generate a summary based on the fine-tuned model.
2. The method according to claim 1, characterized in that The positive examples are constructed in three aspects: character repetition, punctuation replacement, and sentence modification: Step 1a: Character repetition: randomly select some characters in the original text and repeat them once after them, thereby increasing the text length and introducing noise. The repetition ratio is controlled between 0% and 25%; Step 1b: Punctuation replacement: Replace the original punctuation marks in the document with a preset probability. Specifically: Comma: Replace all English commas with Chinese commas, and then replace Chinese commas with other symbols with a 50% probability; Colon: replace all English colons with Chinese colons, and then replace Chinese colons with other symbols with a probability of 50%; Period: Replace all English periods with Chinese periods, and then replace Chinese periods with question marks with a 50% probability; Step 1c: Sentence modification: Exchange, insert, and delete sentences in the original text with a preset probability to generate enhanced text. Specifically: Exchange: Randomly select two sentence indexes to exchange, with a probability of 70% to trigger this event; Insert: Randomly select a sentence from the original text and insert it into any position in the text, with a probability of 90% to trigger the event; Delete: Randomly select sentences up to "number of sentences in the current text - 1" for deletion. The probability of triggering this event is 35%.
3. The method according to claim 1, characterized in that The processing of the online end encoder and the target end encoder in step 2 includes: The enhanced positive example pairs x1 and x2 are input into two encoders respectively, and the text representation is calculated by the following formula: h1=f θ (x1) h2=f ξ (x2) Among them, f θ and f ξ They represent the mapping functions of the online and target encoders respectively, θ and ξ are the corresponding parameter sets, and satisfy the initial condition θ (0) =ξ (0) .
4. The method according to claim 1, characterized in that: The multi-layer perceptron mapping process of step 3 satisfies: g1=g θ (h1) g2=g ξ (h2) g(h i )=W (2) σ(W (1) h i ) where h i is the vector representation obtained previously, g θ , g ξ represents the projection of the online side and the target side, that is, two independent MLPs with the same initial parameters, W (1) , W (2) are two learnable linear layers, σ is the ReLU nonlinear activation function, after the first linear layer W (1) Perform batch normalization on output.
5. The method according to claim 1, characterized in that The contrast loss calculation process of step 4 specifically includes: Normalize the online predictor output vector and the target mapping vector: Calculate the cosine similarity of the normalized vectors and construct the contrast loss function: where q θ Represents the prediction on the online side, that is, an additional predictor on the online side. q1 is the normalized vector after two MLP mappings on the online side, that is, the mappings of the mapper and the predictor respectively. q2 is the normalized vector obtained after one MLP mapping on the target side. The similarity between q1 and q2 is maximized in the comparison process. When the model optimizes and minimizes the loss during training, backpropagation is only performed on the online side. The target side stops gradient updating and uses the EMA mechanism to update parameters.
6. The method according to claim 1, characterized in that The specific implementation of step 5 includes the following steps: Step 5a: Perform back propagation based on the calculated contrast loss and use the gradient descent method to optimize the online side parameters: Step 5b synchronizes the target parameters through the exponential moving average mechanism EMA: The target-side parameter update only depends on the historical trajectory of the online-side parameters and does not participate in the back-propagation calculation; The optimization process of the joint loss function described in step 5c performs back propagation only on the online side, and the parameters on the target side remain frozen.
7. The method according to claim 1, characterized in that The step 6 comprises: Step 6a: Fine-tune using the augmented dataset from step 1 and calculate the prediction summary The cross entropy loss between the reference summary given with the original dataset: in: is the conditional language model probability, X is the input sequence, i.e. the original text, To predict the summary sequence, represents the word at the tth position in the predicted summary sequence, T is the length of the predicted summary sequence, and y i represents the word at the ith position in the reference summary sequence, where one original text corresponds to one reference summary; Step 6b: Update the parameters based on the calculated cross entropy loss and use gradient descent to optimize the encoder and decoder parameters: Step 6c: Control the encoder parameter changes during fine-tuning and set the encoder learning rate lower than the decoder learning rate.
8. A generative legal text automatic summarization device, characterized in that: Includes the following modules: The multi-level data enhancement module is used to modify the original Chinese legal text, implement multi-level data enhancement, generate enhanced text, and form positive example pairs; An encoder structure module, including an online encoder and a target encoder of a BYOL framework for receiving positive example pairs, wherein the online encoder and the target encoder adopt an encoder structure of a pre-trained mBART model and have the same initial parameters; The dimension mapping module uses a multi-layer perceptron to perform dimension mapping on the encoding results of the online end and the target end to obtain a mapping vector; The predictor setting module is used to set the predictor on the online side, normalize the mapping vector, calculate the cosine similarity between the online prediction vector and the target mapping vector, and calculate the contrast loss function; The parameter update module is used to perform gradient updates on the online side parameters. The target side parameters are updated through the EMA mechanism. At the same time, only the encoder of the mBART model is updated, and the decoder parameters are frozen. The model saving and fine-tuning module is used to save the trained online mBART model encoder, extract the mBART model, fine-tune it on the legal text, and set the encoder learning rate to be lower than the decoder learning rate; The summary generation module finally generates a summary based on the fine-tuned model.
9. The device according to claim 1, characterized in that In the multi-level data enhancement module, the positive examples are enhanced from three aspects: character repetition, punctuation replacement, and sentence modification: Step 1a: Character repetition: randomly select some characters in the original text and repeat them once after them, thereby increasing the text length and introducing noise. The repetition ratio is controlled between 0% and 25%; Step 1b: Punctuation replacement: Replace the original punctuation marks in the document with a preset probability. Specifically: Comma: Replace all English commas with Chinese commas, and then replace Chinese commas with other symbols with a 50% probability; Colon: replace all English colons with Chinese colons, and then replace Chinese colons with other symbols with a probability of 50%; Period: Replace all English periods with Chinese periods, and then replace Chinese periods with question marks with a 50% probability; Step 1c: Sentence modification: Exchange, insert, and delete sentences in the original text with a preset probability to generate enhanced text. Specifically: Exchange: Randomly select two sentence indexes to exchange, with a probability of 70% to trigger this event; Insert: Randomly select a sentence from the original text and insert it into any position in the text, with a probability of 90% to trigger the event; Delete: Randomly select sentences up to "number of sentences in the current text - 1" for deletion. The probability of triggering this event is 35%.
10. A storage medium, characterized in that: When the processor executes the program in the storage medium, a generative legal text automatic summarization method as described in any of claims 1-7 is implemented.