Legal Case Element Extraction Method Based on Re-Attention Mechanism and Contrastive Loss
By introducing a reattention mechanism and comparison loss in the extraction of legal case elements, the problems of similar label identification and exposure deviation of Seq2Seq model are solved, and more efficient label extraction performance of legal case elements is achieved.
Patent Information
- Application Number
- CN202310537953.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-15
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-05-15
AI Technical Summary
It is difficult to effectively identify similar legal case elements labels, and the Seq2Seq model has problems with exposure deviations, making it difficult to distinguish legal case elements labels.
The legal case element extraction method based on the reattention mechanism and the comparison loss is adopted, and the legal case element label representation is optimized through joint embedding and comparison loss, and the redundant information is removed and similar labels are distinguished by combining the reattention mechanism.
It effectively improves the extraction performance of legal case element labels, especially in the identification of low-frequency labels and similar labels, and significantly improves the model's distinction ability and extraction effect.
Smart Images

Figure CN116483942B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for extracting legal case elements based on a re-attention mechanism and contrastive loss, belonging to the field of natural language processing. Background Art
[0002] In recent years, the combination of artificial intelligence and justice has gradually become a powerful driving force for promoting the construction of intelligent courts. Legal texts contain important information such as lawsuits and defenses. Correctly using natural language processing technology to extract case element information from legal documents can effectively reduce labor costs. The extracted case element information can assist tasks such as legal text generation and case retrieval. The case elements extracted by previous related work can be divided into two categories: (1) Basic case elements, which appear in legal texts, such as entity information like time, place, and defendant in legal texts. Named entity recognition methods are often used to extract such case elements. (2) Semantic case elements, which need to be extracted from the case element label set by understanding the semantics of legal texts. Most multi-label text classification algorithms can be used for such case element extraction tasks. In the real-world scenario, the lack of datasets makes the legal text data corresponding to many legal case element labels extremely scarce. Simply applying multi-label text classification algorithms to the legal case element extraction task is difficult to achieve satisfactory performance. For this reason, many methods introduce label text information and cross-attention mechanisms to remove redundant information in legal texts to obtain high-quality distinguishable legal text representations. Seq2Seq-based models are also often used in multi-label text classification to improve classification performance by capturing label correlations. However, the current methods still have certain defects:
[0003] It is unable to effectively identify similar legal case element labels. There are similar labels in the legal case element label set, such as "creditor assigns creditor's rights" and "debtor assigns debts". In the attention matrix obtained only by using the cross-attention mechanism, the weight values of similar legal case element labels are also similar, making it difficult for the model to distinguish.
[0004] The Seq2Seq-based model has an exposure bias problem. In the inference stage, there is a great dependence between the predicted legal case element labels, and the prediction error will continue to spread.
[0005] To solve the above problems, we designed a method for extracting legal case elements based on a re-attention mechanism and contrastive loss.
[0006] Glossary:
[0007] Transformer Layer: Transformer is an encoder-decoder model. The RoBERTa pre-trained model used in the present invention, where the Transformer layer mentioned refers to the Transformer encoder, which is mainly composed of a multi-head attention mechanism, a feed-forward neural network, a residual connection, and a normalization layer.
[0008] softmax function: Normalized exponential function.
[0009] L2Norm: That is, L2 Normalization, which divides each value of a vector by the square root of the sum of the squares of the vector. Given a vector z = [z1, z2,..., z l , the calculation process of its L2Norm operation is as follows:
[0010]
[0011] row-wise: Perform calculation operations row by row
[0012] token: When the pre-trained model processes text, it first uses the model-specific tokenizer to split the text into a sequence of tokens, and queries the number of each token according to the vocabulary, so as to convert the text into a vector for model processing. Summary of the Invention
[0013] The present invention designs a legal case element extraction method based on re-attention mechanism and contrast loss, which is used to solve the problem of low extraction performance of similar legal case element labels and low-frequency legal case element labels in the legal case element label extraction task.
[0014] A legal case element extraction method based on re-attention mechanism and contrast loss includes the following steps:
[0015] Step 101: Legal text preprocessing step: Obtain a legal text data set, remove duplicate data in the legal text data set. The legal text data set includes several legal texts and corresponding legal case element label texts;
[0016] Step S102: Train the pre-trained model RoBERTa, and the training steps are as follows:
[0017] Step 1021: Legal text and legal case element label joint embedding step: Input the legal text data set into the pre-trained model RoBERTa together to obtain a fused representation, and form a joint text representation;
[0018] Step 1022: Legal case element label representation extraction step: Extract the legal case element label representation from the spliced text representation and use contrastive loss Optimize the legal case element label representation to obtain the legal text representation;
[0019] Step 1023: Legal text representation refinement step: Extract the legal text representation from the spliced text representation obtained in the legal text and legal case element label joint embedding step, and perform cross-attention twice on the legal text representation and the legal case element representation to obtain the legal text representation after removing redundant information;
[0020] Step 1024: Input the legal text representation O after removing redundant information into the classifier to obtain the probability distribution of each legal case element label and the loss of the classifier Obtain the total loss of the model λ is a hyperparameter between 0 and 1;
[0021] Step 1025: Train until the total loss is less than the preset value to obtain the trained model RoBERTa;
[0022] Step S103: Input the legal text of the legal case element to be extracted into the trained model RoBERTa to extract the corresponding legal case element.
[0023] For further improvement, the specific steps of step 1021 are as follows: Encode the legal text T = [x1, x2,..., x m to obtain its embedding where m is the length of the legal text, and the specific calculation of E T is shown in formula (1), x m represents the m-th token in the legal text, represents the embedding of the m-th token in the legal text; Encode each legal case element label text one by one, and the embedding of the i-th legal case element label text is where n i is the length of the i-th legal case element label text; represents the n i -th token in the legal case element label, represents the embedding of the n i -th token, and take the average of the embeddings E i of each legal case element label text to obtain The specific calculation is shown in formula (2); The legal case element label embedding is c represents the total number of legal case element tags; embedding the legal case element tags into E L and the legal text embedding E T concatenating them and inputting them into the 12-layer Transformer of the RoBERTa model to obtain the joint representation of the text and tags The specific implementation is shown in formula (3), where d represents the dimension of the hidden state, indicating that H is a matrix of dimension (c + m) × d;
[0024]
[0025]
[0026]
[0027] H = TransformerLayers(E L ||E T ) (3)
[0028] where mean() represents the operation of taking the average, || is the concatenation operation, EmbeddingLayer() represents the embedding layer of the pre-trained model RoBERTa, and Transformerlayers() represents the 12-layer Transformer of the pre-trained model RoBERTa.
[0029] For further improvement, the specific steps of step 1022 are as follows:
[0030] Extract the legal text representation and the legal case element tag text representation from the joint representation of the text and tags The calculation process is shown in formula (4), indicating that H L is a matrix of dimension c × d; using the contrastive loss to optimize the legal case element tag representation H L ; d is the dimension of the hidden state, and the calculation of the legal text representation is shown in formula (5), indicating that e T is a vector of dimension d; the specific calculation method of the contrastive loss is shown in formula (6):
[0031]
[0032]
[0033]
[0034]
[0035] Among them, is the set of legal case element labels applicable to the legal text; is the set of legal case element labels not applicable to the legal text; T represents the legal text; e i represents the i-th legal case element label, and both represent the representations after L2Norm() normalization; e p represents the representation of the label applicable to the legal text, e n represents the representation of the label not applicable to the legal text, p represents the label applicable to the legal text, and n represents the label not applicable to the legal text; H [:c] represents taking the first c tokens of H; H [c:] represents taking all tokens after the c-th token in H; τ is the temperature coefficient in the contrastive loss, used to adjust the attention degree to negative samples exp(x) is equivalent to e x , ∪ represents the union; c is the number of legal case element labels, represents the representation of the i-th token in the legal text representation.
[0036] For further improvement, the specific steps of step S1023 are as follows:
[0037] Step 10231: Obtain the legal text representation and the legal case element label text representation
[0038] Step 10232: Multiply the legal text representation H T and the legal case element label representation H L to obtain the attention weight The calculation process is shown in formula (8):
[0039] A = H L ·(H T ) T (8)
[0040] Step 10233: Use the column-wise and row-wise softmax() functions to normalize the two dimensions of the attention weight, and then use the row-wise and column-wise sum() functions for summation operations to obtain the weight of each token in the legal text and the weight of each legal case element label The specific calculation steps are as shown in formula (9):
[0041]
[0042]
[0043] softmax col ( ) indicates calculating softmax( ) by column, and softmax row ( ) indicates calculating softmax( ) by row; sum row ( ) indicates performing sum( ) by row, and sum col ( ) indicates performing sum( ) by column; represents the weight value of the m-th token in the legal text representation H T in; represents the weight value of the c-th label in the legal case element label representation H L in;
[0044] Step 10234: Multiply a T and a L separately with the legal text representation and the legal case element label representation to obtain the refined legal text representation and the legal case element label representation as shown in formula (10).
[0045]
[0046]
[0047] Step 10235: Use the self-attention mechanism selfAttention( ) to process the refined legal text representation and the legal case element label representation and then perform a multiplication operation again to obtain the refined attention weights as shown in formula (11).
[0048]
[0049]
[0050]
[0051] selfAttention( ) represents the self-attention mechanism, and T represents matrix transpose;
[0052] Step 10236: Multiply the refined attention weights Multiply with the original legal text representation to obtain the legal text representation with redundant information removed As shown in formula (12).
[0053]
[0054] For further improvement, the specific steps of step S1024 are as follows:
[0055] Input the legal text representation O with redundant information removed into the classifier to obtain the probability distribution of each legal case element label The loss of the classifier The calculation process of is as shown in formula (13):
[0056]
[0057]
[0058] Among them, FNNi is a feed-forward neural network used to predict the i-th legal case element label; W i and b i are the weight and bias respectively, and the value range of i is between 1 and c; O i indicates that O is a c×d matrix, expressed as O = [O1, O2,... O i ..., O c , where O i is a d-dimensional vector representing the representation related to the i-th label in the legal text; σ() is the sigmoid() activation function, and loss() is the binary cross-entropy loss function; the total loss of the model
[0059] Advantages of the present invention:
[0060] By using joint embedding and contrastive loss, the legal case element label representation fuses the correlation information between labels, which is beneficial to the extraction of low-frequency legal case element labels. In addition, by using the re-attention mechanism, it is possible to distinguish the text representations of similar legal case element labels, enabling the legal text representation to fuse the text information of legal case element labels and remove redundant information. The introduction of the re-attention mechanism and contrastive loss can effectively improve the extraction performance of legal case element labels. Brief Description of the Drawings
[0061] Figure 1 Shows the overall flowchart of the legal case element extraction method based on the re-attention mechanism and contrastive loss.
[0062] Figure 2 Shows the model diagram of the legal case element extraction method based on the re-attention mechanism and contrastive loss.
[0063] Figure 3 It shows the structural diagram of the re-attention mechanism module. Specific implementation mode
[0064] The technical solution of the present invention will be specifically described below in conjunction with examples.
[0065] As Figure 1 shown, the present invention provides a legal case element extraction method based on the re-attention mechanism and contrast loss, Figure 2 which is the model diagram of the legal case element extraction method based on the re-attention mechanism and contrast loss. Figure 3 It is the structural diagram of the re-attention mechanism module. The method includes the following steps.
[0066] 1. Legal text preprocessing step
[0067] Remove the duplicate data in the dataset to avoid affecting the evaluation of the model.
[0068] 2. Joint embedding step of legal text and legal case element labels
[0069] First, we encode the legal text T = [x1, x2,..., x m to obtain its embedding where m is the length of the legal text, and the specific calculation is shown in formula (1). We encode each legal case element label text one by one, and the embedding of the i-th legal case element label text is where n i is the length of the i-th legal case element label text. Next, we calculate the average value of the embeddings of each legal case element label to obtain The specific calculation is shown in formula (2). The embedding of the legal case element label is Finally, we splice the label embedding E L and the legal text embedding E T and input them into a 12-layer Transformer to obtain the joint representation of the text and the label The specific implementation is shown in formula (3):
[0070]
[0071]
[0072]
[0073] H = TransformerLayers(E L ||ET ) (3)
[0074] Among them, mean() represents the operation of calculating the average value, || is the concatenation operation, EmbeddingLayer() represents the embedding layer of the pre-trained model RoBERTa, and TransformerLayers() represents the 12-layer Transformer of the pre-trained model RoBERTa.
[0075] 3. Steps for extracting legal case element label representations
[0076] Extract the legal case element label representation from the joint text representation obtained in the joint embedding step of the legal text and legal case element labels The calculation process is shown in formula (4). In addition, we use the contrastive loss to optimize the legal case element label representation. It should be noted that d is the dimension of the hidden state, and the legal text representation is calculated as shown in formula (5). The specific calculation method of the contrastive loss is shown in formula (6):
[0077]
[0078]
[0079]
[0080]
[0081] Among them, is the set of legal case element labels applicable to the legal text. is the set of legal case element labels not applicable to the legal text. T represents the legal text. and are both representations that have undergone L2Norm() normalization processing. τ is the temperature. c is the total number of legal case element labels.
[0082] 4. Steps for refining the legal text representation
[0083] Step 201: Extract the legal text representation and the legal case element label text representation from the representation obtained in the legal case element label representation extraction step and the legal case element label text representation The specific calculation method is shown in formula (7):
[0084]
[0085]
[0086] Step 202: Multiply the legal text representation H T and the legal case element label representation H L to obtain the attention weights The calculation process is shown in formula (8).
[0087] A = H L · (H T ) T (8)
[0088] Step 203: Use the column-wise and row-wise softmax() functions to normalize the two dimensions of the attention weights, and then use the row-wise and column-wise sum() functions for summation operations to obtain the weights of each token in the legal text and the weights of each legal case element label Figure 3 Intuitively shows this operation process. The specific calculation steps are as shown in formula (9):
[0089]
[0090]
[0091] Step 204: Multiply the two normalized attention weights obtained in Step 203 with the legal text representation and the legal case element label representation respectively to obtain the refined legal text representation and the legal case element label representation as shown in formula (10).
[0092]
[0093]
[0094] Step 205: Use the self-attention mechanism selfAttention() to process the refined legal text representation and the legal case element label representation obtained in Step 204, and then perform another multiplication operation to obtain the refined attention weights as shown in formula (11).
[0095]
[0096]
[0097]
[0098] Step 206: Multiply the refined attention weights obtained in Step 205 by the original legal text representation to obtain the legal text representation with redundant information removed. As shown in formula (12).
[0099]
[0100] 5. Legal Case Element Label Prediction Step
[0101] Input the legal text representation O with redundant information removed obtained in the legal text representation refinement step into the classifier to obtain the probability distribution of each legal case element label. The loss of the classifier The calculation process is as shown in formula (13).
[0102]
[0103]
[0104] where, FNN i is a feed-forward neural network used to predict the i-th legal case element label, W i and b i are the weight and bias respectively. σ() is the sigmoid() activation function, and loss() is the binary cross-entropy loss function.
[0105] In the training stage, use the contrast loss for optimizing the legal case element label representation and the classifier loss to jointly train the model, that is, the total loss of the model where, λ is a hyperparameter between 0 and 1, used to control the proportion of the contrast loss. In the test stage, use the probability distribution of the legal case element labels obtained by the classifier and the set threshold to determine whether the legal text contains legal case element labels.
[0106] In the specific implementation process, we implemented the legal case element extraction method using the Pytorch deep learning framework. To verify the effectiveness of the proposed method, we tested the proposed method on the CAIL2019 legal case element extraction dataset. During the test, we used the same settings as the existing work and adopted multiple evaluation metrics such as Micro-F1, Macro-F1, and Avg-F1 to more comprehensively evaluate the extraction performance of legal case element labels. In Table 1, we give the comparison experiment results between the method of the present invention and other methods. It should be noted that MiF1, MaF1, and AF1 in Table 1 are the abbreviations of Micro-F1, Macro-F1, and Avg-F1 respectively.
[0107] Table 1 Comparison Experiment Results
[0108]
[0109] The CAIL2019 legal case element extraction dataset contains data in three domains: marriage and family, labor disputes, and loan contracts. From the comparative experiments, it can be found that the performance of this method on the datasets in these three domains is better than that of the comparative methods. Especially in the case of fewer datasets, the performance improvement of this method is higher, such as in the labor dispute and loan contract datasets. It can also be found that the MaF1 values of this method on the datasets in the three domains have been greatly improved. It should be noted that the MaF1 value evaluates the average of the F1 values of each legal case element label. The above improvements all indicate that this method can well solve the problems of low performance in extracting low-frequency and similar legal case element labels. The performance of the RoBERTa baseline method far exceeds that of the models based on CNN or RNN, which indicates that obtaining high-quality legal text features can effectively improve the performance of extracting legal case element labels. The LP-MTC method uses a masking mechanism on the basis of the RoBERTa method to capture the correlation between legal case element labels and thus obtains better performance. This shows that capturing the correlation between legal case element labels plays a positive role in improving the performance of legal case element extraction.
[0110] By introducing the re-attention mechanism and contrastive loss to improve the quality of legal text features and capture the correlation between legal case element labels, this method achieves the optimal performance. To better verify the effects of these two modules, we designed ablation experiments, as shown in Table 2. It should be noted that RA and CL represent the re-attention mechanism and contrastive loss respectively. In addition, this method is an improvement based on the RoBERTa method, and the methods without RA and CL in Table 2 represent the RoBERTa method.
[0111] Table 2 Ablation Experiments
[0112]
[0113] As can be seen from Table 2, adding RA or CL to the RoBERTa baseline method can bring performance improvements. Stacking RA and CL on the RoBERTa method can achieve the best performance. These all show that the RA and CL modules are helpful for improving the performance of extracting legal case element labels, and the improvement effect is significant.
Claims
1. A legal case element extraction method based on a re-attention mechanism and contrastive loss, characterized in that It includes the following steps: Step 101: Legal text preprocessing step: Obtain a legal text dataset, remove duplicate data in the legal text dataset. The legal text dataset includes several legal texts and corresponding legal case element label texts; Step S102: Train the pre-trained model RoBERTa. The training steps are as follows: Step 1021: Joint embedding step of legal text and legal case element labels: Input the legal text dataset into the pre-trained model RoBERTa together to obtain a joint text representation; Step 1022: Legal case element label representation extraction step: Extract the legal case element label representation from the joint text representation and use contrastive loss Optimize the legal case element label representation to obtain the legal text representation; Step 1023: Legal text representation refinement step: Extract the legal text representation from the joint text representation obtained in the joint embedding step of legal text and legal case element labels, and perform cross-attention twice between the legal text representation and the legal case element labels to obtain a legal text representation after removing redundant information; Step 1024: Input the legal text representation O with redundant information removed into the classifier to obtain the probability distribution of each legal case element label and the loss of the classifier to obtain the total loss of the model λ is a hyperparameter between 0 and 1; Step 1025: Train until the total loss is less than the preset value to obtain the trained model RoBERTa; Step S103: Input the legal text of the legal case elements to be extracted into the trained model RoBERTa to extract the corresponding legal case elements.
2. The method for extracting legal case elements based on the re-attention mechanism and contrastive loss as described in claim 1, wherein The specific steps of step 1021 are as follows: Encode the legal text T = [x1, x2,..., x m to obtain its embedding where m is the length of the legal text, and E T is specifically calculated as shown in formula (1), x m represents the m-th token in the legal text, represents the embedding of the m-th token in the legal text; Encode each legal case element label text one by one, and the embedding of the i-th legal case element label text is where n i is the length of the i-th legal case element label text; represents the n i -th token in the legal case element label, represents the embedding of the n i -th token, and take the average of the embeddings E i of each legal case element label text to obtain The specific calculation is shown in formula (2); the legal case element label embedding is c represents the total number of legal case element labels; Concatenate the legal case element label embedding E L and the legal text embedding E T and input them into the 12th layer Transformer of the RoBERTa model to obtain the joint representation of the text and the label The specific implementation is shown in formula (3), where d represents the dimension of the hidden state, represents that H is a (c + m) × d-dimensional matrix; H = TransformerLayers(E L || E T ) (3) Among them, mean() represents the operation of calculating the average value, || is the concatenation operation, EmbeddingLayer() represents the embedding layer of the pre-trained model RoBERTa, and TransformerLayers() represents the 12-layer Transformer of the pre-trained model RoBERTa.
3. The method for extracting legal case elements based on the re-attention mechanism and contrastive loss according to claim 1, wherein The specific steps of the said Step 1022 are as follows: Extract the legal text representation from the joint representation of text and tags and the legal case element tag text representation The calculation process is shown in formula (4), where H is a matrix of c×d dimensions; use the contrastive loss L to optimize the legal case element tag representation H where d is the dimension of the hidden state, and the legal text representation L is calculated as shown in formula (5), where e is a vector of d dimensions; the specific calculation method of the contrastive loss is shown in formula (6): T Among them, is the set of legal case element labels applicable to the legal text; is the set of legal case element labels not applicable to the legal text; T represents the legal text; e i represents the representation of the i-th legal case element label, and are both representations that have undergone L2Norm(·) normalization; e p represents the representation of the labels applicable to the legal text, e n represents the representation of the labels not applicable to the legal text, p represents the labels applicable to the legal text, and n represents the labels not applicable to the legal text; H [:c] represents taking the representation of the first c tokens of H; H [c:] represents taking the representation of all tokens after the c-th token in H; τ is the temperature coefficient in the contrastive loss, used to adjust the attention to negative samples ; exp(x) is equivalent to e x , ∪ represents the union; c is the number of legal case element labels, represents the representation of the i-th token in the legal text representation.
4. The method for extracting legal case elements based on the re-attention mechanism and contrastive loss as described in claim 1, wherein The specific steps of the said Step S1023 are as follows: Step 10231: Obtain the legal text representation and the legal case element label text representation Step 10232: Multiply the legal text representation H T by the legal case element label representation H L to obtain the attention weight The calculation process is shown in Equation (8): A = H L ·(H T ) T (8) Step 10233: Normalize the two dimensions of the attention weights using the column-wise and row-wise softmax() functions, and then perform a summation operation using the row-wise and column-wise sum() functions to obtain the weights of each token in the legal text and the weights of each legal case element label The specific calculation steps are shown in formula (9) as follows: softmax col () means calculating softmax() by column, softmax row () means calculating softmax() by row; sum row () means performing sum() by row, sum col () means performing sum() by column; represents the weight value of the m-th token in the legal text representation H T and represents the weight value of the c-th label in the legal case element label representation H L ; Step 10234: Multiply a T and a L with the legal text representation and the legal case element label representation respectively to obtain the refined legal text representation and the legal case element label representation as shown in formula (10); Step 10235: Use the self-attention mechanism selfAttention() to represent the refined legal text and the representation of legal case element labels for processing, and then perform a multiplication operation again to obtain the refined attention weights as shown in formula (11); selfAttention() represents the self-attention mechanism, and T represents matrix transpose; Step 10236: Multiply the refined attention weights by the original legal text representation to obtain a legal text representation with redundant information removed as shown in formula (12); 。 5. The method for extracting legal case elements based on the re-attention mechanism and contrastive loss according to claim 1, wherein The specific steps of the said Step S1024 are as follows: Input the legal text representation O after removing redundant information into the classifier to obtain the probability distribution of each legal case element label The loss of the classifier The calculation process is shown in formula (13): Among them, FNN i is a feedforward neural network used to predict the label of the i-th legal case element; W i and b i are the weight and bias respectively, and the value range of i is between 1 and c; O i indicates that O is a c×d matrix, expressed as O = [O1, O2,... O i ..., O c , where O i is a d-dimensional vector representing the representation related to the i-th label in the legal text; σ() is the sigmoid() activation function, and loss() is the binary cross-entropy loss function; the total loss of the model
Citation Information
Patent Citations
Method for identifying case elements integrated with label information
CN114764913A
Labeling method and apparatus for named entity recognition of legal instrument
US11615247B1