Legal text analysis method based on BERT model

Through the legal text analysis method based on the BERT model, adversarial samples were generated and LMI adjustments were performed, which solved the problem of sample imbalance in the legal data set and the confusing crimes easily, significantly improving the prediction accuracy and overall prediction performance of low-frequency crimes.

CN120012763AActive Publication Date: 2025-05-16GUANGDONG OCEAN UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510148674.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-16
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

The existing legal data sets have sample imbalance, which leads to low-predictive accuracy of low-frequency crimes and is easy to confuse the crimes. The existing methods mainly focus on high-frequency crimes and ignore the handling of the crimes of small samples.

Method used

The legal text analysis method based on the BERT model is adopted to enhance the robustness of the model by generating adversarial samples and adjusting the Logit value (LMI adjustment), alleviating the performance degradation caused by data imbalance, and improving the prediction accuracy of low-frequency charges.

Benefits of technology

Effectively improves the prediction accuracy of low-frequency crimes, reduces confusion among crimes, and improves overall prediction performance, especially when dealing with unbalanced data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012763A_ABST
    Figure CN120012763A_ABST
Patent Text Reader

Abstract

The invention discloses a legal text analysis method based on a BERT model, and belongs to the technical field of data analysis, and the method comprises the following steps: S1, collecting an original input sample, inputting the original input sample into the BERT model, and updating the BERT model; s2, generating an adversarial sample according to the original input sample; and S3, inputting the original input sample and the adversarial sample into the updated BERT model, and performing LMI adjustment to obtain a classification result of the original input sample. According to the method, performance reduction caused by data imbalance is relieved, meanwhile, the prediction accuracy of low-frequency criminal names is improved, interference between easily-confused criminal names is effectively reduced, and therefore the overall prediction performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data analysis, and specifically relates to a legal text analysis method based on a BERT model. Background Art

[0002] Legal charge prediction aims to predict the final charge based on the factual description in the case, and plays an important role in legal intelligence systems. However, due to practical factors, most existing legal data sets have sample imbalance problems, and most existing prediction methods focus on high-frequency charges, while paying less attention to low-frequency charges. For example, according to the statistics of the single-label charge samples of the first_stage train in the CAIL2018 data set, the 10 common charges account for about 78.7%, while the 115 uncommon charges account for only 0.96%. In addition, due to the similarities between legal provisions, the complexity of case plots, and the diversity of judicial interpretations and case facts, the confusion between charges is particularly prominent. Most existing studies focus on the prediction of common charges, ignoring the problem of handling charges with few samples. Summary of the invention

[0003] In order to solve the above problems, the present invention proposes a legal text analysis method based on the BERT model.

[0004] The technical solution of the present invention is: a legal text analysis method based on the BERT model comprises the following steps:

[0005] S1. Collect original input samples, input the original input samples into the BERT model, and update the BERT model;

[0006] S2, generate adversarial samples based on the original input samples;

[0007] S3. Input the original input sample and the adversarial sample into the updated BERT model, perform LMI adjustment, and obtain the classification result of the original input sample.

[0008] Furthermore, S1 includes the following sub-steps:

[0009] S11. Collect the original input sample and determine the prediction probability of the BERT model for the original input sample;

[0010] S12. Generate a cross entropy loss function based on the predicted probability of the original input sample;

[0011] S13. Based on the cross entropy loss function, the gradient descent method is used to update the parameters of the BERT model.

[0012] Furthermore, in S11, the predicted probability of the original input sample is The calculation formula is:

[0013] ;

[0014] In the formula, Represents the original input sample The score value of the class label, Represents the original input sample The score value of the class label, represents the exponential function, Indicates the number of tags.

[0015] Furthermore, in S12, the cross entropy loss function The expression is:

[0016] ;

[0017] In the formula, represents the predicted probability of the original input sample, represents the logarithmic function, Indicates the number of tags, Indicates Indicator function for the class labels.

[0018] Furthermore, in S13, the calculation formula for updating the parameters of the BERT model is:

[0019] ;

[0020] In the formula, represents the loss function, Indicates The updated parameters, Indicates The updated parameters, Represents the learning rate.

[0021] Furthermore, S2 includes the following sub-steps:

[0022] S21, iteratively update the perturbation of the updated BERT model;

[0023] S22. Generate adversarial samples based on the iteratively updated perturbation results.

[0024] Furthermore, in S21, the calculation formula for iteratively updating the disturbance is:

[0025] ;

[0026] In the formula, Indicates The perturbation of the update, Indicates The perturbation of the update, represents the loss function, Represents a projection operation, represents the step length, represents the label corresponding to the original input sample, represents the maximum amplitude of the disturbance, represents the symbolic function, Represents the loss function for The gradient of the perturbation of the update, Represents the original input sample.

[0027] Furthermore, in S22, adversarial samples The expression is:

[0028] ;

[0029] In the formula, Indicates The perturbation of the update, Represents the original input sample.

[0030] Furthermore, in S3, the classification results of the original input samples The calculation formula is:

[0031] ;

[0032] In the formula, Represents the score value of the original input sample.

[0033] The beneficial effects of the present invention are as follows: in response to the problems of sample imbalance and confusion between crimes in legal data sets, the present invention proposes a legal text analysis method based on the BERT model, which aims to alleviate the performance degradation caused by data imbalance by enhancing the robustness of the model, while improving the prediction accuracy of low-frequency crimes and effectively reducing the interference between easily confused crimes, thereby improving the overall prediction performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 This is a flowchart of the legal text analysis method based on the BERT model;

[0035] Figure 2 For function Schematic diagram of

[0036] Figure 3 A schematic diagram of the case. DETAILED DESCRIPTION

[0037] The embodiments of the present invention will be further described below in conjunction with the accompanying drawings.

[0038] like Figure 1As shown, the present invention provides a legal text analysis method based on the BERT model, comprising the following steps:

[0039] S1. Collect original input samples, input the original input samples into the BERT model, and update the BERT model;

[0040] S2, generate adversarial samples based on the original input samples;

[0041] S3. Input the original input sample and the adversarial sample into the updated BERT model, perform LMI adjustment, and obtain the classification result of the original input sample.

[0042] In the task of legal charge prediction, it is usually regarded as a text classification problem. The goal is to describe X={x1,x2,...,x n}, where x i is a word in the vocabulary T, n is the length of the word sequence, and predicts a unique crime label y∈Y, where Y = {y1,y2,...,y m} represents the set of all possible crime labels.

[0043] In this embodiment of the present invention, S1 includes the following sub-steps:

[0044] S11. Collect the original input sample and determine the prediction probability of the BERT model for the original input sample;

[0045] S12. Generate a cross entropy loss function based on the predicted probability of the original input sample;

[0046] S13. Based on the cross entropy loss function, the gradient descent method is used to update the parameters of the BERT model.

[0047] In the embodiment of the present invention, in S11, the predicted probability of the original input sample The calculation formula is:

[0048] ;

[0049] In the formula, Represents the original input sample The score value of the class label, Represents the original input sample The score value of the class label, represents the exponential function, Indicates the number of tags.

[0050] In the embodiment of the present invention, in S12, the cross entropy loss function The expression is:

[0051] ;

[0052] In the formula, represents the predicted probability of the original input sample, represents the logarithmic function, Indicates the number of tags, Indicates Indicator function for the class labels.

[0053] In the embodiment of the present invention, in S13, the calculation formula for updating the parameters of the BERT model is:

[0054] ;

[0055] In the formula, represents the loss function, Indicates The updated parameters, Indicates The updated parameters, Represents the learning rate.

[0056] In this embodiment of the present invention, S2 includes the following sub-steps:

[0057] S21, iteratively update the perturbation of the updated BERT model;

[0058] S22. Generate adversarial samples based on the iteratively updated perturbation results.

[0059] In the embodiment of the present invention, in S21, the calculation formula for iteratively updating the disturbance is:

[0060] ;

[0061] In the formula, Indicates The perturbation of the update, Indicates The perturbation of the update, represents the loss function, Represents a projection operation, represents the step length, represents the label corresponding to the original input sample, represents the maximum amplitude of the disturbance, represents the symbolic function, Represents the loss function for The gradient of the perturbation of the update, Represents the original input sample.

[0062] In the embodiment of the present invention, in S22, the adversarial sample The expression is:

[0063] ;

[0064] In the formula, Indicates The perturbation of the update, Represents the original input sample.

[0065] In the embodiment of the present invention, in S3, the classification result of the original input sample The calculation formula is:

[0066] ;

[0067] In the formula, Represents the score value of the original input sample.

[0068] In the embodiment of the present invention, in the task of legal crime classification, due to the high similarity between certain categories, confusion is easily caused, resulting in classification errors. These confused categories are often shown in the classification output layer as having relatively close raw scores (Logits). In order to reduce the impact of confused categories on classification accuracy, the present invention proposes a method based on adjusting the Logit value (LMI) to dynamically adjust the current highest Logit value. Specifically, the highest Logits value L1 is adjusted using the following formula: ;in, represents the highest adjusted Logit value, Indicates the current highest Logit value, Indicates the current second highest Logit value, Represents a stabilization factor, which is used to adjust the original highest Logit value and the influence weight of the dynamic adjustment part, so as to find a reasonable compromise between reducing confusion between similar categories and maintaining model stability.

[0069] like Figure 2 As shown, the function The role is based on The difference dynamically adjusts the adjustment range.

[0070] The adjustment logic is as follows:

[0071] (1) When the difference is small: When it is small, it means that the two categories in the current classification result are highly confused. In this case, The value of is also small, avoiding the The category represented by may lead to misclassification. The value of decision, thereby reducing the impact of confusion on classification decisions.

[0072] (2) When the difference is moderate: When it is in the moderate range (such as 0.3 to 1.2), The value of increases as the difference increases. At this time, the adjusted The value will be pulled higher, thus further expanding and The larger the gap, the greater the distinction between the two categories will be, thus enhancing the discrimination ability of the classification model.

[0073] (3) When the difference is large: When it is already large, it means that the distinction between the two categories is high and no additional adjustment is needed. The value of decreases as the difference increases. The impact has stabilized.

[0074] Through this method, the adjusted Logits value can effectively reduce the classification errors caused by category confusion and improve the classification accuracy and robustness of the model. In summary, the LMI adjustment method proposed in the present invention enhances the discrimination ability between confused categories by dynamically adjusting the maximum Logits value and maintains stability when the category differences are large, effectively improving the performance of the classification task.

[0075] The following is a detailed description of the performance of the model in the task of predicting confusing charges in conjunction with specific examples. Figure 3 As shown in the figure, in this case, the defendant was ultimately convicted of the crime of "forging, altering, buying and selling official documents, certificates and seals of state organs". However, because the case involves "certificates", "seals" and other noise data related to crimes, traditional models are often prone to misjudging it as "forging the seals of companies, enterprises, institutions and people's organizations". In contrast, the model proposed in the present invention significantly reduces the incidence of such misjudgments through more effective feature extraction and contextual semantic understanding, and successfully predicts the correct crime. The specific results are shown in Table 1.

[0076] Table 1

[0077] Model Prediction results BERT CCP4CI Forgery of seals of companies, enterprises, institutions and people's organizations Forgery, alteration, sale and purchase of official documents, certificates and seals of state organs True label Crime of forging, altering, buying and selling official documents, certificates and seals of state organs

[0078] In order to further verify the effectiveness of the model in dealing with confusing crime labels, the present invention selected several representative confusing crime groups from the entire Charge_L data set for analysis based on the misjudgment in the experiment. Among them, the medium and low frequency confusing crimes include "destroying transportation vehicles" and "destroying transportation facilities", "destroying supervision order" and "abusing supervised persons", and "smuggling" and "smuggling goods and articles prohibited from import and export by the state", with a sample number of less than 115, and the high frequency confusing crimes include "producing and selling food that does not meet safety standards", "producing and selling fake and inferior products" and "producing and selling toxic and harmful food", "abusing power" and "negligence in duty", "opening a casino" and "gambling", and "forging, altering, buying and selling official documents, certificates, and seals of state organs" and "forging seals of companies, enterprises, institutions, and people's groups" with a sample number of about 800. As shown in Table 2, by analyzing the F1 values ​​of these crime groups, it can be found that the performance of the model has improved in both low-frequency and high-frequency confusing crimes, which shows that the model has strong robustness and effectiveness in dealing with confusing crimes.

[0079] Table 2

[0080] Charge Type Mid-low frequency High frequency BERT CCP4CI 74.4092.73(⬆18.33%) 95.4996.51(⬆1.02%)

[0081] In order to further verify the advantages of the model in processing unbalanced data sets, the present invention conducts a comparative analysis of the F1 values ​​of different frequency categories of the entire Charge_L data set. Specifically, according to the distribution of the number of crime samples, the crimes are divided into three categories: low-frequency crimes with a sample number less than or equal to 30, medium-frequency crimes with a sample number between 30 and 150, and high-frequency crimes with a sample number between 150 and 800. The experimental results are shown in Table 3. The model performs better than the baseline method in the low-frequency, medium-frequency, and high-frequency crime categories, especially in the low-frequency crime category. The F1 value of the model is improved by more than 43% compared with the baseline method. This significant improvement shows that the model is effective in dealing with the problem of unbalanced sample distribution.

[0082] Table 3

[0083] Charge Type Low frequency Medium frequency High frequency Charge Number 20 57 116 BERT CCP4CI 40.3983.75(⬆43.36%) 92.7895.07(⬆2.29%) 96.8496.93(⬆0.09%)

[0084] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.

Claims

1. A legal text analysis method based on the BERT model, characterized in that: The following steps are involved: S1. Collect original input samples, input the original input samples into the BERT model, and update the BERT model; S2, generate adversarial samples based on the original input samples; S3. Input the original input sample and the adversarial sample into the updated BERT model, perform LMI adjustment, and obtain the classification result of the original input sample.

2. The legal text analysis method based on the BERT model according to claim 1 is characterized in that: The S1 comprises the following sub-steps: S11. Collect the original input sample and determine the prediction probability of the BERT model for the original input sample; S12. Generate a cross entropy loss function based on the predicted probability of the original input sample; S13. Based on the cross entropy loss function, the gradient descent method is used to update the parameters of the BERT model.

3. The legal text analysis method based on the BERT model according to claim 2 is characterized in that: In S11, the predicted probability of the original input sample The calculation formula is: ; In the formula, Represents the original input sample The score value of the class label, Represents the original input sample The score value of the class label, represents the exponential function, Indicates the number of tags.

4. The legal text analysis method based on the BERT model according to claim 2 is characterized in that: In S12, the cross entropy loss function The expression is: ; In the formula, represents the predicted probability of the original input sample, represents the logarithmic function, Indicates the number of tags, Indicates Indicator function for the class labels.

5. The legal text analysis method based on the BERT model according to claim 2 is characterized in that: In S13, the calculation formula for updating the parameters of the BERT model is: ; In the formula, represents the loss function, Indicates The updated parameters, Indicates The updated parameters, Represents the learning rate.

6. The legal text analysis method based on the BERT model according to claim 1, characterized in that: The S2 comprises the following sub-steps: S21, iteratively update the perturbation of the updated BERT model; S22. Generate adversarial samples based on the iteratively updated perturbation results.

7. The legal text analysis method based on the BERT model according to claim 6 is characterized in that: In S21, the calculation formula for iteratively updating the disturbance is: ; In the formula, Indicates The perturbation of the update, Indicates The perturbation of the update, represents the loss function, Represents a projection operation, represents the step length, represents the label corresponding to the original input sample, represents the maximum amplitude of the disturbance, represents the symbolic function, Represents the loss function for The gradient of the perturbation of the update, Represents the original input sample.

8. The legal text analysis method based on the BERT model according to claim 6, characterized in that: In S22, the adversarial sample The expression is: ; In the formula, Indicates The perturbation of the update, Represents the original input sample.

9. The legal text analysis method based on the BERT model according to claim 1, characterized in that: In S3, the classification result of the original input sample The calculation formula is: ; In the formula, Represents the score value of the original input sample.

Citation Information

Patent Citations

  • Adversarial sample defense method based on significance adversarial training

    CN112766401A

  • BERT short text sentiment analysis method for improving training mode

    CN114757182A

  • Method and device for pruning deep neural network

    CN114969340A

  • Criminal name prediction method and system, readable storage medium and computer equipment

    CN116308898A

  • Text processing model training method, and text processing method and apparatus

    US20220180202A1