Prediction method for intellectual property criminal responsibility, applicable laws and legal provisions and criminal period
By constructing a hybrid expert model based on a pre-trained model and a hierarchical heterogeneous low-rank adaptive method, the problems of limited training samples and high costs in intellectual property criminal cases are solved, and efficient prediction of criminal liability, applicable legal provisions, and sentence is achieved.
Patent Information
- Application Number
- CN202510974584.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-10-21
AI Technical Summary
In predicting judgments in intellectual property criminal cases, existing technologies suffer from limited training sample data, high training time costs, and an inability to effectively utilize existing case data. Furthermore, existing models have a large number of structural parameters, resulting in high prediction costs and making it difficult to achieve efficient predictions of criminal liability, applicable legal provisions, and sentence duration.
We adopt a low-rank adaptive method based on pre-trained models and hierarchical heterogeneity. By constructing a hybrid expert model structure and designing a low-rank matrix calculation layer, we reduce the number of model parameters and use existing case data as a pre-trained model to optimize the model's multi-task prediction capability.
It enables efficient and low-cost prediction of criminal charges, applicable legal provisions, and sentences in intellectual property criminal cases, improving the model's predictive performance and training efficiency while reducing overall training costs.
Smart Images

Figure CN120821844A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of natural language processing and machine learning technology, and relates to a method for predicting intellectual property criminal liability, applicable laws and regulations, and prison sentences based on a pre-training model and a hierarchically heterogeneous low-rank adaptive method. Background Art
[0002] Currently, research on deep learning models focuses on large language models (LLMs), often referred to as Large Language Models (LLMs) due to their numerous encoder and decoder modules, larger parameters, and enhanced predictive reasoning capabilities. However, due to the sheer number of parameters and the complex training and updating process, updating large models incurs significant energy and time costs, making intelligent technologies prohibitive for organizations and institutions that urgently need these models. Therefore, to optimize the training and updating process for large models while reducing training costs, researchers have proposed a variety of new technologies and approaches. Among these, combining pre-trained models with Low-Rank Adaptation (LoRA) techniques has become a hot topic in research and application for achieving low-cost and efficient parameter updates for large models. This approach updates both pre-trained and large models in a relatively limited amount of time by training an "additional model" with a relatively small number of parameters.
[0003] In the field of sentencing budget research for intellectual property criminal cases, while intellectual property civil disputes and cases have been increasing in recent years, criminal cases remain relatively rare and contain limited descriptive information, making it difficult to support the dual requirements of quantity and quality in training corpora for machine learning methods. Consequently, the proportion of intellectual property criminal cases in publicly available judgment prediction datasets is relatively small compared to other types of cases involving infringements on people's rights to life and property. Therefore, the starting point and goal of this methodological research is to leverage the large number of intellectual property civil cases to enhance the training efficiency and predictive power of models for intellectual property criminal cases.
[0004] Among existing technical solutions, Yue et al. separated the information related to the crime from the irrelevant information in the case description. They then used a fusion method to predict the crime and applicable legal provisions. They also used a graph-based computational approach to calculate the similarity between legal provisions, thereby improving the accuracy of legal provision prediction (reference: L. Yue et al., “NeurJudge: A Circumstance-aware Neural Framework for Legal Judgment Prediction,” in Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, Virtual Event Canada: ACM, July 2021, pp. 973–982. doi:10.1145 / 3404835.3462826.). This method requires manual pre-processing to identify similarities between legal provisions and crime descriptions, and still requires the assistance of experts with certain expertise. Furthermore, the model structure is relatively complex and has a high number of parameters, which means that the prediction process still takes a certain amount of time. Furthermore, it is not possible to use existing models as pre-trained models, and a large amount of data must be used to train the model from scratch.
[0005] Bhambhoria et al. proposed a label fusion method for charge prediction. They employed a curriculum learning strategy to gradually learn from simple samples to complex ones, achieving good charge prediction results with low resource consumption (reference: R. Bhambhoria, H. Liu, S. Dahan, and X. Zhu, “Interpretable Low-Resource Legal Decision Making,” AAAI, vol. 36, no. 11, pp. 11819–11827, Jun. 2022, doi:10.1609 / aaai.v36i11.21438). This method employed a label fusion joint attention mechanism to extract key semantic information from case descriptions and a curriculum learning strategy to achieve good prediction performance with low sample size. However, the joint attention mechanism employed tends to capture semantic information that is positioned later, making it susceptible to the influence of semantic information at different positions on the prediction performance. At the same time, experts are required to first judge the difficulty of the data samples, and then determine the order of training to maximize the effectiveness of the course learning mechanism. Therefore, the overall cost of training is still difficult to effectively reduce. Summary of the Invention
[0006] To address the challenges of the existing technology, the present invention aims to provide a method for predicting intellectual property criminal liability, applicable laws and regulations, and prison sentences based on a pre-trained model and a hierarchically heterogeneous, low-rank adaptive approach. This method utilizes deep neural network natural language processing technology to predict civil torts and multiple applicable laws and regulations. The present invention designs and constructs a multi-task prediction model based on a hybrid expert model architecture. This model optimizes the prediction of applicable laws and regulations while providing infringement predictions, thereby enhancing the persuasiveness and interpretability of infringement predictions.
[0007] This invention proposes specific solutions to the problems of limited training sample data in the prediction of judgments in intellectual property criminal cases, and the difficulty in predicting criminal liability, applicable laws and regulations, and prison sentences.
[0008] First, a model design method for predicting the adjudication of intellectual property criminal cases based on a pre-trained model is proposed. The intellectual property criminal liability prediction model can be trained based on the existing civil case adjudication prediction model or the existing criminal case adjudication prediction model, so as to solve the problem that the case data set has a small sample size and the training time cost is high, and it is impossible to fully utilize the existing case data or the existing judgment prediction model as the training basis.
[0009] Second, a low-rank adaptive model structure was designed. For deep learning models with multi-layered structures, computational parameter matrices with varying ranks are designed based on the model's depth. This design pattern reduces the overall number of model parameters while maintaining the model's predictive power.
[0010] The technical solution of the present invention is:
[0011] A method for predicting intellectual property criminal liability, applicable laws and regulations, and sentence length, comprising the following steps:
[0012] 1) For each case in the criminal case dataset, label the crime, applicable criminal law and judicial interpretation, and sentence to obtain a training sample;
[0013] 2) Constructing a judgment prediction model, using the training samples to train the judgment prediction model for predicting the crime, applicable criminal law provisions and legal interpretations, and sentence in a criminal case; wherein the judgment prediction model includes an encoder module, a pre-trained model, a criminal crime multi-classification prediction module, a criminal law provision and judicial interpretation multi-label prediction module, and a sentence multi-classification prediction module; adding a low-rank matrix calculation layer to each level of a multi-level encoder or decoder model based on a Transformer architecture to obtain the pre-trained model; the method for training the judgment prediction model is:
[0014] 21) The encoder module encodes the input training sample to obtain the case description text vector X of the training sample doc_em , crime text vector X accusation_em , applicable criminal law articles and judicial interpretation text vector X law_em And normalize them respectively to get X doc 、X accusation 、X law ; Among them, X doc is the normalized case description text vector, X accusation is the normalized criminal charge text vector, X law is the normalized legal article and judicial interpretation text vector; k is the number of words contained in the text after word segmentation, and emb is the word embedding space dimension;
[0015] 22) The pre-trained model is based on the input X doc 、X accusation 、X law , through the formula Calculate logits n ; Among them, for X doc 、X accusation 、X law Do linear calculation to get vector X LoRA ; ΔW j The low-rank matrix calculation layer added for the jth level, j = 1 to n, n is the total number of levels of the pre-trained model;
[0016] 23) The criminal offense multi-classification prediction module is based on logits n Predict the output result S of the crime classification multilayer perceptron accusation The multi-label prediction module of criminal law articles and judicial interpretations is based on logits n Predict the output results S of the multi-layer perceptron of criminal law and judicial interpretation law The sentence multi-classification prediction module is based on logits n The predicted sentence classification multilayer perceptron output result S time ; Then, the loss value is calculated based on the prediction results and the annotation information to optimize the referee prediction model;
[0017] 3) For a criminal case to be predicted, the case facts are input into the trained judgment prediction model to obtain the crime, applicable criminal law provisions and judicial interpretations, and sentence of the criminal case to be predicted.
[0018] Furthermore, through the loss function L = λloss accusation +βloss law +γloss timeThe loss value L is calculated; where λ, β, and γ are linear combination coefficients, λ+β+γ=1; Ω neg is the set of non-applicable laws in the prediction results, Ω pos is the set of applicable laws in the prediction results, is the true value of the i-th category in the criminal charges, For S accusation The prediction component for the i-th category, N accusation is the total number of crime categories; loss time is the sentence loss function, is the true value of the annotation of the i-th category in the sentence, For S time The prediction component for the i-th category, N time is the total number of sentence categories, For S law The predicted component for the i-th category in .
[0019] Further, B j is the dimension reduction matrix, A j is the dimension-raising matrix, r is the rank of the matrix, and α is the hyperparameter.
[0020] Furthermore, the criminal charge multi-classification prediction module is Predict the output result S of the crime classification multilayer perceptron accusation ;Linear accusation1 and Linear accusation2 It is the linear calculation unit in the criminal charge classification prediction module, and ReLU is the linear rectification function activation function.
[0021] Furthermore, the criminal law and judicial interpretation multi-label prediction module is Predict the output results S of the multi-layer perceptron of criminal law and judicial interpretation law ;Linear law1 and Linear law2 It is the linear calculation unit in the multi-label prediction module of criminal law provisions and judicial interpretations.
[0022] Furthermore, the sentence multi-classification prediction module is The predicted sentence classification multilayer perceptron output result S time ;Linear time1 and Linear time2 It is the linear calculation unit in the sentence classification prediction module.
[0023] Furthermore, the encoder module is a pre-trained language model.
[0024] Furthermore, the criminal case dataset includes copyright criminal cases, trademark criminal cases, and patent criminal cases.
[0025] A server, characterized in that it includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the above method.
[0026] A computer-readable storage medium stores a computer program thereon, wherein the computer program implements the above method when executed by a processor.
[0027] The advantages of the present invention are as follows:
[0028] This method adopts a prediction model for intellectual property criminal liability, applicable laws and sentences based on a pre-training model and a hierarchical heterogeneous low-rank adaptive method. It can predict criminal charges, applicable criminal law provisions, judicial interpretations, and sentences in intellectual property criminal cases. It can effectively utilize existing judgment prediction models as pre-training models to reduce overall training costs.
[0029] At the same time, the application of a hierarchical heterogeneous low-rank matrix design further reduces the overall model parameter count while ensuring predictive performance. This approach can beneficially enhance the transfer learning and task adaptability of pre-trained models in multi-layer models with limited parameter increments. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a flow chart of the method of the present invention. DETAILED DESCRIPTION
[0031] The present invention will be described in further detail below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.
[0032] This paper addresses the challenges of predicting adjudication in intellectual property criminal cases, including the small number of training samples, the high time and cost of training models from scratch, the uniform dimensions of multi-level low-rank adaptive matrices, and the overly simple design structure that prevents the full application potential of the model structure. This paper proposes a method for predicting intellectual property criminal liability, applicable laws, and prison sentences based on a pre-trained model and a hierarchically heterogeneous low-rank adaptive method. By using an already trained model as a pre-trained model and designing a hierarchically heterogeneous low-rank matrix, a multi-task model capable of predicting adjudication in intellectual property criminal cases is constructed.
[0033] This proposal covers the following key aspects: First, the main process of constructing a dataset of intellectual property infringement cases, including copyright, trademark, and patent data, defining two prediction subtasks: infringement behavior and applicable laws and regulations. Furthermore, the project completes the data annotation of the six criminal offenses and 70 applicable laws and regulations involved in each case in the dataset. Second, it discusses the design of the model's main structure and the formulas involved. While focusing on the model training process, it also covers the category judgment process during the model prediction process.
[0034] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present invention. Figure 1 Flowchart of the technical method implemented by the present invention, which includes the following steps:
[0035] (1) Processing the intellectual property criminal case dataset, constructing a dataset that includes copyrights, trademarks, and patents, defining three prediction subtasks: the first prediction subtask is the crime, the second prediction subtask is the applicable criminal law and judicial interpretation, and the third prediction subtask is the sentence. The dataset also completes the data annotation of the 6 intellectual property crimes, 70 applicable laws and regulations, and 68 types of sentences involved in various cases in the dataset; that is, annotating the crime, applicable criminal law and judicial interpretation, and sentence of each case in the dataset, thus obtaining a sample dataset;
[0036] (2) Construct a judgment prediction model based on a pre-trained model and a hierarchically heterogeneous low-rank adaptive method, so that the judgment prediction model has the ability to predict criminal charges, applicable criminal law provisions and legal interpretations, and sentence length;
[0037] (3) Define the combined loss function of the referee prediction model and specify different training target loss functions for each subtask. After iterative training on the training set, a runnable model is formed.
[0038] (4) The judgment prediction model uses the input case fact description information to output three prediction results: the predicted criminal charge, the applicable criminal law provisions and judicial interpretations, and the sentence.
[0039] Step 1 includes the following sub-steps:
[0040] Step 1.1: Collect and organize the public copyright criminal cases, trademark criminal cases, and patent criminal cases on the Judgment Documents Network, and mark all cases with the crimes involved, applicable laws and judicial interpretations, and sentences. The dataset is composed as shown in formula (1). criminal is a criminal case dataset, S copyright is a copyright infringement case dataset, S trademarkis a trademark infringement case dataset, S patent It is a dataset of patent infringement cases.
[0041] S criminal =S copyright ∪S trademark ∪S patent (1)
[0042] Step 1.2: Filter the non-IP criminal case data within the case data to select data samples that can be used for data training and testing. Extract and annotate the criminal charges determined in the verdicts, as well as the criminal law provisions and legal interpretation clauses, and encode them. Applicable laws include the Criminal Law of the People's Republic of China, the Interpretation of the Supreme People's Court and the Supreme People's Procuratorate on Several Issues Concerning the Specific Application of Laws in Handling Intellectual Property Criminal Cases (I), the Interpretation of the Supreme People's Court and the Supreme People's Procuratorate on Several Issues Concerning the Specific Application of Laws in Handling Intellectual Property Criminal Cases (II), and the Interpretation of the Supreme People's Court and the Supreme People's Procuratorate on Several Issues Concerning the Specific Application of Laws in Handling Intellectual Property Criminal Cases (III).
[0043] Step 1.3: Define the task of predicting intellectual property criminal liability judgments into three subtasks: criminal charge prediction, applicable criminal law provisions and judicial interpretations, and sentence prediction. cirminal Defined as the overall task of criminal responsibility prediction, Task accusation For the criminal charge prediction subtask, Task article For the application of criminal law provisions and legal interpretation prediction subtasks, Task imprisonment For the sentence prediction subtask.
[0044] Task cirminal =Task accusation +Task article +Task imprisonment (3)
[0045] And the classification task is modeled as:
[0046]
[0047] Where F is the prediction model, and the case description text sequence is defined as w1,w1,...,w n As the input of the model, the model calculates the predicted output results of the two subtasks, namely the predicted results of criminal charges. Prediction results of applicable criminal law provisions and legal interpretations and the predicted results of sentence length
[0048] Step 1.3: Use manual labeling to label the criminal offense, applicable criminal law provisions and legal interpretations, and sentence.
[0049] Step 2 includes the following sub-steps:
[0050] Step 2.1: Define the overall structure of the judgment prediction model. The judgment prediction model primarily consists of an encoder module, a pretrained model, a multi-classification prediction module for criminal offenses, a multi-label prediction module for criminal law provisions and judicial interpretations, and a multi-classification prediction module for sentence lengths. In this invention, a pretrained language model is used as the encoder module.
[0051] Step 2.2, define the input and data annotation labels of the judgment prediction model. The case description information of the intellectual property criminal case (including the criminal behavior and evidence in the plaintiff's indictment, the explanation of the behavior in the defense, and the "this court believes" part of the judgment) is used as the input of the judgment prediction model. The input W can be defined as W = {w1,w2,w3,..,w n}, where w1,w2,w3,..,w n It is a phrase after word segmentation, and its maximum length is defined as n.
[0052] When the number of phrases is less than n, fill with blank characters to align the input information. Define the marked crime as The quantity is i, and the superscript a is used to identify the crime label.
[0053] All applicable laws and judicial interpretations marked in the case are defined as a text label group of k The above superscript e is the label of the legal provisions and judicial interpretations based on which the case was tried. Since the application of legal provisions and judicial interpretations is a multi-label classification task, the applicable legal provisions and judicial interpretations marked in each case can be defined as This array is L e Where m is the number of applicable laws and judicial interpretations. If the number of applicable laws and judicial interpretations marked in the marked case is less than m, the value of "-1" is used to supplement it, so that the label of all cases in the dataset is e The length of each is m. All sentences marked in the case are defined as a classification label group with a number of h n is the number of sentence categories.
[0054] Step 2.3, define the encoder module and its input and output. To improve the training efficiency of the judgment prediction model and its ability to capture key semantic information, the semantic information of the label is incorporated into the input data samples during training. The case description text, crime charges, criminal law articles, and judicial interpretation text are encoded into text vectors through the pre-trained language model. The case description text vector is defined as X doc_em∈R k×emb , the criminal charge text vector is defined as X accusation_em ∈R k ×emb , legal provisions and judicial interpretation text vectors are defined as X law_em ∈R k×emb , where k is the number of words contained in the text after word segmentation, and emb is the word embedding space dimension. The pre-trained language model is defined as f pt , the above process is defined by the following formula:
[0055]
[0056] In order to make the model converge more easily during training, the above text encoding is normalized layer by layer, where f LN is the layer normalization function, X doc 、X accusation 、X law are the normalized case description text vector, criminal charge text vector, and legal provisions and judicial interpretation text vector:
[0057]
[0058] Step 2.4: To solve the problem that the small number of training samples for intellectual property criminal cases can easily lead to slow model convergence and weak generalization prediction ability of the trained model, the present invention adopts a transfer learning method that uses existing models trained in other fields or cases as pre-training models to speed up the training of this model and improve the robustness of the model. The pre-training model is defined as model pre-train , and X doc 、X accusation 、X law Three text vectors as models pre-train The input of the Transformer architecture is used for selection and use in the training and prediction process, which also makes the range of models that can be selected by the present invention wider. In current research, the encoder or decoder multi-level model of the Transformer architecture is generally used to convert the model pre-train Defined as:
[0059] model pre-train ={layer1,....,...,layer n} (7)
[0060] Among them, layer n Represented as model pre-train The nth level in . During training and testing, the calculation process of each level can be expressed as:
[0061]
[0062] where logits1, logits2, logits n Calculate results for each layer.
[0063] Step 2.5: Starting from the pre-trained multi-level structure, add a low-rank matrix calculation layer ΔW to each level i , where i is the number of the level. ΔW i The calculation formula is as follows:
[0064]
[0065] Among them, the dimension reduction matrix B i ∈R k×r , the dimension-raising matrix A i ∈R r×emb , r is the rank of the set matrix, α is a hyperparameter that is set to the same as r by default during training and can be adjusted according to the actual training situation. In order to reduce the overall parameter amount of the model, r is set to different values at different levels, r = {8, 16, 32, 64}, and increases from low level to high level. At initialization, set B i =0, A i =N(0,σ 2 ), that is, the matrix B corresponding to the i-th level i The initial value is set to 0, and the initial value does not increase with the level change. i The mean is 0 and the variance is σ 2 The initial value is set by the standard Gaussian distribution of B. i With A i Both are trainable parameter matrices, which are updated according to the returned gradients during training.
[0066] X doc 、X accusation 、X law The three text vectors are linearly calculated at the ratio of 0.5, 0.3, and 0.2 to form the input X of the first low-rank matrix calculation layer. LoRA .
[0067] X LoRA =0.5X doc +0.3X accusation +0.2X law (10)
[0068] In training and testing, the calculation process of combining the pre-trained model layer and the low-rank matrix calculation layer is upgraded from formula (8) to formula (11), where logits1, logits2, logits n The calculation results of each layer are:
[0069]
[0070] Step 2.6, define the criminal offense multi-classification prediction module. As shown in formula (12), Linear accusation1 and Linear accusation2 It is the linear calculation unit in the criminal charge classification prediction module, and ReLU is the linear rectification function activation function. These two formula symbols will also be used in the subsequent steps and will not be explained again. accusation-linear1 is the output result of the first linear calculation unit in the crime classification multilayer perceptron module, S accusation-relu To output the result after activation operation, S accusation Output of the multilayer perceptron for crime classification:
[0071]
[0072] Step 2.7, define the multi-label prediction module of criminal law articles and judicial interpretations. As shown in formula (13), Linear law1 and Linear law2 It is a linear calculation unit in the multi-label prediction module of criminal law articles and judicial interpretations. law-linear1 is the output result of the first linear calculation unit in the multi-label prediction module of criminal law articles and judicial interpretations, S law-relu To output the result after activation operation, S law Output of the multilayer perceptron for crime classification:
[0073]
[0074] Step 2.7, define the sentence multi-classification prediction module. As shown in formula (14), Linear time1 and Linear time2 It is the linear calculation unit in the sentence classification prediction module. time-linear1 is the output result of the first linear calculation unit in the sentence classification prediction module, S time-relu To output the result after activation operation, S time Output of the multilayer perceptron for crime classification:
[0075]
[0076] Step 3 includes the following sub-steps:
[0077] Step 3.1, define the loss function of each subtask in the model and the overall loss function. The criminal charge prediction subtask and the sentence prediction subtask use the cross entropy loss function, loss accusation is the criminal charge loss function, is the true value of the criminal offense, where i is the component value of the i-th category, For S accusation The prediction component for the i-th category, N accusation is the total number of crime categories. time is the sentence loss function, is the true value of the annotation of the i-th category in the sentence, For S time The predicted component for the jth category, N time is the total number of sentence categories.
[0078]
[0079] The subtask of applying criminal law provisions and judicial interpretations is calculated using the following loss function, where loss law To apply the criminal law and judicial interpretation loss function, For S law The prediction component for the i-th category, Ω neg is the set of non-applicable laws in the prediction results, Ω pos It is the set of laws applicable to the prediction results.
[0080]
[0081] The model uses a combined loss function that includes three sub-task loss functions.
[0082]
[0083] Where λ, β and γ are the linear combination coefficients in the combined loss function.
[0084] Step 3.3: Based on the algorithm model and loss function defined in the above steps, apply the sample data set formed in step 1 to carry out training.
[0085] Step 4 includes the following sub-steps:
[0086] Step 4.1: Define the prediction process of the model. After inputting the case description information string into the model, the model outputs the criminal charge and sentence according to the well-known function Softmax according to the output of formula (12), (13), and (14), and determines the maximum classification result number as shown in formula (18). select The classification prediction result number of the crime, time select Number the sentence classification prediction results:
[0087]
[0088] Regarding the application of criminal law and judicial interpretation, by judging S lawThe i-th component in is greater than 0. If it is greater than zero, the category represented by the component is selected, otherwise it is not selected. As shown in formula (19). Predict classification numbers for criminal law articles and judicial interpretations.
[0089]
[0090] Step 4.2: Based on the accusation output in step 4.1 select 、time select as well as They are converted into the names of crimes, sentences, criminal law articles and judicial interpretations they represent.
[0091] While specific embodiments of the present invention have been disclosed for illustrative purposes, intended to facilitate understanding and implementation of the present invention, those skilled in the art will appreciate that various substitutions, variations, and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the disclosure of the preferred embodiments, and the scope of protection claimed in the present invention shall be determined by the scope of the claims.
Claims
1. A method for predicting intellectual property criminal liability, applicable laws and regulations, and sentence length, comprising the following steps: 1) For each case in the criminal case dataset, label the crime, applicable criminal law and judicial interpretation, and sentence to obtain a training sample; 2) Constructing a judgment prediction model, using the training samples to train the judgment prediction model for predicting the crime, applicable criminal law provisions and legal interpretations, and sentence in a criminal case; wherein the judgment prediction model includes an encoder module, a pre-trained model, a criminal crime multi-classification prediction module, a criminal law provision and judicial interpretation multi-label prediction module, and a sentence multi-classification prediction module; adding a low-rank matrix calculation layer to each level of a multi-level encoder or decoder model based on a Transformer architecture to obtain the pre-trained model; the method for training the judgment prediction model is: 21) The encoder module encodes the input training sample to obtain the case description text vector X of the training sample doc_em , crime text vector X accusation_em , applicable criminal law articles and judicial interpretation text vector X law_em And normalize them respectively to get X doc 、X accusation 、X law ; Among them, X doc is the normalized case description text vector, X accusation is the normalized criminal charge text vector, X law is the normalized legal article and judicial interpretation text vector; k is the number of words contained in the text after word segmentation, and emb is the word embedding space dimension; 22) The pre-trained model is based on the input X doc 、X accusation 、X law , through the formula Calculate logits n ; Among them, for X doc 、X accusation 、X law Do linear calculation to get vector X LoRA ; ΔW j The low-rank matrix calculation layer added for the jth level, j = 1 to n, n is the total number of levels of the pre-trained model; 23) The criminal offense multi-classification prediction module is based on logits n Predict the output result S of the crime classification multilayer perceptron accusation The multi-label prediction module of criminal law articles and judicial interpretations is based on logits n Predict the output results S of the multi-layer perceptron of criminal law and judicial interpretation law The sentence multi-classification prediction module is based on logits n The predicted sentence classification multilayer perceptron output result S time ; Then, the loss value is calculated based on the prediction results and the annotation information to optimize the referee prediction model; 3) For a criminal case to be predicted, the case facts are input into the trained judgment prediction model to obtain the crime, applicable criminal law provisions and judicial interpretations, and sentence of the criminal case to be predicted.
2. The method according to claim 1, characterized in that Through the loss function L = λloss accusation +βloss law +γloss time The loss value L is calculated; where λ, β, and γ are linear combination coefficients, λ+β+γ=1; Ω neg is the set of non-applicable laws in the prediction results, Ω pos is the set of applicable laws in the prediction results, is the true value of the i-th category in the criminal charges, For S accusation The prediction component for the i-th category, N accusation is the total number of crime categories; loss time is the sentence loss function, is the true value of the annotation of the i-th category in the sentence, For S time The prediction component for the i-th category, N time is the total number of sentence categories, For S law The predicted component for the i-th category in .
3. The method according to claim 1, characterized in that B j is the dimension reduction matrix, A j is the dimension-raising matrix, r is the rank of the matrix, and α is the hyperparameter.
4. The method according to claim 1, 2 or 3, characterized in that: The criminal offense multi-classification prediction module is Predict the output result S of the crime classification multilayer perceptron accusation ;Linear accusation1 and Linear accusation2 It is the linear calculation unit in the criminal charge classification prediction module, and ReLU is the linear rectification function activation function.
5. The method according to claim 1, 2 or 3, characterized in that: The criminal law and judicial interpretation multi-label prediction module is Predict the output results S of the multi-layer perceptron of criminal law and judicial interpretation law ;Linear law1 and Linear law2 It is the linear calculation unit in the multi-label prediction module of criminal law provisions and judicial interpretations.
6. The method according to claim 1, 2 or 3, characterized in that: The sentence multi-classification prediction module is The predicted sentence classification multilayer perceptron output result S time ;Linear time1 and Linear time2 It is the linear calculation unit in the sentence classification prediction module.
7. The method according to claim 1, characterized in that The encoder module is a pre-trained language model.
8. The method according to claim 1, characterized in that The criminal case dataset includes copyright criminal cases, trademark criminal cases, and patent criminal cases.
9. A server, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.