Judicial case progress dynamic prediction method based on time sequence prediction model
Patent Information
- Application Number
- CN202510427542.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
Smart Images

Figure CN120336714A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of time series prediction models, and in particular to a dynamic prediction method for the progress of judicial cases based on a time series prediction model. Background Art
[0002] In the practice of judicial case trial and management, the uncertainty of the case flow progress has always been a key factor affecting judicial efficiency and resource allocation. With the development of social economy and the improvement of the legal awareness of the people, the number of cases accepted by judicial organs has been increasing, and the types, complexity levels and process nodes of cases also show a highly diversified trend. Under this background, how to scientifically and accurately predict the progress of judicial cases at different process nodes has become an important topic for improving the level of judicial administration, rationally allocating trial resources and enhancing public satisfaction.
[0003] In the prior art, the prediction of the progress of judicial cases mainly relies on rule-driven models or traditional statistical methods, such as linear regression, support vector regression, time series analysis, etc. These methods usually assume that the process of case progress has a linear or stationary modeling trend. However, in reality, judicial cases are often affected by various factors such as case types, court levels, regional differences, parties' behaviors, and judge scheduling, resulting in complex non-linear dynamic changes of cases at different time nodes. It is difficult for models based on linear assumptions or fixed structures to effectively capture these dynamic characteristics, and there are significant limitations in the accuracy and stability of prediction results.
[0004] In addition, some emerging artificial intelligence methods have gradually been introduced into the judicial field to assist in tasks such as case text classification, legal provision matching, and judgment result prediction. However, applying deep learning technology to case progress time prediction still faces many challenges. On the one hand, the progress of judicial cases has significant time series characteristics, and there are strong dependencies between different stages. It is difficult for traditional recurrent neural networks or one-dimensional convolutional networks to take into account both long-term dependencies and feature importance modeling. On the other hand, existing models often ignore the structured correlation information between different features in the process of case progress. For example, the interaction relationship between different variables in the case trial link is difficult to describe with linear or independent assumptions.
[0005] In response to the above problems, some studies have tried to introduce graph neural networks to model the structural relationships in judicial data. However, most methods only model static graphs or structural data, and it is difficult to meet the joint modeling requirements of time series dynamic characteristics and graph structures. At the same time, in terms of hyperparameter optimization, existing studies mostly adopt general methods such as grid search or Bayesian optimization. Although they have a certain search efficiency, they lack the modeling and control of the uncertainty and stage characteristics unique to the judicial scenario, resulting in the final model performance being easily affected by local optima and it being difficult to achieve stable generalization in complex judicial environments.
[0006] In addition, although the current partial research has introduced the attention mechanism to enhance the model's attention ability to key time points, most of them are based on single-scale time series modeling, ignoring the multi-time scale information existing in the case progress. For example, in a long-term criminal case, the change patterns in different time periods may be completely different, and the existing attention mechanism is difficult to effectively fuse among multiple time scales, restricting the improvement of prediction accuracy.
[0007] To sum up, the existing judicial case progress prediction technologies generally have the following defects: First, there is a lack of a joint modeling mechanism for dynamic time series features and the structural relationship between features, making it difficult to comprehensively extract the potential associations and dependence patterns in complex case information; Second, the robustness in the parameter training and optimization process is insufficient, the model is sensitive to perturbations and fluctuations, and there is a lack of robustness control; Third, the existing methods fail to combine the multi-stage characteristics and uncertainty factors of case progress for hyperparameter adjustment and dynamic optimization, making it difficult to adapt to different case types and process stages; Fourth, the model structure lacks a multi-scale time dependence information fusion mechanism, restricting the expression ability for complex progress patterns.
[0008] Based on the above technical background and existing deficiencies, it is urgent to propose a new method that combines graph structure modeling, time series feature encoding, adaptive attention mechanism and optimization strategy fusion, which can not only accurately model the multi-dimensional complex time series features in the process of judicial case progress, but also improve the model's adaptability to uncertainty and sample changes, so as to achieve dynamic and high-precision prediction of the time nodes in each stage of judicial cases, and help the judicial organs to achieve scientific scheduling, intelligent early warning and auxiliary decision-making.
[0009] Therefore, how to provide a dynamic prediction method for judicial case progress based on a time series prediction model is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0010] An object of the present invention is to propose a dynamic prediction method for judicial case progress based on a time series prediction model. The present invention comprehensively introduces an edge condition adaptive graph convolution algorithm, a multi-scale time series encoding network, a hierarchical adaptive cross-attention fusion mechanism, a sensitivity perception minimization algorithm, and an improved simulated annealing hyperparameter optimization strategy, and details the whole process of structurally modeling, time dynamic analysis and stage prediction of historical data of judicial cases, and has the advantages of high prediction accuracy, strong model generalization ability, and good adaptability to different case types and stages.
[0011] A dynamic prediction method for judicial case progress based on a time series prediction model according to an embodiment of the present invention includes the following steps:
[0012] A dynamic prediction method for the progress of judicial cases based on a time series prediction model, characterized by including the following steps:
[0013] S1. Preprocess the historical data of judicial cases, including data missing filling and normalization processing, to generate static features and dynamic time series features;
[0014] S2. Encode the normalized static features to generate static hidden vectors;
[0015] S3. Construct a dynamic graph for the normalized dynamic time series features, use the edge-conditioned adaptive graph convolution algorithm to perform feature fusion on each node in the dynamic graph, generate a fusion feature matrix, and use a multi-scale time series encoding network based on dilated convolution and residual connection to perform time series encoding on the fusion feature matrix to obtain a time series hidden state;
[0016] S4. Use a hierarchical adaptive cross-attention fusion module for the time series hidden state to perform local interaction and weighted fusion on the hidden states of different time scales to form a context vector;
[0017] S5. Concatenate the static hidden vector and the context vector, and generate a prediction output through a fully connected layer;
[0018] S6. Calculate the loss between the prediction output and the target label, and use the sensitivity-aware minimization algorithm to optimize the model parameters;
[0019] S7. During the model training process, use an improved simulated annealing algorithm to adjust and optimize the hyperparameters;
[0020] S8. Apply the trained model to new judicial case data and output the time prediction values for each stage.
[0021] Optionally, the specific steps of S2 include:
[0022] S21. Set the dimension of the static feature vector to n, and denote the static feature vector as S ′ ;
[0023] S22. Define the static hidden vector as:
[0024] Z S = W S ·S ′ + b S ;
[0025] Where W S is the encoding weight matrix, b S is the bias term, m is the dimension of the static hidden vector, and Z S is the output static hidden vector.
[0026] Optionally, S3 specifically includes:
[0027] S31. Construct the normalized dynamic time series features into a time series matrix X = [X1, X2, …, X T , where T is the number of time steps, X t represents the dynamic feature vector at the t-th time step, and d is the feature dimension of each time step;
[0028] S32. Use each dynamic feature vector X t as the node set V t = {x t,1 , x t,2 , …, x t,d} to construct a graph structure G t = (V t , E t ), where each edge e t in the edge set E t,ij corresponds to the dependency relationship between the feature nodes x t,i and x t,j ;
[0029] S33. Apply the edge-conditioned adaptive graph convolution algorithm to the graph structure G t for feature fusion, including through the edge mapping function:
[0030] e t,ij = ReLU(W e [x t,i ‖ x t,j + b e );
[0031] Generate the edge feature e t,ij , where W e , b e , [x t,i ‖ x t,j represents node concatenation, ReLU represents the activation function, and input the edge feature into the edge parameter generation function Φ(·) to obtain a learnable convolution kernel:
[0032] W t,ij = Φ(e t,ij );
[0033] Based on the adjacency relationship, perform weighted fusion on the adjacent nodes of the node x t,i to obtain a fused feature vector:
[0034]
[0035] Among them, represents the adjacency set of the node x t,j ;
[0036] Combine all the fused node vectors to obtain the fused feature matrix at time step t:
[0037]
[0038] S34. Input the time series fused feature into the multi-scale time series coding network based on dilated convolution and residual connection, and use causal convolution kernels under multiple dilation factor combinations to perform multi-scale convolutional coding to obtain the time series hidden state sequence H = [H1, H2, …, H T , where H T represents the hidden state at the t-th time step.
[0039] Optionally, the S4 specifically includes:
[0040] S41. Use the time series hidden state sequence H = [H1, H2, …, H T as the input and input it into the hierarchical adaptive cross-attention fusion module;
[0041] S42. Downsample the time series hidden state sequence H according to the preset multi-time scale set to obtain the multi-scale hidden state subsequence set where each subsequence represents the sequence length at this scale;
[0042] S43. For any two different scale subsequences and where s i , s i ≠s j , perform cross-attention fusion calculation, and the fused representation of the hidden state corresponding to the t-th time step is defined as:
[0043]
[0044] where the attention weight coefficient α t,l is calculated by the following formula:
[0045]
[0046] where W q , W k , are learnable query, key, and value transformation matrices respectively, d a is the attention hidden space dimension, and exp represents the exponential function;
[0047] S44. For all scales Attention output Perform upsampling to restore to the original time step length T, perform weighted fusion, and obtain a unified context representation sequence C = [C1, C2, …, C T , where each context vector is:
[0048]
[0049] Among them, the fusion weight β s Satisfy β s ≥0;
[0050] S45. Use the obtained context vector sequence C = [C1, C2, …, C T for concatenation with the static hidden vector and enter the subsequent prediction output generation step.
[0051] Optionally, the S5 specifically includes: Concatenate the static hidden vector Z S generated in step S2 with the context vector sequence C = [C1, C2, …, C T at each time step t to obtain a fusion vector F t = [Z S ‖ C t , and input the fusion vector F t into the prediction output generation network to generate a prediction output through a fully connected transformation where W o is the output weight matrix, b o is the bias term, is the time prediction result at time step t, and finally obtain a prediction output sequence
[0052] Optionally, the S6 specifically includes:
[0053] S61. Optimize the model parameter set θ based on the sensitivity-aware minimization algorithm. The parameter set θ includes: the convolution kernel parameters in the multi-scale time series encoding network, the query matrix W q , key matrix W k , value matrix W v , the weight matrix W o of the prediction output fully connected layer, and the bias term b o . When the initial model parameter set is θ, the prediction output sequence is denoted as
[0054] S62. Concatenate the prediction output sequence with the corresponding true target label sequence Y = [Y1, Y2, …, YT Perform loss calculation and define the basic loss function as the mean square error function:
[0055]
[0056] Among them, represents the predicted output sequence the predicted output at time step t in, Y t the output at time step t in the true target label sequence Y, T represents the total number of time steps of the prediction sequence;
[0057] S63. For the current model parameter θ, calculate the local perturbation direction ∈ * , and find the local steepest point of the loss function in the parameter space. The calculation formula is as follows:
[0058]
[0059] Among them, argmax represents the function to take the maximum value of the function, ρ is the preset perturbation radius, ‖·‖2 represents the Euclidean norm constraint, and ∈ represents the perturbation vector with the same dimension as the model parameter θ;
[0060] S64. Substitute the perturbation direction into the gradient of the loss function and execute the sensitivity-aware optimization strategy to update the model parameters. The update formula is as follows:
[0061]
[0062] Among them, represents the gradient of the loss function L base with respect to the original parameter θ, and η is the learning rate.
[0063] Optionally, the S7 specifically includes:
[0064] S71. Set the initial hyperparameter vector λ (0) =[η (0) ,b (0) ,l (0) , where η represents the learning rate, b represents the batch size, l represents the number of network layers, and set the initial temperature T0, temperature decay factor k>0, maximum number of iterations N, and minimum temperature threshold T min ;
[0065] S72. In the i-th iteration, use the current hyperparameter vector λ (i) to construct a model, and after training and validation, obtain the loss value L (i) and the variance of the predicted output sequence
[0066] S73. According to the progress distribution density function corresponding to the current prediction stage of the judicial case Introduce stage perception regulation weights for the perturbation components to generate a stage perception perturbation vector Δλ (i) , which is defined as:
[0067]
[0068] where, γ is the perturbation ratio coefficient, Σ is the covariance matrix, which can be obtained from the statistical distribution of the case stage time;
[0069] S74. Generate a candidate hyperparameter vector λ (i+1) = λ (i) + Δλ (i) , and retrain the model accordingly to obtain a new loss value L (i+1) and the predicted output variance
[0070] S75. Define an improved acceptance probability function P (i) as follows:
[0071]
[0072] where α is the uncertainty regulation factor, T i = T0·e -k·i is the temperature in the i-th round, and exp is the exponential function;
[0073] S76. According to the above acceptance probability P (i) , accept or reject λ (i+1) as the hyperparameter vector for the next round of iteration in a probabilistic manner;
[0074] S77. Repeat steps S72 to S76 until the stopping condition is met: the number of iteration rounds i ≥ N or the current temperature T i ≤ T min , and finally select the hyperparameter vector λ (*) corresponding to the lowest loss value L * in the acceptance history as the optimal hyperparameter configuration of the final trained model.
[0075] The beneficial effects of the present invention are:
[0076] (1) By introducing the edge condition adaptive graph convolution algorithm, the present invention can dynamically construct the graph structure relationship between case features at each time step, and combine the multi-scale dilation residual convolution network to deeply encode the temporal information, thereby effectively capturing the long-term dependencies and complex non-linear dynamic changes existing in the progress of judicial cases, and significantly improving the prediction accuracy of the arrival time of case process nodes.
[0077] (2) The present invention optimizes the loss function by introducing a sensitivity perception minimization algorithm, and not only considers the loss value itself during the model training process, but also evaluates its sensitivity to parameter perturbations, enabling the model to still have stable prediction performance when facing different samples, different case types, and process stages, effectively improving the anti-perturbation ability and generalization performance of the model.
[0078] (3) The present invention constructs an improved simulated annealing algorithm, adds a prediction uncertainty control term to the acceptance criterion, and combines the judicial case stage distribution to guide the perturbation direction to dynamically search for the optimal hyperparameter combination, making the model more in line with the process characteristics of judicial cases, avoiding falling into local optima, and improving the practicality and flexibility of deployment in the actual judicial environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0080] Figure 1 is the overall flowchart of a dynamic prediction method for judicial case progress based on a time series prediction model proposed by the present invention;
[0081] Figure 2 is the structural schematic diagram of the judicial case data preprocessing module in the present invention;
[0082] Figure 3 is the flowchart of dynamic graph construction and edge condition adaptive graph convolution processing in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0083] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0084] Refer to Figures 1-3 , a dynamic prediction method for judicial case progress based on a time series prediction model, comprising the following steps:
[0085] S1. Preprocess the historical data of judicial cases, including data missing filling and normalization processing, to generate static features and dynamic time series features;
[0086] S2. Encode the normalized static features to generate static hidden vectors;
[0087] S3. Construct a dynamic graph for the normalized dynamic time series features, use the edge condition adaptive graph convolution algorithm to perform feature fusion on each node in the dynamic graph, generate a fused feature matrix, and use a multi-scale time series encoding network based on dilated convolution and residual connection to perform time series encoding on the fused feature matrix to obtain a time series hidden state;
[0088] S4. Use a hierarchical adaptive cross-attention fusion module for the time series hidden state to perform local interaction and weighted fusion on the hidden states at different time scales to form a context vector;
[0089] S5. Concatenate the static hidden vector and the context vector, and generate a prediction output through a fully connected layer;
[0090] S6. Calculate the loss between the prediction output and the target label, and use the sensitivity-aware minimization algorithm to optimize the model parameters;
[0091] S7. During the model training process, use an improved simulated annealing algorithm to adjust and optimize the hyperparameters;
[0092] S8. Apply the trained model to new judicial case data and output the time prediction values for each stage.
[0093] In this embodiment, the specific steps of S2 include:
[0094] S21. Set the dimension of the static feature vector to n, and denote the static feature vector as S ′ ;
[0095] S22. Define the static hidden vector as:
[0096] Z S = W S ·S ′ + b S ;
[0097] where W S is the encoding weight matrix, b S is the bias term, m is the dimension of the static hidden vector, and Z S is the output static hidden vector.
[0098] In this embodiment, the specific steps of S3 include:
[0099] S31. Construct the normalized dynamic time series features into a time series matrix X = [X1, X2, …, X T , where T is the number of time steps, X t represents the dynamic feature vector at the t-th time step, and d is the feature dimension of each time step;
[0100] S32. For each dynamic feature vector X tFor the node set V t ={x t,1 , x t,2 , …, x t,d}, construct the graph structure G t =(V t , E t ), where each edge e t in the edge set E t,ij corresponds to the dependency relationship between the feature nodes x t,i and x t,j ;
[0101] S33. Apply the edge condition adaptive graph convolution algorithm to the graph structure G t for feature fusion, including through the edge mapping function:
[0102] e t,ij =ReLU(W e [x t,i ‖x t,j +b e );
[0103] Generate the edge feature e t,ij , where W e , b e , [x t,i ‖x t,j represents node concatenation, ReLU represents the activation function, and input the edge feature into the edge parameter generation function Φ(·) to obtain the learnable convolution kernel:
[0104] W t,ij =Φ(e t,ij );
[0105] Based on the adjacency relationship, perform weighted fusion on the adjacent nodes of the node x t,i to obtain the fused feature vector:
[0106]
[0107] where, represents the adjacency set of the node x t,j ;
[0108] Combine all the fused node vectors to obtain the fused feature matrix at time step t:
[0109]
[0110] S34. Input the time series fused feature into the multi-scale time series coding network based on dilated convolution and residual connection, and use the causal convolution kernels under multiple combinations of dilation factors to Perform multi-scale convolutional encoding to obtain a temporal hidden state sequence H = [H1, H2, …, H T , where H T represents the hidden state at the t-th time step.
[0111] In this embodiment, S4 specifically includes:
[0112] S41. Take the temporal hidden state sequence H = [H1, H2, …, H T as the input and input it into the hierarchical adaptive cross-attention fusion module;
[0113] S42. Downsample the temporal hidden state sequence H according to a preset multi-time scale set to obtain a multi-scale hidden state subsequence set where each subsequence represents the sequence length at this scale;
[0114] S43. For any two different-scale subsequences and where s i , s i ≠ s j , perform cross-attention fusion calculation, and the fused representation of the hidden state corresponding to the t-th time step is defined as:
[0115]
[0116] where the attention weight coefficient α t,l is calculated by the following formula:
[0117]
[0118] where W q , W k , are learnable query, key, and value transformation matrices respectively, d a is the attention hidden space dimension, and exp represents the exponential function;
[0119] S44. Upsample the attention outputs at all scales back to the original time step length T, perform weighted fusion, and obtain a unified context representation sequence C = [C1, C2, …, C T , where each context vector is:
[0120]
[0121] where the fusion weight βs Meet β s = 1, β s ≥ 0;
[0122] S45. Use the obtained context vector sequence C = [C1, C2, …, C T for concatenating with the static hidden vector and enter the subsequent prediction output generation step.
[0123] In this embodiment, S5 specifically includes: Concatenate the static hidden vector Z S generated in step S2 with the context vector sequence C = [C1, C2, …, C T at each time step t to obtain the fusion vector F t = [Z S ‖C t , and input the fusion vector F t into the prediction output generation network to generate the prediction output through a fully connected transformation where W o is the output weight matrix, b o is the bias term, is the time prediction result at time step t, and finally obtain the prediction output sequence
[0124] In this embodiment, S6 specifically includes:
[0125] S61. Optimize the model parameter set θ based on the sensitivity-aware minimization algorithm. The parameter set θ includes: the convolution kernel parameters in the multi-scale time series encoding network, the query matrix W q , key matrix W k , value matrix W v , the weight matrix W o of the prediction output fully connected layer, and the bias term b o . When the initial model parameter set is θ, the prediction output sequence is denoted as
[0126] S62. Calculate the loss between the prediction output sequence and the corresponding true target label sequence Y = [Y1, Y2, …, Y T . Define the basic loss function as the mean square error function:
[0127]
[0128] where, represents the prediction output at time step t in the prediction output sequence Yt The output at time step t in the true target label sequence Y, where T represents the total number of time steps in the prediction sequence;
[0129] S63. For the current model parameters θ, calculate the local perturbation direction ∈ * , and find the locally steepest point of the loss function in the parameter space. The calculation formula is as follows:
[0130]
[0131] where argmax represents the function to take the maximum value of the function, ρ is the preset perturbation radius, ‖·‖2 represents the Euclidean norm constraint, and ∈ represents the perturbation vector with the same dimension as the model parameters θ;
[0132] S64. Substitute the perturbation direction into the gradient of the loss function, and execute the sensitivity-aware optimization strategy to update the model parameters. The update formula is as follows:
[0133]
[0134] where represents taking the gradient of the loss function L base with respect to the original parameters θ, and η is the learning rate.
[0135] In this embodiment, the S7 specifically includes:
[0136] S71. Set the initial hyperparameter vector λ (0) = [η (0) , b (0) , l (0) , where η represents the learning rate, b represents the batch size, l represents the number of network layers, and set the initial temperature T0, the temperature decay factor k>0, the maximum number of iterations N, and the minimum temperature threshold T min ;
[0137] S72. In the i-th iteration, use the current hyperparameter vector λ (i) to construct a model. After training and validation, obtain the loss value L (i) and the variance of the predicted output sequence
[0138] S73. According to the progress distribution density function corresponding to the current prediction stage of the judicial case, introduce a stage-aware regulation weight for the perturbation component to generate a stage-aware perturbation vector Δλ (i) , which is defined as:
[0139]
[0140] where γ is the perturbation ratio coefficient, Σ is the covariance matrix, It can be obtained from the statistical analysis of the time distribution of case stages;
[0141] S74. Generate a candidate hyperparameter vector λ (i+1) = λ (i) + Δλ (i) , and retrain the model accordingly to obtain a new loss value L (i+1) and the prediction output variance
[0142] S75. Define an improved acceptance probability function P (i) as follows:
[0143]
[0144] where α is an uncertainty regulation factor, T i = T0·e -k·i is the temperature at the i-th round, and exp is the exponential function;
[0145] S76. According to the above acceptance probability P (i) , accept or reject λ (i+1) with a certain probability as the hyperparameter vector for the next round of iteration;
[0146] S77. Repeat steps S72 to S76 until the stopping condition is met: the number of iteration rounds i ≥ N or the current temperature T i ≤ T min . Finally, select the hyperparameter vector λ (*) corresponding to the lowest loss value L * in the acceptance history as the optimal hyperparameter configuration for the final trained model.
[0147] Example 1:
[0148] To verify the feasibility of the present invention in implementation, the present invention is applied to the construction of a criminal case trial process prediction system of a provincial high people's court in 2023. Experimental verification is carried out on the criminal case data accepted and closed by the court during the period from 2022 to 2023 to evaluate the application effect and performance of the "dynamic prediction method for judicial case progress based on a time series prediction model" proposed by the present invention in actual business.
[0149] In the past, when courts arranged case trial plans and allocated trial forces, they generally predicted case time by using empirical rules plus historical average time estimates. However, due to the different complexities of cases themselves and individual differences in case process rhythms, the accuracy of this empirical method in time prediction is only about 58%, and problems such as node expiration delays, scheduling overlaps, and resource conflicts often occur, seriously affecting the judicial operation efficiency and public satisfaction.
[0150] To solve this problem, the present invention combines the case trial log data within the court (including case type, filing time, actual completion time of each stage, presiding judge, number of involved parties, trial method, etc.) with structured auxiliary data (such as court schedule, case document processing time, parties' attendance, etc.) to construct static features and dynamic time series features. Through cleaning, standardization, and model training on 18,094 criminal case data that were actually closed during the period from January 2022 to June 2023, and validation on the data of newly accepted cases to be predicted during the period from July 2023 to December 2023.
[0151] In the model implementation, the system first performs missing value filling and normalization on the original data. The extracted static features include case category encoding, applicable procedure, trial level, case complexity label, whether it is a public trial, judge's working years, number of people in the collegial panel, etc., totaling 14 dimensions; the dynamic time series features include the intermediate time nodes and their historical cumulative durations at 7 stages such as the time from filing to court scheduling, the time from court scheduling to the first hearing, the time from the hearing to the completion of cross-examination, the time from the end of the trial to deliberation, the time from deliberation to the formation of the judgment, and the time from the judgment to the generation of the judgment document. These features construct the graph structure between nodes through the dynamic graph modeling method proposed by the present invention, use the edge-conditioned adaptive graph convolutional network for feature fusion, encode the time series features through multi-scale dilated residual convolution, and then fuse the case dynamic trends at different time scales through the cross-attention mechanism, and finally output the predicted values for the completion time of each stage.
[0152] The model training uses the sensitivity-aware minimization algorithm to optimize the parameters, effectively improving the generalization ability of the model among different case samples. At the same time, by introducing an improved simulated annealing algorithm based on the judicial stage distribution to dynamically adjust hyperparameters such as the learning rate and the number of layers, the convergence performance of the model is further optimized. In the model testing stage, compared with traditional time series prediction models such as LSTM, GRU, TCN, and XGBoost, the average prediction error of the present invention is significantly reduced, and the adaptability to different types of cases is excellent.
[0153] The following is the performance of the model of the present invention on the test set, an evaluation data table for comparison with the current method:
[0154] Table 1: Comparison table of prediction accuracy (MAE) and mean absolute percentage error (MAPE) of the method of the present invention and existing models in different case types
[0155]
[0156] It can be seen from the above comparison table that the dynamic prediction method for judicial case progress based on the time series prediction model proposed in the present invention has a significant performance advantage over the existing mainstream methods in the prediction tasks of various criminal cases. By comparing with common time series prediction models such as LSTM, TCN and XGBoost, the model of the present invention has extremely strong stability and accuracy in error control.
[0157] First, from the perspective of mean absolute error (MAE), the invention controls all case types within 5 days, and most types remain at around 3 days. The lowest error is only 2.65 days in the "obstruction of public service" case, and it is also stable at around 3 days in high-frequency case types. This shows that the model of the invention can effectively learn the key node features in the progress of the case and has a strong time series capture capability.
[0158] Secondly, from the perspective of mean absolute percentage error (MAPE), the MAPE of the present invention is controlled within 8% in all six typical case types, with an average of 6.44%, which is significantly better than LSTM (12.47%), TCN (10.65%) and XGBoost (13.19%). For example, in "drug-related cases", the MAPE of the model of the present invention is 5.34%, which is about 5.7 percentage points lower than LSTM, and the error reduction rate is more than 50%; in the more complex "duty crime" cases, the MAPE of the present invention is 7.55%, which is 5.3 and 7.6 percentage points lower than TCN and XGBoost respectively.
[0159] Further analysis of the tabular data shows that the model of the present invention exhibits extremely high robustness in dealing with differences in case types and imbalanced data scales. For example, in the cases of "obstructing public service" and "gang fighting" with fewer test samples, despite the sparse data, the present invention still maintains low error and high stability, showing that it can effectively mine hidden structures and key trends between data through graph structure modeling and multi-scale time series encoding, thereby avoiding the defect of traditional models being overly sensitive to small sample fluctuations.
[0160] In summary, the experimental data fully demonstrates that the method of the present invention has significant practical advantages and theoretical value in the problem of predicting the dynamic progress of judicial cases. It not only improves the accuracy and consistency of predictions, but also enhances the generalization ability of different case types, especially in dealing with complex processes, uncertain time structures and staged changes. It has obvious advantages and provides reliable support for intelligent decision-making in the judicial system.
[0161] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered within the protection scope of the present invention.
Claims
1. A dynamic prediction method for the progress of judicial cases based on a time series prediction model, characterized in that, It includes the following steps: S1. Preprocess the historical data of judicial cases, including filling missing data and normalizing the data, to generate static features and dynamic time-series features; S2. Encode the normalized static features to generate static hidden vectors; S3. Construct a dynamic graph for the normalized dynamic time-series features, use the edge-conditioned adaptive graph convolutional algorithm to perform feature fusion on each node in the dynamic graph, generate a fused feature matrix, and use a multi-scale time-series encoding network based on dilated convolution and residual connection to perform time-series encoding on the fused feature matrix to obtain a time-series hidden state; S4. Use a hierarchical adaptive cross-attention fusion module for the time-series hidden state to perform local interaction and weighted fusion on the hidden states at different time scales to form a context vector; S5. Concatenate the static hidden vector and the context vector, and generate a prediction output through a fully connected layer; S6. Calculate the loss between the prediction output and the target label, and use the sensitivity-aware minimization algorithm to optimize the model parameters; S7. During the model training process, use an improved simulated annealing algorithm to adjust and optimize the hyperparameters; S8. Apply the trained model to new judicial case data and output the time prediction values at each stage.
2. The dynamic prediction method for the progress of judicial cases based on a time series prediction model according to claim 1, wherein The specific content of S2 includes: S21. Set the dimension of the static feature vector to n, and denote the static feature vector as S ′ ; S22. Define the static hidden vector as Z S .
3. A dynamic prediction method for the progress of judicial cases based on a time series prediction model according to claim 1, characterized in that The specific content of S3 includes: S31. Construct the normalized dynamic time-series features into a time-series matrix X; S32. With each dynamic feature vector X t as the node set V t , construct a graph structure G t =(V t , E t ), where each edge e t in the edge set E t,ij corresponds to the dependency relationship between the feature nodes x t,i and x t,j ; S33. Apply the edge condition adaptive graph convolution algorithm to the graph structure G t for feature fusion to generate edge features e t,ij , input the edge features into the edge parameter generation function Φ(·) to obtain a learnable convolution kernel, and perform weighted fusion on the adjacent nodes of node x t,i to obtain a fused feature vector Combine all the fused node vectors to obtain the fused feature matrix at time step t S34. Input the time series fusion features into a multi-scale time series encoding network based on dilated convolution and residual connection, and use causal convolution kernels under multiple dilation factor combinations to perform multi-scale convolutional encoding to obtain a time series hidden state sequence H = [H1, H2, …, H T , where H T represents the hidden state at the t-th time step.
4. A dynamic prediction method for the progress of judicial cases based on a time series prediction model according to claim 1, characterized in that The specific content of S4 includes: S41. Use the time-series hidden state sequence H as the input and input it into the hierarchical adaptive cross-attention fusion module; S42. Downsample the sequential hidden state sequence H according to a preset multi-time scale set to obtain a multi-scale hidden state subsequence set S43. For any two subsequences with different scales in the subsequence set and perform cross-attention fusion calculation; S44. Upsample the attention outputs at all scales to restore them to the original time step length T, and perform weighted fusion to obtain a unified context representation sequence C.
5. A dynamic prediction method for the progress of judicial cases based on a time series prediction model according to claim 1, characterized in that The S5 specifically includes: concatenating the static hidden vector Z S with the context vector sequence C at each time step t to obtain a fusion vector F t , and inputting the fusion vector F t into the prediction output generation network to generate a prediction output through a fully connected transformation, and integrating the prediction outputs to obtain a prediction output sequence 6. The dynamic prediction method for the progress of judicial cases based on a time series prediction model according to claim 1, wherein, The specific content of S6 includes: S61. Define the model parameter set θ based on the acuity perception minimization algorithm, and predict the output sequence when the initial model parameter set is θ denoted as S62. Calculate the loss between the predicted output sequence and the corresponding true target label sequence Y. Define the basic loss function as the mean squared error function: Among them, represents the predicted output sequence The predicted output at time t in, Y t The output at time t in the true target label sequence Y, and T represents the total number of time steps of the prediction sequence; S63. Calculate the local perturbation direction ∈ for the current model parameter θ * , and find the locally steepest point of the loss function in the parameter space. The calculation formula is as follows: Among them, argmax represents the function to take the maximum value of the function, ρ is a preset perturbation radius, ‖‖2 represents the Euclidean norm constraint, and ∈ represents a perturbation vector with the same dimension as the model parameter θ; S64. Substitute the perturbation direction into the loss function gradient, and perform the sensitivity-aware optimization strategy to update the model parameters. The update formula is as follows: Among them, represents taking the gradient of the loss function L base with respect to the original parameter θ, and η is the learning rate.
7. A dynamic prediction method for the progress of judicial cases based on a time series prediction model according to claim 1, characterized in that, The specific content of S7 includes: S71. Set the initial hyperparameter vector λ (0) ; S72. In the i-th iteration, use the current hyperparameter vector λ (i) to construct a model. After training and validation, obtain the loss value L (i) and the variance of the predicted output sequence S73. According to the progress distribution density function corresponding to the current prediction stage of the judicial case Introduce the stage perception regulation weight for the perturbation component to generate the stage perception perturbation vector Δλ (i) ; S74. Generate a candidate hyperparameter vector λ (i+1) = λ (i) + Δλ (i) , and retrain the model accordingly to obtain a new loss value L (i+1) and the predicted output variance S75. Define the improved acceptance probability function P (i) as follows: where exp is the exponential function, α is the uncertainty regulation factor, and T i is the temperature at the i-th round; S76. According to the above acceptance probability P (i) , accept or reject λ probabilistically (i+1) as the hyperparameter vector for the next iteration; S77. Repeatedly screen until the stopping condition is met: the number of iteration rounds \(i\geq N\) or the current temperature \(T\) i \(\leq T\) min ; Finally, select the hyperparameter vector \(\lambda\) (*) corresponding to the lowest loss value \(L\) in the acceptance history * as the optimal hyperparameter configuration of the final trained model.