Electric charge recovery dynamic risk assessment method and system based on dual-channel integrated learning and dynamic PID regulation and control
Through the dual-channel integrated learning and dynamic PID regulation, the coordinated training of arrears risk identification and amount prediction is solved, accurate risk assessment and resource allocation for power users are realized, and the efficiency of electricity bill recycling and refined management capabilities are improved.
Patent Information
- Application Number
- CN202510353992.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-25
AI Technical Summary
In the recovery of electricity bills, the existing technology has the inability to coordinate the identification of arrears risk and the prediction of amounts cannot be trained in a coordinated manner, resulting in insufficient adaptability of the model and the inability to achieve refined management. Moreover, traditional models are difficult to adapt to the unbalanced scenarios of small samples and positive and negative samples, resulting in insufficient identification of high-risk users or excessive collection of low-risk users.
The dynamic risk assessment method of electricity bill recovery based on dual-channel integrated learning and dynamic PID regulation is adopted. Through the integration of XGBoost and LightGBM model, combined with quantile random forest regression and logistic regression, the two-dimensional prediction of the arrear probability and amount is realized, and the PID controller is introduced to dynamically adjust the risk score threshold and optimize the collection strategy.
It realizes accurate identification and amount quantification of users with arrears, improves the accuracy and adaptability of risk assessment, and can promptly warn of large amounts of arrears, optimizes resource allocation, reduces invalid collection costs, and improves collection efficiency.
Smart Images

Figure CN120373840A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a power user overdue payment risk prediction system, and specifically to a dynamic risk assessment method and system for electricity bill collection based on dual-channel integrated learning and dynamic PID regulation. Background Art
[0002] Electricity bill collection, as an important link in the management of power enterprises, is directly related to the economic benefits and sustainable development of the enterprises. Timely and full electricity bill collection can ensure that power enterprises have sufficient funds for power grid construction, equipment maintenance and renewal, and provide continuous and stable power supply services.
[0003] The prediction of electricity bill collection risk has always been a key area of research in the power industry. Many scholars and practitioners have deeply analyzed the risk monitoring of electricity bill collection from different perspectives and achieved fruitful research results. The research has gone through three stages of development:
[0004] (1) Traditional statistical model stage: In the early research, Logistic regression was mainly used to analyze the relationship between user attributes and overdue payment behaviors. Although such methods are highly interpretable, they cannot capture the non-linear time-series characteristics of electricity consumption behaviors.
[0005] (2) Single machine learning model stage: Algorithms such as random forest and support vector machine were introduced, significantly improving the prediction accuracy. However, these models only output the overdue payment probability and ignore the need for amount quantification, and cannot output both the overdue payment probability and amount assessment simultaneously.
[0006] (3) Big data hybrid model stage: Deep learning and integrated learning began to be applied to electricity bill collection analysis. However, such methods rely on millions of samples, are prone to overfitting in small data scenarios, and lack business interpretability.
[0007] There are two major pain points in the traditional electricity bill collection mode. On the one hand, traditional overdue payment risk identification mainly relies on manual investigation after the fact. Usually, it starts to intervene in verification and processing only when the customer has already had an overdue payment, or even when the overdue amount is large or the overdue time is long. This passive identification method not only consumes a large amount of human resources but also misses the best intervention window period, making it difficult to take effective collection or risk control measures in a timely manner.
[0008] In addition, traditional models usually rely only on the probability of arrears or the amount of arrears for risk assessment, and use fixed rules for risk classification. They fail to dynamically adjust the risk grading based on historical collection data, changes in user behavior, and industry characteristics, resulting in insufficient adaptability of the model in different scenarios and prone to problems such as insufficient identification of high-risk users or over-collection of low-risk users. The collection strategy often adopts a "one-size-fits-all" or simple grading strategy, lacking refined management for users with different risk levels. High-risk users with large arrears do not receive sufficient investment in collection resources, while medium- and low-risk users are affected in terms of experience due to excessive collection frequency. Moreover, traditional prediction models are all trained based on big data in the millions, making it difficult to adapt to scenarios with small samples and unbalanced positive and negative samples.
[0009] The existing technology mainly relies on a single classification model for the risk assessment of power user arrears, and identifies potential risks by predicting the probability of users having arrears. Although it can identify users who may have arrears, it cannot quantify the possible arrears amount of arrear users in the future. This means that even if high-risk users can be effectively identified, there is no accurate prediction technically for the actual financial loss size faced by these users, resulting in the power enterprise being unable to adopt targeted recovery strategies for users with large arrears. This single-dimensional risk assessment cannot reflect the complexity of users' arrears behavior, making the enterprise lack sufficient data support when formulating differentiated collection measures and risk control strategies.
[0010] For arrears risk identification and arrears amount prediction, the existing technology usually trains two models independently. One is responsible for predicting the probability of users having arrears, and the other focuses on predicting the possible arrears amount of users in the future. There is a lack of effective information interaction between the two models, making it difficult to achieve information complementarity. Moreover, the separately trained models are difficult to operate collaboratively and cannot be optimized as a whole through a unified objective function. As a result, when facing complex and changeable user behavior patterns, the models cannot fully capture the internal connection between the two, thus reducing the overall prediction accuracy. Summary of the Invention
[0011] In view of the problem of power user arrears risk prediction, the present invention proposes a dynamic risk assessment method and system for electricity bill recovery based on dual-channel integrated learning and dynamic PID regulation.
[0012] Glossary:
[0013] XGBoost: Full name Extreme Gradient Boosting, is an integrated learning algorithm based on the gradient boosting framework. Its core lies in iteratively constructing a decision tree model and gradually correcting the prediction residuals of the previous model to minimize the loss function. By optimizing the objective function through second-order Taylor expansion, it can effectively handle high-dimensional sparse data.
[0014] Its objective function is composed of the loss function and the regularization term:
[0015]
[0016] Where n is the number of samples; y i ∈{0,1} represents the true label of sample i (0 is a normal user, 1 is a user in arrears); is the predicted value of sample i after the tth iteration, f t is the t-th tree, T k is the number of leaf nodes of the kth tree, w is the leaf weight, γ and λ are hyperparameters, which are the leaf node complexity penalty coefficient and L2 regularization parameter respectively; l(·) is the loss function, which is calculated as:
[0017]
[0018] in, is the predicted probability of sample i. XGBoost optimizes the loss function approximately through the second-order Taylor expansion, and uses the regularization term to prevent overfitting and improve the generalization ability of the model. The greedy algorithm is used to calculate the optimal split gain:
[0019]
[0020] G L , G R is the sum of the first-order gradients of the left and right child nodes, H L , H R It is the sum of the second-order gradients of the left child node and the right child node. The denominator is the regularized second-order gradient statistic, which is used to balance loss optimization and model complexity.
[0021] LightGBM model: The LightGBM model uses a histogram-based decision tree algorithm, a gradient unilateral sampling method that only retains samples with larger absolute gradient values, and a mutually exclusive feature bundling method to merge sparse features, quickly process electricity consumption data, and mine potential patterns and features in the data. The feature splitting criteria are:
[0022]
[0023] Among them A l , A r is the sample set of left and right child nodes after splitting; g x ,h x It is a heterogeneous integration strategy for first-order and second-order gradients.
[0024] First, the dynamic risk assessment method for electricity fee recovery based on dual-channel integrated learning and dynamic PID control includes the following processes:
[0025] S100: Data Selection; Select the electricity data of a certain number of enterprise users in the past period for analysis, including four dimensions: basic information, electricity consumption, payment, and arrears. Through SHAP value analysis, screen out the top-15 high-contribution features; Use whether there is an arrear and the expected arrear amount in the next month as labels, and adopt SMOTE-ENN hybrid sampling: First, generate arrear samples based on 5-nearest neighbor interpolation, and then remove the abnormal data among the 3-nearest neighbors to amplify the arrear data. The positive-negative sample ratio is optimized from 1:36 to 1:3, and the training set and test set are divided according to 7:3;
[0026] S200: Predict the user's arrear probability by integrating XGBoost and LightGBM;
[0027] S300: Take the arrear probability output by the above classification model as the input of the quantile random forest regression model, and predict the arrear amount of the arrear user by outputting the predicted values corresponding to different quantiles;
[0028] S400: The arrear probability obtained by the classification model in S200 and the quantile predicted value of the regression model in S300 are concatenated into a feature vector and input into the logistic regression model for training to obtain the prediction result of dual-channel fusion.
[0029] Preferably, S200 includes the following process:
[0030] S210: Set the learning rate of XGBoost to 0.05, the maximum depth limit of the tree to 6, and the other parameters to the default values. The number of leaf nodes of LightGBM is 31, and each leaf node contains at least 20 samples. Set class_weight='balanced' to let LightGBM automatically calculate the class weights and handle the class imbalance problem;
[0031] S220: Use the feature data of electricity users to train XGBoost and LightGBM respectively to obtain two independent classification models; Adopt the method of ensemble learning to fuse the prediction results of the two models to obtain the final classification result, that is, judge whether the electricity user will have an arrear.
[0032] Preferably, in S220, the Stacking ensemble method is used to fuse the prediction results of XGBoost and LightGBM. By using logistic regression as the meta-learner to dynamically adjust the model weights, the generalization ability of the model is improved, and thus the final classification result is obtained to judge whether the user will have an arrear;
[0033]
[0034] In the formula is the weight coefficient, which is optimized by grid search. are the probabilities of user arrears predicted by XGBoost and LightGBM respectively.
[0035] In the classification task, the performance of the model is comprehensively evaluated by three indicators: user recall Recall, false positive rate FPR, and AUC value ROC curve. The recall rate is the proportion of the number of arrears users successfully identified by the model to the actual arrears users.
[0036] Preferably, in S300, to balance business requirements and model performance, τ∈{0.5, 0.75, 0.9} is selected as the quantile regression target; the prediction results of the three are fused by weighting.
[0037] Preferably, in S300, the model construction steps are as follows:
[0038] S310: Decision tree construction; randomly draw samples from the training set with replacement by Bootstrap sampling to construct 200 independent decision trees, and the maximum depth of each tree is 8.
[0039] S320: Feature random selection; each tree randomly selects some features for splitting, selects the optimal splitting point to minimize the mean square error, and sets min_samples_split = 10 to ensure that each node contains at least 10 samples before it can continue to split.
[0040]
[0041] where N L , N R are the sample numbers of the left and right child nodes, D L , D R are the left and right subsets after splitting, is the subset mean;
[0042] S330: Quantile aggregation; each tree outputs the quantile prediction value, and the final result is the weighted average of the prediction values of all trees. The mathematical expression is:
[0043]
[0044] where B represents the number of trees, and b represents the b-th tree.
[0045] Preferably, in S400, the arrears probability obtained by the classification model and the quantile prediction value of the regression model are concatenated into a feature vector and input into the logistic regression model for training to obtain the prediction result of dual-channel fusion. The objective function is:
[0046]
[0047] where θ is the model parameter; λ is the L2 regularization coefficient; is the Sigmoid function, representing the final overdue risk score R of power users:
[0048]
[0049] The Sigmoid function converts the linear combination result into a probability value R ∈ (0, 1). The closer the R value is to 1, the higher the overdue risk of the user. By dynamically adjusting the meta-learner parameter θ, the model can adaptively fuse the classification and regression results, adapt to the risk characteristics of different users, and improve the prediction accuracy of the model.
[0050] Preferably, the following process is connected after S400:
[0051] S500: Introduce the PID control algorithm. By monitoring the current error, cumulative error, and error change trend between the actual value and the target value of the collection response rate, generate a control signal to dynamically adjust the risk score threshold and achieve the adaptive optimization of risk classification. The output of the PID controller is as follows:
[0052]
[0053] where e(t) is the collection response rate error of current level-I high-risk users, that is, the difference between the actual response rate and the target response rate;
[0054] K p e(t) is the proportional control term P, which is linearly adjusted according to the magnitude of the current e(t). When the target response rate is much higher than the actual response rate, it quickly responds to the error, reduces the threshold to expand the scope of high-risk users, and provides a basic control effect;
[0055] is the integral control term I, which eliminates the steady-state error according to the cumulative historical collection response rate error and keeps the system at the target value for a long time. When the response rate is lower than the target value for a continuous period of time, the weight coefficient of the meta-learner is automatically optimized to enhance the contribution of the regression model;
[0056] is the derivative control term D, which predicts the future trend of the error according to the change rate of the collection response rate error and takes measures in advance to prevent overshoot.
[0057] The PID controller combines P, I, and D to achieve precise control. P is responsible for adjusting the current collection response rate error, which is fast but may have residual errors. I is responsible for accumulating errors for long-term correction, eliminating steady-state errors but may overshoot. D is responsible for predicting the error trend, reducing oscillations, and improving stability. Through the PID dynamic regulation mechanism, the system can automatically adjust the R threshold, adaptively optimize the risk grading and collection strategy, and achieve more accurate resource allocation.
[0058] In a second aspect, a dynamic risk assessment system for electricity bill collection based on dual-channel integrated learning and dynamic PID regulation includes:
[0059] A dataset module; used to store and analyze the electricity data of a certain number of enterprise users in the past period, including four dimensions: basic information, electricity consumption, payment, and arrears; screening out the top-15 high-contribution features through SHAP value analysis; using the SMOTE-ENN hybrid sampling method with whether there is an arrearage in the next month and the expected arrearage amount as labels: first generating arrearage samples based on 5-nearest neighbor interpolation, and then removing abnormal data among the 3-nearest neighbors to amplify the arrearage data, optimizing the positive-negative sample ratio from 1:36 to 1:3, and dividing the training set and test set according to 7:3;
[0060] A user arrearage probability prediction module; predicting the user arrearage probability by integrating XGBoost and LightGBM;
[0061] A user arrearage amount prediction module; used to take the arrearage probability output by the above classification model as the input of the quantile random forest regression model, and predict the arrearage amount of the arrearage user by outputting the predicted values corresponding to different quantiles;
[0062] A dual-channel fusion prediction module; used to splice the arrearage probability obtained by the classification model in S200 and the quantile prediction value of the regression model in S300 into a feature vector and input it into a logistic regression model for training to obtain the prediction result of dual-channel fusion.
[0063] Preferably, the system further includes:
[0064] A PID controller, which generates a control signal by monitoring the current error, accumulated error, and error change trend between the actual value and the target value of the collection response rate, corresponding to the proportional term, integral term, and differential term, and dynamically adjusts the risk score threshold, that is, the R threshold, to achieve adaptive optimization of risk grading.
[0065] Advantages of the present invention over the prior art:
[0066] (1) The present invention breaks through the limitation that the traditional method cannot co-train the overdue risk identification and overdue amount prediction, and proposes a dual-channel integrated learning dynamic warning model based on a classification model + regression model. Using logistic regression as the meta-learner to fuse the classification and regression results, it realizes the two-dimensional dynamic evaluation of overdue probability prediction and overdue amount quantitative analysis. This method predicts the user's overdue probability through a heterogeneous integrated classification model, gradient boosting decision tree and light gradient boosting machine, accurately identifies overdue users, and improves the accuracy of early warning. At the same time, combined with the quantile random forest regression model to quantify the potential overdue amount, it realizes the two-dimensional evaluation of risk probability and loss scale, and realizes the information complementarity and collaborative optimization between models through cross-model information fusion, providing a refined risk management basis for power enterprises. Through elastic net regression, the weights of classification and regression are adaptively adjusted according to the risk characteristics of different users, improving the adaptability to different types of users and enhancing the reliability and interpretability of risk scores.
[0067] (2) In the embodiment, the present invention uses the quantile random forest regression to predict the overdue amount, takes the overdue probability as the input of the model, and outputs the predicted values corresponding to different quantiles by dynamically adjusting the quantiles, accurately predicting the overdue risk level and realizing the hierarchical quantification of the overdue amount. The present invention not only covers the benchmark overdue indicators of low risks, but also can identify moderately risky users, especially accurately capture the tail risks. The prediction results cover 95% of the large overdue users, which can help power enterprises give early warnings and prevent the financial impact brought by large overdue amounts, and realize the key control of high-risk, low-probability but major-loss events.
[0068] (3) In the embodiment, the present invention constructs a dynamic risk scoring mechanism. Based on the user's overdue probability and overdue amount predicted by the dual-channel risk assessment model, a risk score value R is generated for each power user, and a PID dynamic risk regulation mechanism is introduced, which combines the automatic control theory with the risk assessment, and adjusts the risk score threshold and collection strategy parameters in real time according to the historical collection response rate, improving the adaptability of the model to user behavior changes and realizing the dynamic balance between risk warning and collection efficiency. The real-time optimization of the boundary value comprehensively considers the historical collection cost and risk coverage efficiency. The high-risk level I ensures that 95% of the large overdue users are accurately identified, and the medium-risk level II balances the model sensitivity and misjudgment rate, avoiding excessive interference with low-risk users. Combining the real-time calculation of the R score and the dynamic update ability of the R threshold, it automatically triggers collection actions for different risk levels, and intelligently allocates collection resources according to the proportion of high-risk users and the overdue amount coverage rate, realizing the principle of "precise policy implementation and priority handling".
[0069] (4) In the embodiment, to address the problem of imbalanced positive and negative samples, the present invention utilizes the SMOTE-ENN hybrid sampling technique to improve the data distribution. By combining the SMOTE oversampling technique for generating synthetic samples and the ENN undersampling technique for removing noise samples, the class imbalance problem is alleviated. This enables the model to enhance the representativeness of minority class samples while maintaining the high quality and diversity of the data, improving the classification performance, reducing the high false negative rate, ensuring that high-risk overdue users can be identified in a timely manner, and providing more accurate risk warnings and collection decision-making support for the enterprise. Description of the Drawings
[0070] Figure 1 It is a schematic flow diagram of the dynamic risk assessment method for electricity bill recovery based on dual-channel integrated learning and dynamic PID regulation of the present invention.
[0071] Figure 2 In the embodiment, the SHAP value analysis is used to screen out the Top-15 high-contribution features.
[0072] Figure 3 It is a comparison chart of the ROC curves of the dual-channel model and three baseline models, namely the LR, RF, and single XGBoost classifiers, in the embodiment.
[0073] Figure 4 It is the user risk rating result calculated based on the risk score R in the embodiment. Detailed Embodiment
[0074] The present invention defines overdue users as users whose bills are not settled when due and whose liquidated damages are greater than 0. 2000 enterprise users' electricity data for the past year are selected for analysis, including four dimensions: basic information, electricity consumption, payment, and overdue. The SHAP value analysis is used to screen out the Top-15 high-contribution features, as Figure 2 shown. Using whether there is an overdue payment in the next month and the estimated overdue amount as labels, the proportion of overdue samples is 2.7%. To address the small sample and class imbalance problems, the SMOTE-ENN hybrid sampling is adopted: first, overdue samples are generated based on 5-nearest neighbor interpolation, and then the abnormal data among the 3-nearest neighbors are removed, expanding the dataset to 3000 records. The positive-negative sample ratio is optimized from 1:36 to 1:3, and the training set and test set are divided at a ratio of 7:3.
[0075] The present invention predicts the probability of users' arrears by integrating XGBoost and LightGBM. The learning rate of XGBoost is set to 0.05, the maximum depth of the tree is limited to 6, and the remaining parameters are default values. The number of leaf nodes of LightGBM is 31, and each leaf node contains at least 20 samples. Set class_weight='balanced' to let LightGBM automatically calculate the class weights and handle the class imbalance problem. First, use the feature data of power users to train XGBoost and LightGBM respectively to obtain two independent classification models. Then, adopt the method of ensemble learning to fuse the prediction results of the two models to obtain the final classification result, that is, to judge whether the power user will be in arrears.
[0076] The present invention uses the Stacking ensemble method to fuse the prediction results of XGBoost and LightGBM, dynamically adjusts the model weights by using logistic regression as the meta-learner, improves the generalization ability of the model, and thus obtains the final classification result to judge whether the user will be in arrears.
[0077]
[0078] In the formula is the weight coefficient, which is optimized by grid search. are the probabilities of users' possible arrears predicted by XGBoost and LightGBM respectively. Heterogeneous integration can effectively make up for the bias of a single model and improve the classification robustness.
[0079] In the classification task, the performance of the model is comprehensively evaluated by three indicators: user recall rate Recall, false positive rate FPR, and AUC value ROC curve.
[0080] The recall rate is the proportion of the number of arrears users successfully identified by the model to the actual arrears users. Recall = TP / (TP + FN). The higher the value, the lower the risk of missed judgment.
[0081] The false positive rate refers to the proportion of normal users misjudged as arrears users. FPR = FP / (FP + TN), which reflects the degree of injury to normal users by the model.
[0082] As shown in the confusion matrix in Table 1, the dual-channel model reaches Recall = 89.3% on the test set, that is: 201 / 225, successfully identifying 201 arrears users, and FPR = 2.67%, that is: 18 / 675, that is, only 2.66 out of every 100 normal users are misjudged.
[0083] Table 1 Confusion matrix table
[0084] Predicted delinquent user Predicted normal user Actual delinquent user TP = 201 FN = 24 Actual normal user FP = 18 TN = 657
[0085] AUC, whose full name is Area Under ROC Curve, is a core metric for measuring a model's ability to distinguish between positive and negative samples, especially suitable for scenarios with imbalanced samples. The ROC curve uses the true positive rate Recall as the vertical axis and the false positive rate FPR as the horizontal axis. The AUC value is the area under the curve, and its value range is [0, 1]. The closer the value is to 1, the stronger the model's discrimination ability. For example Figure 3 is a comparison graph of the ROC curves of the dual-channel model and three baseline models, namely the LR, RF, and single XGBoost classifiers. The AUC value of the dual-channel model is 0.88, and its ROC curve is significantly above those of the other models, indicating that this model has significant advantages in risk user identification. In contrast, the AUC of the traditional logistic regression model is 0.82, and its curve is closer to the diagonal line, showing weaker discrimination ability.
[0086] In the present invention, the overdue probability output by the above classification model is used as the input of the quantile random forest regression model, and the overdue amounts of overdue users are predicted by outputting the predicted values corresponding to different quantiles. To balance business requirements and model performance, τ ∈ {0.5, 0.75, 0.9} is selected as the quantile regression target. τ = 0.5 reflects the typical overdue scale and is used to formulate the benchmark collection strategy; τ = 0.75 identifies medium-risk users; τ = 0.9 covers most overdue scenarios and can quantify the risk of large overdue amounts at the tail, supporting the enterprise's risk reserve planning. By weighted fusion of the prediction results of the three, the model can dynamically adapt to different risk scenarios and support the formulation of refined collection strategies.
[0087] Steps for model construction:
[0088] Decision tree construction: Randomly sample with replacement from the training set through Bootstrap sampling to construct 200 independent decision trees, and the maximum depth of each tree is 8.
[0089] Random feature selection: Each tree randomly selects some features for splitting, selects the optimal split point to minimize the mean squared error, and sets min_samples_split = 10 to ensure that each node contains at least 10 samples before it can continue to split.
[0090]
[0091] where, N L , N R are the number of samples in the left and right child nodes, D L , D R are the left and right subsets after splitting, is the subset mean.
[0092] Quantile aggregation: Each tree outputs the quantile predicted value, and the final result is the weighted average of the predicted values of all trees. The mathematical expression is:
[0093]
[0094] Among them, B represents the number of trees, and b represents the b-th tree.
[0095] In the present invention, the performance evaluation of the regression model uses the quantile loss value. As shown in Table 2, when the quantile is selected as 0.9, it can effectively prevent the financial impact brought by large overdue fees. Power enterprises need to focus on such high-risk, low-probability but large-loss events. By predicting the 90% quantile through the quantile regression model, the tail risk can be effectively captured, which conforms to the prudence principle of the enterprise. When the quantile is 0.9, the quantile loss value of the dual-channel model is the lowest, and the coverage rate of customers with large overdue fees is the highest, thus proving the accuracy of the dual-channel model in predicting the overdue fee amount of users. The experiment shows that this method can cover 95% of the users with large overdue fees, that is, users with an overdue fee amount > 50,000 yuan, supporting the effectiveness of the 90% quantile in the business scenario. The high recall rate of the classification model ensures that as many overdue users as possible are identified, while the high coverage rate of the regression model further ensures the accurate quantification of the tail risk of large overdue fees. The two together improve the comprehensiveness of risk assessment.
[0096] Table 2 Quantile Regression Performance Table (τ = 0.9)
[0097] Model Quantile loss Large delinquent coverage (>¥50,000) Ordinary random forest (RF) 8 68% Single-channel quantile regression (QRF) 7.2 82% Dual-channel model 6.5 95%
[0098] The present invention selects logistic regression as the meta-learner for fusion, optimizes the weight coefficients through elastic net regression, dynamically adjusts the weights of classification and regression, and at the same time introduces the L2 regularization term to prevent overfitting and enhance sparsity. The overdue probability obtained from the classification model and the quantile prediction value of the regression model are concatenated into a feature vector and input into the logistic regression model for training to obtain the prediction result of dual-channel fusion. The objective function is:
[0099]
[0100] where θ is the model parameter; λ is the L2 regularization coefficient; is the Sigmoid function, representing the final overdue risk score R of power users:
[0101]
[0102] The Sigmoid function converts the linear combination result into a probability value R ∈ (0, 1). The closer the R value is to 1, the higher the overdue risk of the user. By dynamically adjusting the meta-learner parameter θ, the model can adaptively fuse the classification and regression results, adapt to the risk characteristics of different users, and improve the prediction accuracy of the model.
[0103] Based on the output results of the dual-channel integrated learning model, the present invention combines the overdue probability predicted by the classification model with the large overdue risk quantified by the regression model through a dynamic fusion mechanism to construct a comprehensive risk score R, and uses the Sigmoid function to map the linear combination of the classification and regression results to the interval (0, 1), intuitively reflecting the overdue risk intensity of users and realizing the continuous dynamic quantification of risks. To further support the differential collection decision-making of power enterprises, according to business requirements and cost-benefit analysis, users are initially divided into three risk levels, as shown in Table 1:
[0104] Level I: R > 0.7; Level II: 0.4 ≤ R ≤ 0.7; Level III: R < 0.4.
[0105] Among them, the setting of R needs to consider: ensuring that as many large overdue users as possible are accurately identified; being able to balance the model sensitivity and misjudgment rate, and avoiding excessive interference with low-risk users. Power enterprises dynamically adjust the warning strategies and collection measures for users according to the level of risk. For high-risk users at Level I, "high frequency + multi-channel" intervention is adopted, aiming to attract the attention of users immediately and reduce the possibility of overdue payments. For users at Level II and Level III, relatively mild warning strategies are adopted, which not only avoid over-disturbing low-risk users but also ensure that users can be effectively reminded to pay when necessary.
[0106] To optimize the setting of the risk level parameter R, the present invention introduces a PID control algorithm. By monitoring the current error, cumulative error, and error change trend between the actual value and the target value of the collection response rate, a control signal is generated to dynamically adjust the risk score threshold, the R threshold, to achieve the adaptive optimization of risk classification. Its core idea is to balance the risk coverage ability and collection efficiency through continuous feedback and regulation, and ensure that resources are preferentially allocated to high-value users. The output of the PID controller is as follows:
[0107]
[0108] Among them, e(t) is the collection response rate error of high-risk users at Level I currently, that is, the difference between the actual response rate and the target response rate.
[0109] K p e(t) is the proportional control term P, which is linearly adjusted according to the magnitude of the current e(t). When the target response rate is much higher than the actual response rate, it responds quickly to the error, reduces the threshold to expand the scope of high-risk users, and provides a basic control effect.
[0110] It is the integral control term I. According to the accumulated historical collection response rate error, it eliminates the steady-state error and enables the system to maintain the target value in the long term. When the response rate is lower than the target value for a continuous period of time, the weight coefficient of the meta-learner is automatically optimized to enhance the contribution of the regression model.
[0111] It is the differential control term D. According to the change rate of the collection response rate error, it predicts the future trend of the error and takes measures in advance to prevent overshoot.
[0112] The PID controller combines P, I, and D for precise control. P is responsible for adjusting the current collection response rate error, which is fast but may have residual errors. I is responsible for long-term correction of the accumulated error, eliminating the steady-state error but may overshoot. D is responsible for predicting the error trend, reducing oscillations, and improving stability. Through the PID dynamic regulation mechanism, the system can automatically adjust the R threshold, adaptively optimize the risk grading and collection strategy, and achieve more accurate resource allocation.
[0113] Table 3 Risk Dynamic Early Warning Strategy Table
[0114] Risk level Early warning strategy Level I Day 1 SMS + APP push; Day 3 automated call; Day 5 account manager visit Level II Day 3 SMS, Day 7 APP pop-up, Day 15 phone call back Level III Regular automated SMS reminder 7 days before bill due date
[0115] The present invention comprehensively evaluates the business value of the dynamic risk rating through two indicators: risk coverage ability and collection efficiency. The user risk rating results calculated based on the risk score R are as Figure 4 shown. In the test set, the proportion of users with level I risk is 9.6%, the proportion of users with level II risk is 39.0%, and the proportion of users with level III risk is 51.4%. High-risk users account for less than 10% of the total sample, but cover 95% of the large-amount overdue cases, that is, cases with overdue amount > 50,000 yuan. This indicates that the model can effectively focus on the tail risk and provide a key basis for resource allocation. By analyzing the characteristics of users in each risk level, it is found that users at level I generally have the following commonalities: the historical maximum overdue amount exceeds 2.5 times the industry average, the overdue frequency in the past 3 months ≥ 3 times, and the industry annual average overdue rate is higher than 25%. Such users are usually in high-volatility industries, such as manufacturing and construction, and the electricity consumption volatility σ > 0.3 is significantly higher than other groups.
[0116] By comparing with the traditional single-dimensional model that only relies on the overdue probability or amount, the dual-channel model shows significant advantages in collection efficiency. Because the model accurately locates large-amount overdue users, the collection response rate of level I users has increased to 78%, while the traditional model is 62%; the misjudgment rate of level III users has dropped to 4.3%, while the traditional model is 12%, effectively reducing the ineffective collection cost; the average collection cycle of high-risk users has been shortened to 7 days, while the traditional model is 11 days.
Claims
1. A dynamic risk assessment method for electricity bill collection based on dual-channel integrated learning and dynamic PID regulation, characterized in that It includes the following processes: S100: Data selection; Select the electricity data of a certain number of enterprise users in the past period for analysis, including four dimensions: basic information, electricity consumption, payment, and arrears. Through SHAP value analysis, select the top-15 high-contribution features; Use whether there is an arrear and the expected arrear amount in the next month as labels, and adopt SMOTE-ENN hybrid sampling: First, generate arrear samples based on 5-nearest neighbor interpolation, and then remove the abnormal data among the 3-nearest neighbors to amplify the arrear data. The positive-negative sample ratio is optimized from 1:36 to 1:3, and the training set and test set are divided according to 7:3; S200: Predict the user's arrear probability by integrating XGBoost and LightGBM; S300: Take the arrear probability output by the above classification model as the input of the quantile random forest regression model, and predict the arrear amount of the arrear user by outputting the predicted values corresponding to different quantiles; S400: Combine the overdue probability obtained from the classification model in S200 and the quantile prediction value of the regression model in S300 to form a feature vector and input it into a logistic regression model for training to obtain a prediction result of dual-channel fusion.
2. The dynamic risk assessment method for electricity bill collection based on dual-channel integrated learning and dynamic PID regulation according to claim 1, characterized in that The S200 includes the following processes: S210: Set the learning rate of XGBoost to 0.05, the maximum depth limit of the tree to 6, and the other parameters to the default values. The number of leaf nodes of LightGBM is 31, and each leaf node contains at least 20 samples. Set class_weight='balanced' to let LightGBM automatically calculate the class weights and handle the class imbalance problem; S220: Use the feature data of electricity users to train XGBoost and LightGBM respectively to obtain two independent classification models; Adopt the method of ensemble learning to fuse the prediction results of the two models to obtain the final classification result, that is, judge whether the electricity user will have an arrear.
3. The dynamic risk assessment method for electricity bill collection based on dual-channel integrated learning and dynamic PID regulation according to claim 2, characterized in that In the S220, the Stacking ensemble method is used to fuse the prediction results of XGBoost and LightGBM. By using logistic regression as the meta-learner to dynamically adjust the model weights, the generalization ability of the model is improved, and the final classification result is obtained to judge whether the user will have an arrear; where is the weight coefficient, optimized by grid search are the probabilities of users' possible arrears predicted by XGBoost and LightGBM respectively 4. The dynamic risk assessment method for electricity bill collection based on dual-channel integrated learning and dynamic PID regulation according to claim 1, wherein In the S300, to balance the business requirements and model performance, select τ∈{0.5, 0.75, 0.9} as the quantile regression target; Weightedly fuse the prediction results of the three; 5. The dynamic risk assessment method for electricity bill collection based on dual-channel integrated learning and dynamic PID regulation according to claim 1, characterized in that In the S300, the model construction steps: S310: Decision tree construction; Randomly draw samples from the training set with replacement through Bootstrap sampling to construct 200 independent decision trees, and the maximum depth of each tree is 8; S320: Feature random selection; Each tree randomly selects some features for splitting, selects the optimal splitting point to minimize the mean squared error, and set min_samples_split = 10 to ensure that each node contains at least 10 samples before it can continue to split; Among them, N L , N R is the number of samples of the left and right child nodes, D L , D R are the left and right subsets after splitting, is the subset mean; S330: Quantile aggregation; Each tree outputs the quantile predicted value, and the final result is the weighted average of the predicted values of all trees. The mathematical expression is: where B represents the number of trees, and b represents the b-th tree.
6. The dynamic risk assessment method for electricity bill collection based on dual-channel integrated learning and dynamic PID regulation according to claim 1, wherein In the S400, the overdue probability obtained by the classification model and the quantile prediction value of the regression model are concatenated into a feature vector and input into the logistic regression model for training to obtain the prediction result of dual-channel fusion; the objective function is: where θ is the model parameter; λ is the L2 regularization coefficient; is the Sigmoid function, representing the final overdue risk score R of electricity users: The Sigmoid function converts the result of the linear combination into a probability value \(R\in(0,1)\). The closer the \(R\) value is to 1, the higher the risk of user arrears. By dynamically adjusting the parameters \(\theta\) of the meta-learner, the model can adaptively fuse the classification and regression results, adapt to the risk characteristics of different users, and improve the prediction accuracy of the model.
7. The dynamic risk assessment method for electricity charge recovery based on dual-channel integrated learning and dynamic PID regulation according to any one of claims 1-6, characterized in that, The following process is connected after the S400: S500: Introduce the PID (Proportional-Integral-Derivative) control algorithm. By monitoring the current error (proportional term), cumulative error (integral term), and error change trend (derivative term) between the actual value and the target value of the collection response rate, generate a control signal, and dynamically adjust the risk score threshold (\(R\) threshold) to achieve adaptive optimization of risk classification. The output of the PID controller is as follows: where \(e(t)\) is the collection response rate error of the current level-I high-risk users, that is, the difference between the actual response rate and the target response rate; K p e(t) is the proportional control term P, which is linearly adjusted according to the magnitude of the current e(t). When the target response rate is much higher than the actual response rate, it quickly responds to the error, reduces the threshold to expand the scope of high-risk users, and provides a basic control function; This is the integral control term I. Based on the cumulative historical collection response rate error, it eliminates the steady-state error and enables the system to maintain the target value in the long term. When the response rate is lower than the target value for a continuous period of time, the weight coefficient of the meta-learner is automatically optimized to enhance the contribution of the regression model; It is the differential control term D. According to the change rate of the collection response rate error, it predicts the future trend of the error and takes measures in advance to prevent overshoot.
8. The dynamic risk assessment system for electricity bill collection based on dual-channel integrated learning and dynamic PID regulation is characterized in that including: Dataset module; Used to store and analyze the power data of a certain number of enterprise users over a certain period of time, including four dimensions: basic information, electricity consumption, payment, and arrears; Select the top-15 high-contribution features through SHAP value analysis; use the next month's arrears status and the expected arrears amount as labels, and adopt SMOTE-ENN hybrid sampling: first generate arrears samples based on 5-nearest neighbor interpolation, and then remove the abnormal data in the 3-nearest neighbors to amplify the arrears data. The positive-negative sample ratio is optimized from 1:36 to 1:3, and the training set and test set are divided according to 7:3; User arrears probability prediction module; predict the user arrears probability by integrating XGBoost and LightGBM; User arrears amount prediction module; used to take the arrears probability output by the above classification model as the input of the quantile random forest regression model, and predict the arrears amount of the arrears users by outputting the predicted values corresponding to different quantiles; Dual-channel fusion prediction module; used to combine the overdue probability obtained by the classification model in the S200 and the quantile prediction value of the regression model in the S300 to splice into a feature vector and input it into a logistic regression model for training to obtain the prediction result of dual-channel fusion.
9. The dynamic risk assessment system for electricity bill collection based on dual-channel integrated learning and dynamic PID regulation according to claim 8, characterized in that, It also includes: PID controller, by monitoring the current error (proportional term), cumulative error (integral term), and error change trend (derivative term) between the actual value and the target value of the collection response rate, generate a control signal, and dynamically adjust the risk score threshold (\(R\) threshold) to achieve adaptive optimization of risk classification.
Citation Information
Cited By
Big data problem clue mining method based on mutual exclusiveness rule
CN120929771A
A big data problem clue mining method based on mutual exclusivity rule
CN120929771B
Coronary artery CT imaging reexamination time window prediction method fused with multi-source information
CN121075710A
Industrial component defect self-adaptive calibration evaluation method and system based on information entropy
CN121302219A
An Adaptive Calibration and Evaluation Method and System for Defects in Industrial Components Based on Information Entropy
CN121302219B