Arrearage risk management and control method, device and equipment

Through a hybrid density network and LightGBM algorithm, combined with the characteristic data of power users and step-by-step electricity price information, accurately predict the probability of arrears and payment probability, and generate a personalized collection reminder strategy, which solves the problem of unrefined control of arrears in the existing technology and improves the control effectiveness of arrears risks.

CN120373859APending Publication Date: 2025-07-25NEUSOFT CORP
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510459036.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing technology fails to effectively consider the uncertainty and complexity of the time scale of users' arrears, resulting in insufficient refinement of the risk control of arrears and poor control effectiveness.

Method used

The mixed density network predicts the probability distribution of arrears and payment probability distribution of the target risk power users, determines the time node of arrears and risk warning of arrears and generates a personalized collection prompt strategy, and uses the LightGBM algorithm to build a credit evaluation model to identify high-risk users, and combines feature data and ladder electricity price information for refined management and control.

Benefits of technology

It realizes accurate prediction of the probability of user arrears and payment behaviors on the time scale, generates a personalized collection reminder strategy, improves the control effectiveness of arrears risks, and achieves more refined arrears risk management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373859A_ABST
    Figure CN120373859A_ABST
Patent Text Reader

Abstract

The invention discloses an arrearage risk management and control method, device and equipment. The method comprises the steps of firstly determining a target risk power consumer; predicting arrearage probability distribution and payment probability distribution of the target risk power consumer through a hybrid density network based on the feature data, the account balance and the step tariff use information of the target risk power consumer; determining an arrearage risk early warning time node corresponding to the target risk power user based on the arrearage probability distribution, and generating a collection prompt strategy corresponding to the target risk power user based on the payment probability distribution; and before the time reaches the arrearage risk early warning time node, performing electric charge collection prompting on the target risk power user based on the collection prompting strategy. According to the technical scheme, based on the arrearage probability distribution and the payment probability distribution predicted by the mixed density network, the occurrence probability of the arrearage behavior and the payment behavior of the user is predicted on the time scale, so that arrearage risk management and control are more finely realized, and the management and control effectiveness of the arrearage risk is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information processing technologies, and particularly to a method, device, and equipment for delinquent fee risk control. Background Art

[0002] In the power industry, electricity charges are the main source of income for power supply companies. In addition, electricity charges are also the material guarantee for the infrastructure construction of power supply facilities. The recovery of electricity charges is related to the sustainability and stability of the healthy development of power supply companies. However, the situation where some users default on electricity charges occurs from time to time. Due to the public welfare nature of the power industry, power supply companies cannot achieve electricity charge collection through arbitrary power outage operations. Therefore, the collection of electricity charges has always been a difficult problem in the industry. With the rise of big data and machine learning technologies, researchers in the power industry have also begun to explore the use of advanced technologies to analyze and model user data, thereby improving the power supply company's ability to identify and control user delinquent fee risks.

[0003] Currently, power supply companies usually only use the delinquent behavior information of users in the past to control delinquent fee risks, without considering the uncertainty and complexity of users' delinquent behaviors on the time scale, resulting in difficulty in achieving refined delinquent fee risk control, and thus the effectiveness of delinquent fee risk control is not good. Summary of the Invention

[0004] Based on the above problems, the present application provides a method, device, and equipment for delinquent fee risk control, aiming to achieve refined delinquent fee risk control and improve the effectiveness of delinquent fee risk control.

[0005] The embodiments of the present application disclose the following technical solutions:

[0006] The first aspect of the present application provides a method for delinquent fee risk control, and the method includes:

[0007] Determine target risk power users;

[0008] Based on the characteristic data, account balance, and ladder electricity price usage information of the target risk power users, predict the delinquent probability distribution and payment probability distribution of the target risk power users through a mixture density network; wherein, the delinquent probability distribution reflects the change of the delinquent probability over time, and the payment probability distribution reflects the change of the payment probability over time; the characteristic data of the target risk power users at least includes: the delinquent behavior characteristics and payment behavior characteristics of the target risk power users;

[0009] Based on the delinquent probability distribution, determine the delinquent risk warning time node corresponding to the target risk power users, and generate a collection reminder strategy corresponding to the target risk power users based on the payment probability distribution;

[0010] Before the time reaches the overdue risk warning time node, perform electricity bill collection reminders on the target risk power users based on the collection reminder strategy.

[0011] In an alternative implementation, the determination of the target risk power users includes:

[0012] Generate a credit risk probability value for the target power user based on the characteristic data of the target power user;

[0013] If the credit risk probability value is higher than the first probability threshold, determine that the target power user is a target risk power user;

[0014] After determining that the target power user is a target risk power user, the method further includes:

[0015] Use the SHAP algorithm to calculate the characteristic importance representation values of each characteristic in the characteristic data of the target power user for the credit risk probability value.

[0016] In an alternative implementation, after using the SHAP algorithm to calculate the characteristic importance representation values of each characteristic in the characteristic data of the target power user for the credit risk probability value, the method further includes:

[0017] Sort the characteristic data of the target power user according to the characteristic importance representation values, and display the characteristic sorting result list; and / or,

[0018] Determine the top N characteristics with the largest characteristic importance representation values in the characteristic data of the target power user, and display the determined top N characteristics and the corresponding characteristic importance representation values; and / or,

[0019] Determine multiple characteristics in the characteristic data of the target power user whose characteristic importance representation values are greater than the preset threshold, and display the multiple characteristics and the corresponding characteristic importance representation values.

[0020] In an alternative implementation, the overdue risk control method further includes:

[0021] Establish sub-categories of the user group according to the characteristic sorting result lists corresponding to multiple different power users; each sub-category corresponds to a type of overdue risk;

[0022] Based on the characteristic sorting result list corresponding to the target power user, identify the sub-category corresponding to the target power user, and determine the target overdue risk type corresponding to the target power user.

[0023] In an alternative implementation, based on the overdue probability distribution, determining the overdue risk warning time node corresponding to the target risk power user includes:

[0024] Determine the overdue risk warning time node of the target power user based on the overdue probability distribution of the target power user and the target overdue risk type corresponding to the target power user;

[0025] and / or,

[0026] The generating the collection reminder strategy corresponding to the target risk power user based on the payment probability distribution includes:

[0027] Generate the collection reminder strategy corresponding to the target power user based on the payment probability distribution and the target overdue risk type corresponding to the target power user.

[0028] In an alternative implementation, the overdue risk control method further includes:

[0029] Adjust the first probability threshold based on the target overdue risk type corresponding to the target power user to obtain a second probability threshold;

[0030] The performing, before the time reaches the overdue risk warning time node, a collection reminder for the electricity bill of the target risk power user based on the collection reminder strategy includes:

[0031] If the credit risk probability value of the target power user is higher than the second probability threshold, then before the time reaches the overdue risk warning time node, perform a collection reminder for the electricity bill of the target risk power user based on the collection reminder strategy.

[0032] In an alternative implementation, before generating the credit risk probability value of the target power user based on the characteristic data of the target power user, the method further includes:

[0033] Construct a credit evaluation model using the LightGBM algorithm;

[0034] The generating the credit risk probability value of the target power user based on the characteristic data of the target power user includes:

[0035] Use the characteristic data of the target power user as the input of the credit evaluation model, analyze the characteristic data of the target power user through the credit evaluation model, and output the credit risk probability value of the target power user by the credit evaluation model.

[0036] In an alternative implementation, the overdue risk control method further includes:

[0037] Select multiple features as historical training features from the historical characteristic dataset of power users through the SelectKBest method;

[0038] Using the historical training features and overdue payment date information of power users, an overdue payment prediction model based on a mixture density network is constructed; and, based on the historical training features and payment date information of power users, a payment prediction model based on a mixture density network is constructed;

[0039] Based on the characteristic data, account balance, and tiered electricity price usage information of the target risk power users, the overdue payment probability distribution and payment probability distribution of the target risk power users are predicted through a mixture density network, including:

[0040] Inputting the characteristic data, account balance, and tiered electricity price usage information of the target risk power users screened by the SelectKBest method into the overdue payment prediction model to obtain the overdue payment probability distribution output by the overdue payment prediction model; and inputting the characteristic data, account balance, and tiered electricity price usage information of the target risk power users screened by the SelectKBest method into the payment prediction model to obtain the payment probability distribution output by the payment prediction model.

[0041] In an optional implementation manner, the model evaluation indexes of the overdue payment prediction model and the model evaluation indexes of the payment prediction model both include: an accuracy index and a mean absolute error index;

[0042] The accuracy index is: a value in the range of [0, 1] obtained by converting the negative log-likelihood loss of the model with a preset value as the loss upper limit;

[0043] The mean absolute error index is: obtained by calculating the mean absolute error between the expected value of the predicted probability distribution and the actual sample label.

[0044] The second aspect of this application provides an overdue payment risk control device, and this device includes:

[0045] A risk user determination module, used to determine target risk power users;

[0046] A probability distribution prediction module, used to predict the overdue payment probability distribution and payment probability distribution of the target risk power users through a mixture density network based on the characteristic data, account balance, and tiered electricity price usage information of the target risk power users; wherein, the overdue payment probability distribution reflects the change of the overdue payment probability over time, and the payment probability distribution reflects the change of the payment probability over time; the characteristic data of the target risk power users at least includes: the overdue payment behavior characteristics and payment behavior characteristics of the target risk power users;

[0047] A time node determination module, used to determine the overdue payment risk warning time node corresponding to the target risk power users based on the overdue payment probability distribution;

[0048] A strategy generation module, configured to generate a collection reminder strategy corresponding to the target risk power user based on the payment probability distribution;

[0049] A reminder module, configured to perform an electricity bill collection reminder on the target risk power user based on the collection reminder strategy before the time reaches the overdue risk warning time node.

[0050] A third aspect of the present application provides an overdue risk control device, which includes: a memory and a processor;

[0051] The memory is used to store a computer program;

[0052] The processor is configured to run the computer program, and when the computer program runs, it executes the steps of the overdue risk control method introduced in any implementation manner in the first aspect.

[0053] Compared with the prior art, the present application has the following beneficial effects:

[0054] The present application provides an overdue risk control method, device and equipment. In the method, first, a target risk power user is determined; based on the characteristic data, account balance and ladder electricity price usage information of the target risk power user, the overdue probability distribution and payment probability distribution of the target risk power user are predicted through a mixture density network; based on the overdue probability distribution, the overdue risk warning time node corresponding to the target risk power user is determined, and a collection reminder strategy corresponding to the target risk power user is generated based on the payment probability distribution; before the time reaches the overdue risk warning time node, an electricity bill collection reminder is performed on the target risk power user based on the collection reminder strategy. The technical solution of the present application is based on the overdue probability distribution and payment probability distribution predicted by the mixture density network. The overdue probability distribution reflects the change of the overdue probability over time, and the payment probability distribution reflects the change of the payment probability over time, so as to accurately predict the occurrence probabilities of the user's overdue behavior and payment behavior on the time scale. Therefore, the overdue risk control is realized more refinedly, and the control effectiveness of the overdue risk is improved. Description of the Drawings

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0056] Figure 1 It is a flowchart of an overdue risk control method provided by an embodiment of the present application;

[0057] Figure 2 Schematic diagram of the structure of a mixed density network provided by an embodiment of the present application;

[0058] Figure 3 Schematic diagram of the internal structure of a hidden layer provided by an embodiment of the present application;

[0059] Figure 4 Schematic diagram of another structure of a mixed density network provided by an embodiment of the present application;

[0060] Figure 5 Architecture diagram for predicting overdue payment probability distribution and payment probability distribution through a mixed density network provided by an embodiment of the present application;

[0061] Figure 6 Schematic diagram of the content of a feature group provided by an embodiment of the present application;

[0062] Figure 7 Display effect diagram of the feature importance results of a certain user;

[0063] Figure 8 Overall flowchart of an overdue payment risk control method provided by an embodiment of the present application;

[0064] Figure 9 Schematic diagram of the payment probability distribution of a certain high-risk electricity user;

[0065] Figure 10 Schematic diagram of the prediction situation taking the payment probability distribution as an example;

[0066] Figure 11 Schematic diagram of the structure of an overdue payment risk control device provided by an embodiment of the present application. Detailed implementation manners

[0067] Currently, in terms of electricity bill collection, due to the public welfare nature of the power industry, power supply companies are often passive and it is difficult to effectively control users' overdue payment situations. Although the problem of electricity bill collection can be alleviated with the rise of big data and machine learning technologies, since the current technologies do not take into account the uncertainty and complexity of users' overdue payment behaviors on the time scale, the refinement degree of overdue payment risk control is lacking and the control effectiveness is still not ideal.

[0068] To this end, through research, the inventor has proposed a targeted solution: providing a method, device, and equipment for delinquent fee risk control. In this solution, based on the characteristic data, account balance, and tiered electricity price usage information of target risk power users, a mixture density network is used to predict the delinquent probability distribution and payment probability distribution of target risk power users. Based on the delinquent probability distribution, determine the delinquent risk warning time node corresponding to the target risk power user, and generate a collection reminder strategy corresponding to the target risk power user based on the payment probability distribution; before the time reaches the delinquent risk warning time node, give a reminder for electricity bill collection to the target risk power user based on the collection reminder strategy. The technical solution of this application, with the characteristic data, account balance, and tiered electricity price usage information of target risk power users as the data basis, can achieve personalized analysis of users and is more flexible and closer to user behavior. In addition, by means of the mixture density network, it is possible to predict the changes in delinquent probability and payment probability on the time scale, which is conducive to more accurately judging the delinquent risk warning time node and generating a personalized collection reminder strategy. Furthermore, it is possible to achieve more refined delinquent risk control and improve the control effectiveness of delinquent risks.

[0069] In order to enable those skilled in the art of this technology to better understand the solution of this application, the following will clearly and completely describe the technical solution in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.

[0070] See Figure 1 , which is a flowchart of a method for delinquent fee risk control provided by an embodiment of this application. As Figure 1 shown, this method includes:

[0071] S101. Determine the target risk power user.

[0072] Here, the target risk power user refers to a user with a relatively high risk of delinquency. Such users can be identified and screened through data analysis. For such users, it is necessary to give a reminder for electricity bill collection at the necessary moment in combination with the actual situation.

[0073] The following introduces an optional implementation for determining the target risk power user. In the optional implementation, the process of determining the target risk power user may include: generating a credit risk probability value of the target power user based on the characteristic data of the target power user; if the credit risk probability value is higher than the first probability threshold, determine the target power user as the target risk power user.

[0074] The characteristic data of the target power users generally includes at least the arrears behavior characteristics and payment behavior characteristics of the target power users themselves. Therefore, based on the characteristic data of the target power users, the personal credit risk of the users can be assisted in analysis. In this application, the credit risk is reflected in the form of a probability value within the range of 0 to 1. In practical applications, it can also be characterized by numerical values or indicators within other range intervals, which is not limited here.

[0075] In an alternative implementation, before generating the credit risk probability value of the target power user based on the characteristic data of the target power user, the method further includes: constructing a credit evaluation model using the LightGBM algorithm. Generating the credit risk probability value of the target power user based on the characteristic data of the target power user specifically includes: taking the characteristic data of the target power user as the input of the credit evaluation model, analyzing the characteristic data of the target power user through the credit evaluation model, and outputting the credit risk probability value of the target power user by the credit evaluation model.

[0076] LightGBM, whose full English name is Light Gradient Boosting Machine, is a distributed gradient boosting framework based on the gradient boosting decision tree algorithm. LightGBM uses a histogram-based decision tree algorithm, supports efficient parallel training, has lower memory usage and higher accuracy. The main features of LightGBM include histogram difference acceleration, LLeaf-wise leaf growth strategy, category feature support, and parallel learning. By adjusting parameters such as the number of leaf nodes num_leaves and the minimum number of samples in each leaf node min_data_in_leaf, overfitting can be effectively prevented and the model performance can be improved. The growth strategy of the LightGBM leaf nodes is to preferentially select the leaf nodes with the largest information gain for splitting. The following shows the calculation formulas for information entropy, conditional entropy, and information gain. These three calculation formulas (1) to (3) are of great significance for constructing decision trees in the credit evaluation model.

[0077]

[0078] In formula (1), K represents the total number of result classifications, k represents the kth category among the K classifications, D represents the total number of samples, C k represents the number of samples belonging to a certain kth category, and H(D) represents information entropy.

[0079]

[0080] In formula (2), H(D|A) represents conditional entropy, which is the weighted average entropy of the divided data set under the condition of feature A and is used to represent the remaining uncertainty of the divided data set. D iDenote the \(i\)-th subset. The dataset is partitioned into \(n\) subsets using feature \(A\). Each subset corresponds to a certain value of feature \(A\). For example, the feature of gender can partition the dataset into two subsets: male and female.

[0081] It represents calculating the entropy of each subset and performing a weighted average based on the number of samples in each subset. There are multiple samples within each subset, and the information entropy of the subset can be calculated in the same way as in formula (1). D ik Denote the number of samples belonging to the \(k\)-th category in the \(i\)-th subset.

[0082] IGain(S,g)=H(D)-H(D|A) Formula (3)

[0083] In formula (3), IGain(S, g) represents the information gain. As shown in this formula, the information gain is the difference between the information entropy and the conditional entropy.

[0084] In the LightGBM algorithm, for each decision tree, the goal is to minimize the following objective function:

[0085]

[0086] Formula (4) is the objective function used for optimizing the decision tree, and the goal is to minimize the value of \(L\). In formula (4), \(l(y\) i , \(f(x\) i )) is the prediction error loss; \(\Omega(f)\) is the regularization term, which is used to prevent overfitting of the credit evaluation model. Each time the information gain is calculated, the position with the maximum information gain on the decision tree is calculated and the node is split at this position, so as to minimize the value of the objective function.

[0087] The initial parameters of the credit rating model include: learning rate, number of trees, maximum depth of the tree, minimum number of samples for splitting, minimum number of samples for leaves, subsample ratio, and feature selection ratio. Optionally, 10-fold cross-validation is adopted during training, grid search is performed, all possible combinations of hyperparameters are traversed, and the combination with the best performance is selected. Furthermore, more model parameters are adjusted, such as the coefficient of the L1 regularization term, the coefficient of the L2 regularization term, etc. Finally, the model is retrained using the optimal hyperparameters, and the accuracy, precision, recall, and F1 score of the final model are recorded. In practical applications, the effectiveness of the model can be reflected by these evaluation indicators such as accuracy, precision, recall, and F1 score.

[0088] In the embodiments of the present application, the credit evaluation model is trained based on the LightGBM algorithm. The model architecture includes multiple decision trees, and the nodes in the decision trees are formed by splitting at the position with the maximum information gain, and the optimization objective is to minimize the objective function shown in formula (4). Applying this mechanism to the constructed credit evaluation model can obtain a lower loss value under the same number of iterations, thereby efficiently improving the model accuracy. The established credit evaluation model has the characteristics of accurately evaluating the credit risk and electricity bill arrears risk of power users. The accuracy of the evaluation of the credit risk and electricity bill payment risk of power users is closely related to the control effectiveness of the technical solution of the present application for the arrears risk. Therefore, the credit evaluation model established based on the LightGBM algorithm can assist in improving the control efficiency of the arrears risk and facilitate the power supply company to achieve refined management of the arrears risk.

[0089] In addition, the LightGBM algorithm supports large-scale data parallel computing. For a power supply company, the associated power user group is very large, and as time goes by, the amount of data generated by users is also very large. Therefore, it is necessary for the credit evaluation model to frequently process a large amount of data to achieve risk assessment. Facing this demand, the LightGBM algorithm provides the necessary performance support for the timeliness of the arrears risk control.

[0090] It can be understood that the magnitude of the credit risk probability value output by the credit evaluation model is positively correlated with the credit risk. In addition, the higher the credit risk, the worse the credit, and the higher the risk of electricity bill arrears. Therefore, the credit risk probability value is also positively correlated with the electricity bill arrears risk. That is to say, the higher the credit risk probability value, the higher the credit risk and the higher the electricity bill arrears risk.

[0091] As an example, set the first probability threshold to 50%. If the credit risk probability value of a certain target power user is higher than 50%, it can be determined that its credit risk probability value is too high, so it is determined as a target risk power user. The first probability threshold can be set according to the actual situation and actual needs, and the specific value is not limited here. For example, if it is necessary to more strictly identify the target risk power users, a lower first probability threshold can be set.

[0092] In practical applications, the arrears behavior pattern of target power users may be stable within a certain period of time. Considering this possible pattern, the operation of generating the credit risk probability value for target power users through a credit evaluation model can be performed periodically. For example, the credit risk of target power users can be predicted periodically on a monthly basis. As an example, in January 2025, through the analysis of the credit evaluation model, the credit risk probability value of this target power user is higher than the first probability threshold, so it is determined as a target risk power user in January 2025; while in February 2025, through the analysis of the credit evaluation model, the credit risk probability of this target power user is lower than the first probability threshold, so it is determined as a non-target risk power user in February 2025.

[0093] S102. Based on the characteristic data, account balance, and tiered electricity price usage information of the target risk power user, predict the arrears probability distribution and payment probability distribution of the target risk power user through a mixture density network.

[0094] On the premise that a certain power user has been determined as a target risk power user, in the embodiments of the present application, step S102 is continued to predict the arrears probability distribution and payment probability distribution of this user. Before prediction, obtain the characteristic data, account balance, and tiered electricity price usage information of this power user as the data basis for prediction.

[0095] Here, the characteristic data of the target risk power user at least includes: the arrears behavior characteristics and payment behavior characteristics of the target risk power user.

[0096] In a possible implementation manner, the present application pre-constructs two prediction models based on mixture density networks (MDN), one as an arrears prediction model and the other as a payment prediction model. MDN is a deep learning architecture that combines a neural network and a probability density model. Different from traditional neural networks, the output of MDN is not a single predicted value, but the entire probability distribution. Therefore, MDN can provide a richer output for continuous variable prediction problems because MDN not only predicts the target value, but also predicts the probability distribution of the target value, thereby introducing the uncertainty of the model.

[0097] The core idea of MDN is to represent the probability density of the target value as a linear combination of multiple kernel functions, where each kernel function usually follows a Gaussian distribution. The Gaussian distribution has the following characteristics: the total area under the distribution is 1; the probability of obtaining an expected value within a certain range can be obtained; it extends to positive and negative infinity on both sides; each Gaussian distribution is determined by two parameters, the mean and the variance. The parameters (mean, variance) of these Gaussian distributions and the mixing coefficients (weights of each Gaussian distribution) are learned by a neural network and optimized using the maximum likelihood estimation as the loss function. A probability distribution is a mathematical tool that describes the possible values of a random variable and their corresponding probabilities, and it can help people understand the uncertainty and variability of data. The Gaussian distribution, also known as the normal distribution, is a common continuous probability distribution that can describe random variables in many natural and social phenomena. The mixture Gaussian distribution is a composite distribution composed of multiple Gaussian distributions. It superimposes different Gaussian distributions through weights to more flexibly fit the distribution characteristics of complex data, especially showing advantages when dealing with multimodal distributions. The mixture Gaussian distribution can theoretically approximate any probability distribution by combining multiple Gaussian probability distributions. Therefore, in the mixture Gaussian distribution, each distribution is determined by three parameters: weight, mean, and variance.

[0098] In this way, the overdue payment prediction model and the payment prediction model constructed in the embodiments of the present application can take into account the uncertainty and complexity of users' overdue payment behaviors and payment behaviors on the time scale. Through the overdue payment probability distribution output by the overdue payment prediction model and the payment probability distribution output by the payment prediction model, where the overdue payment probability distribution reflects the change of the overdue payment probability over time, and the payment probability distribution reflects the change of the payment probability over time, the fluctuations of the overdue payment probability and the payment probability can be grasped from the time scale, thereby improving the accuracy and flexibility of the prediction.

[0099] In an optional implementation manner, the structure of the mixture density network used to construct the overdue payment prediction model and the payment prediction model includes: an input layer, a hidden layer, and an output layer. Figure 2 shows the structure of a mixture density network. In Figure 2 the example, the mixture density network is specifically configured with 4 hidden layers, namely the first hidden layer, the second hidden layer, the third hidden layer, and the fourth hidden layer arranged in sequence along the input-to-output direction.

[0100] Among them, the input layer inputs the data that has been pre-trained.

[0101] The structure of each hidden layer includes a feed-forward fully connected layer, a batch normalization layer, and an activation layer. The structure of a single hidden layer is as Figure 3As shown in the figure. The hidden layer is used to increase the complexity of the network. The number of neurons in each feed-forward fully-connected layer of multiple hidden layers is halved layer by layer. The feed-forward fully-connected layer can enhance the ability of the model. The batch normalization layer accelerates the learning convergence speed, prevents overfitting, and avoids the vanishing gradient. The activation layer linearizes the multi-layer network so that the multi-layers of the network are meaningful.

[0102] The output layer has multiple sets of parameters (mean, variance) of the Gaussian distribution and mixing coefficients (weights of each Gaussian distribution). Multiple Gaussian distributions are combined into a mixture Gaussian distribution. Figure 4 shows another structure of the mixture density network, different from Figure 2 the structure shown in the figure, showing modules named Linear_mu, Linear_Sigma, and Linear_pi at the end, representing the mean, variance of the Gaussian distribution, and the weights of multiple Gaussian distributions respectively.

[0103] As an example, in the MDNs adopted by the overdue payment prediction model and the payment prediction model in the embodiments of the present application, both include Figure 4 the 4 hidden layers shown in the figure, namely the first hidden layer to the fourth hidden layer. Among them, the first hidden layer includes 200 neurons; the second hidden layer includes 100 neurons; the third hidden layer includes 50 neurons; the fourth hidden layer includes 25 neurons. There are 5 Gaussian distributions mixed in the MDN.

[0104] Figure 5 This is an architecture diagram for predicting the overdue payment probability distribution and the payment probability distribution through a mixture density network provided by the embodiments of the present application. As Figure 5 shown in the figure, the characteristic data of the target risk power users, data such as account balance and ladder electricity price usage information are used as input data, and enter the overdue payment prediction model and the payment prediction model respectively. The overdue payment prediction model outputs the overdue payment probability distribution of the target risk power users, and the payment prediction model outputs the payment probability distribution of the target risk power users. That is to say, the overdue payment prediction model and the payment prediction model share the same input data.

[0105] The overdue payment probability distribution and the payment probability distribution can be displayed in the form of a curve. In the displayed distribution curve, the horizontal axis represents time and the vertical axis represents probability. Regarding the network structure in the overdue payment prediction model and the payment prediction model, reference can be made to Figures 2 to 4 the relevant introduction, which will not be elaborated here.

[0106] S103. Determine the overdue payment risk warning time node corresponding to the target risk power users based on the overdue payment probability distribution, and generate a collection reminder strategy corresponding to the target risk power users based on the payment probability distribution.

[0107] Since the overdue payment probability distribution of the target-risk power users reflects the change of the overdue payment probability over time, the time node with the highest overdue payment probability can be identified based on the overdue payment probability distribution. Since the overdue payment probability distribution is analyzed and predicted based on the characteristic data of the target-risk power users themselves, account balances, and tiered electricity price usage information, the overdue payment probability distribution effectively integrates the characteristics of the target-risk power users themselves, the real-time account situation, and the electricity usage process. Thus, the overdue payment probability distribution can more accurately reflect the time node with the highest overdue payment probability of the target-risk power users on the time scale. This time node can effectively assist the power supply company in managing the overdue payment risk of the target-risk power users. In this application, this time node can be referred to as the overdue payment risk warning time node corresponding to the target-risk power users. For example, if it is predicted that the overdue payment risk probability of the target-risk power users reaches the highest value on March 16, 2025, then March 16 is regarded as the overdue payment risk warning time node. In the above embodiments, the time node when the overdue payment probability reaches the highest point is used as the overdue payment risk warning time node. In other embodiments, the time node when the overdue payment probability first exceeds the preset probability can also be used as the overdue payment risk warning time node. For example, it is predicted that the target-risk power users are at a relatively high overdue payment probability level from March 16 to March 19. Among them, March 16 is the first time node when the overdue payment probability exceeds the preset probability during this period, so March 16 is used as the overdue payment risk warning time node corresponding to the target-risk power users.

[0108] Similarly, since the payment probability distribution of the target-risk power users reflects the change of the payment probability over time, the payment behavior pattern of the target-risk users can be identified based on the payment probability distribution, which can also be understood as identifying the payment rules of the target-risk users. Taking this payment probability distribution as the data basis, it is convenient to generate a targeted collection reminder strategy for the target-risk power users. Thus, before the target-risk power users are overdue, personalized and effective collection reminders can be sent to them through appropriate collection reminder strategies.

[0109] S104. Before the time reaches the overdue payment risk warning time node, send a reminder for electricity bill collection to the target-risk power users based on the collection reminder strategy.

[0110] The following are several examples:

[0111] Example 1: Through model prediction, it is found that user A has a high payment probability in the morning on specific dates at the beginning or middle of the month. Then, an electricity bill collection text message can be sent in the morning on similar dates before the predicted balance is exhausted (before the time reaches the overdue payment risk warning time node). Thus, after user A sees the text message, there is a high probability of paying the bill in time.

[0112] Example 2: Through model prediction, it is found that User B has a high probability of paying bills at night on weekends. Then, electricity bill reminder messages can be sent in the afternoon on similar dates before the predicted balance is exhausted (before the time reaches the overdue risk warning time node). In this way, after User B sees the message, there is a high probability of paying the bill in time.

[0113] Example 3: User C does not have obvious and specific periodic patterns similar to User A or User B, but there is still an implicit payment logic that can be captured by the model. The model will generate a payment probability distribution. The power supply company can choose to conduct electricity bill reminders on the 1-2 days with the highest payment probability for User C. In this way, after User C sees the message, there is a high probability of paying the bill in time. For example, if there are multiple peaks in the predicted payment probability distribution, a time chain for electricity bill reminders can be formed based on the combination of these multiple peaks, and a reminder will be sent once when the time arrives and the bill has not been paid.

[0114] In the method embodiments introduced above, based on the characteristic data, account balance, and ladder electricity price usage information of target risk power users, personalized analysis of users can be realized, which is closer to user behavior more flexibly. In addition, with the help of the mixture density network, the changes in the overdue probability and payment probability can be predicted on the time scale, which is conducive to more accurately judging the overdue risk warning time node and generating personalized collection reminder strategies. Furthermore, more refined overdue risk control can be achieved, and the control effectiveness of overdue risk can be improved.

[0115] From the above analysis, it is not difficult to see that in the embodiments of this application, the electricity bill collection warning mechanism for target risk power users can be automatically triggered, and the predicted overdue probability distribution and payment probability distribution are of great significance for the power supply company to formulate collection reminder strategies, bringing great convenience to the development of power marketing business.

[0116] In the above introduction, the characteristic data of power users was mentioned, and the characteristic data of power users is also part of the input data of the MDN. In practical applications, the characteristic data input into the MDN can be screened. For example, collect the original data of users, clean the data after feature extraction, and then use the SelectKBest method for feature screening, and input the screened characteristic data into the MDN. Similarly, before training the overdue prediction model and payment prediction model, feature screening also needs to be performed through a similar feature screening method.

[0117] For example, the overdue payment risk control method further includes: screening out multiple features from the historical feature dataset of power users through the SelectKBest method as historical training features; constructing an overdue payment prediction model based on the mixture density network using the historical training features of power users and the overdue payment date information; and constructing a payment prediction model based on the mixture density network based on the historical training features of power users and the payment date information. Among them, the historical feature dataset contains the original data of some historical periods of power users.

[0118] Exemplarily, the data of power users in a certain power consumption area from January 2022 to March 2024 are selected, including original data such as user profile tables, card meter profile tables, ladder usage situation tables, SMS information tables, fee control balance snapshots, receivable tables, overdue payment reminder tables, work order tables, user electricity price information, subsistence allowance and five-guarantee information, user warning thresholds, etc., features are extracted, and multiple feature groups are constructed. These feature groups are respectively: customer profile feature group, power consumption behavior feature group, payment behavior feature group, overdue payment reminder behavior feature group, overdue payment behavior feature group, as Figure 6 shown in. Next, feature data cleaning is performed on the feature groups in all aspects to determine the absence of overdue payment reminder behavior, the absence of payment behavior, abnormal feature fluctuations, missing marked data, etc. Based on the actual environment, in the embodiments of the present application, the information of using SMS is selected as the data marking source. The specific logic is: starting from a balance warning message until the balance is zero (i.e., there is an overdue payment behavior) or a payment is made (i.e., there is no overdue payment behavior), it is used as the basis for user credit rating and marked respectively. For example, through the above marking method, the user feature data is marked as two categories: high risk and low risk. Finally, about 12.62 million groups of samples are obtained for model training, and the total number of features is 380. Some of the features are shown in Table 1.

[0119] Table 1

[0120] Feature identifier Feature name AC_HALF_YEAR Number of reminders for negative balance in the last six months TODAY_BALANCE Controlled-fee balance PC_TWO_YEAR Payment count in the last two years AC_ONE_YEAR Number of reminders for negative balance in the last year ZERO_WPR_HALF_YEAR Proportion of reminder for overdue text messages in the previous payment in the last six months PAYMENT_INTERVAL_HALF_YEAR Average payment interval in the last six months (days) PAYMENT_HALF_YEAR_MAX Maximum payment amount in six months

[0121] SelectKBest is a feature selection function in the scikit-learn library, which is used to select the k best features from the dataset. It can select and rank features according to the given evaluation function and score. By performing feature screening through the SelectKBest method, the data dimension of the dataset can be reduced, only the most important features are retained, and in addition, the importance ranking of each feature can be obtained, which is convenient for further analysis and use. As an example, considering eliminating interference and optimizing performance, in the present application, the SelectKBest algorithm can be used to screen out 120 relatively important features from all 380 features for the training of the MDN-based models (i.e., the overdue payment prediction model and the payment prediction model).

[0122] When it is necessary to predict target risk power users, based on the characteristic data, account balance and tiered electricity price usage information of the target risk power users, the overdue probability distribution and payment probability distribution of the target risk power users are predicted through a mixture density network, which may specifically include:

[0123] Input the characteristic data, account balance and tiered electricity price usage information of the target risk power users screened by the SelectKBest method into the overdue payment prediction model to obtain the overdue payment probability distribution output by the overdue payment prediction model; and input the characteristic data, account balance and tiered electricity price usage information of the target risk power users screened by the SelectKBest method into the payment prediction model to obtain the payment probability distribution output by the payment prediction model.

[0124] That is to say, in the early stage, in order to train the model to construct training samples and actually use the model for prediction, the SelectKBest method can be used to screen the characteristic data. The use of the SelectKBest method effectively reduces the redundant large number of characteristic dimensions, and combines the importance of the characteristics to assist in establishing a more accurate prediction model and assist in realizing a more accurate prediction of the overdue payment probability distribution and payment probability distribution.

[0125] As introduced in the previous embodiments, in an optional implementation manner, determining the target risk power users includes: generating a credit risk probability value of the target power user based on the characteristic data of the target power user; if the credit risk probability value is higher than the first probability threshold, determining the target power user as the target risk power user. For the generated credit risk probability value or related conclusion, in the embodiments of the present application, the importance of each feature can be further characterized by the shap algorithm, so as to reasonably explain the contribution (or importance) of each feature to the credit risk probability value or related conclusion.

[0126] Specifically, after determining that the target power user is the target risk power user, the overdue risk control method further includes: using the shap algorithm to calculate the characteristic importance characterization value of each feature in the characteristic data of the target power user for the credit risk probability value. The shap algorithm can explain the prediction result, and can obtain the contribution degree information of each feature to the prediction result of each power user, which solves the problem that it is difficult for business personnel to understand machine learning algorithms to a certain extent, helps business personnel quickly understand, and efficiently carry out work.

[0127] The core of the shap algorithm is based on the Shapley value, which is a distribution method in game theory used to measure the marginal contribution of each participant in a cooperative game. For a machine learning model, the Shapley value helps measure the contribution of each feature to the model prediction result.

[0128] Specifically, for a given feature, the Shapley value is calculated by considering the average change (weighted average) in the model prediction for that feature across all possible subsets of features, as per the following calculation formula:

[0129]

[0130] The coefficient consisting of three factorials in formula (5) represents the number of times different coalitions appear, that is, the weights corresponding to different coalitions. Simply put, the Shapley value measures the change in the model prediction value when a certain feature is added to a group of other features. This change is relative to the baseline value, which is usually the average prediction value across the entire dataset.

[0131] Suppose there is a model for predicting apartment prices and there are four features: park, size, floor, and cat. If we want to know how the feature cat = banned affects the price prediction of a specific apartment, we need to calculate its marginal contribution in all possible coalitions. For example, if two features park = nearby and size = 50 are combined, that is, when the coalition is formed, the predicted price is 320,000 euros, and it becomes 310,000 euros after adding cat = banned, then the marginal contribution of cat = banned in this coalition (park = nearby and size = 50) is -10,000 euros. This process needs to be repeated for all possible coalitions, and the weighted average of all marginal contributions is taken as the final Shapley value of this feature.

[0132] In the embodiments of this application, after calculating the feature importance characterization values of each feature in the feature data of the target power user through the shap algorithm, the feature data of the target power user can also be sorted according to the feature importance characterization values, and a list of the feature sorting results can be displayed; and / or, determine the top N features with the largest feature importance characterization values in the feature data of the target power user, and display the determined top N features and the corresponding feature importance characterization values; and / or, determine multiple features in the feature data of the target power user whose feature importance characterization values are greater than a preset threshold, and display the multiple features and the corresponding feature importance characterization values. Among them, the sorting can be performed in descending order or ascending order. The sorting rule is not limited here. In addition, the way to display the determined features and the corresponding feature importance characterization values can be in graphical form or in list form. The display method can be the default setting or user-defined setting.

[0133] Whether it is to display the feature importance characterization value (Shapley value) of all feature data involved in the calculation, or to display the top N features with the largest feature importance characterization values, or to display multiple features whose feature importance characterization values exceed a preset threshold, it can intuitively reflect to business personnel the importance of features to the prediction result, so as to facilitate business personnel to intuitively understand the prediction result of the LightGBM algorithm and trust the prediction result. It can be seen that the embodiment of the present application uses the shap algorithm to realize the interpretation of the prediction result. Figure 7 It is an effect diagram showing the feature importance result of a certain user. In the left area of the figure, the feature ranking result list corresponding to the user is shown, and in the right area, 20 features with the top 20 feature importance characterization values and their corresponding feature importance characterization values are shown in the form of a bar chart.

[0134] In the embodiment of the present application, the feature importance characterization value of each feature in the feature data of the target power user is calculated through the shap algorithm, and it can be further applied to determine the overdue risk warning time node, or applied to generate a collection reminder strategy, or applied to formulate a credit risk review threshold for electricity bill collection.

[0135] The above several scenarios can be implemented on the premise that the determined target power user is a target risk power user and after determining the target overdue risk type corresponding to the target power user. The implementation method for determining the target overdue risk type corresponding to the target power user will be described below. In the embodiment of the present application, the subcategories of the user group can be established first according to the feature ranking result lists corresponding to multiple different power users; each subcategory corresponds to an overdue risk type. Then, based on the feature ranking result list corresponding to the target power user, identify the subcategory corresponding to the target power user and determine the target overdue risk type corresponding to the target power user. That is to say, under the condition of knowing the feature ranking result lists corresponding to multiple different power users, it is possible to clarify which features have a higher importance for the prediction of high overdue risk of power users based on the ranking results, and then divide the user group into subcategories. The subcategory corresponds to the overdue risk type. For example, the overdue risk type corresponding to category 1 is: poor payment habit; the overdue risk type corresponding to category 2 is weak financial stability. Through the above method, multiple subcategories are established. Then, when the feature ranking result list corresponding to the target power user is known, the target power user can be directly divided into its corresponding subcategory based on the first few features with higher feature importance characterization values in its feature ranking result list, and based on the one-to-one correspondence between the subcategory and the overdue risk type, further determine the target overdue risk type corresponding to the target power user.

[0136] Further, in combination with the determined target risk overdue payment types, the embodiments of the present application can design differentiated services or intervention measures accordingly. For example:

[0137] In an alternative implementation, determining the overdue payment risk warning time node corresponding to the target risk power user based on the overdue payment probability distribution includes:

[0138] Determining the overdue payment risk warning time node of the target power user based on the overdue payment probability distribution of the target power user and the target overdue payment risk type corresponding to the target power user.

[0139] In the previous embodiments, it is proposed that the overdue payment risk warning time of the target power user can be determined based on the overdue payment probability distribution predicted by the overdue payment prediction model. In the embodiments of the present application, it is further emphasized that after the target overdue payment risk type corresponding to the target power user has been determined, the target overdue payment risk type and the above-mentioned overdue payment probability distribution can be combined to determine the overdue payment risk warning time node of the target power user. For example:

[0140] Based on the overdue payment probability distribution, the preliminary warning time node is determined as t1 (for example, the overdue payment probability at t1 reaches the peak), but since the target risk overdue payment type is poor payment habit, the determined preliminary warning time node is advanced by 5 days, and the warning time node t2 is determined as the overdue payment risk warning time node of the target power user.

[0141] As introduced in the previous embodiments, the collection reminder strategy corresponding to the target risk power user can be generated according to the payment probability distribution predicted by the payment prediction model. In practical applications, on the basis of having determined the target overdue payment risk type corresponding to the target power user, the target overdue payment risk type and the above-mentioned payment probability distribution can also be combined to generate the collection reminder strategy corresponding to the target power user. For example:

[0142] The target power user has a high probability of paying the bill in the morning on specific dates at the beginning or middle of the month, and its corresponding target overdue payment risk type is weak financial stability. After analyzing and comparing the financial stability at the beginning and middle of the month, a date period with higher stability can be determined before its balance is exhausted to send an electricity bill collection reminder text message. In this way, there is a high probability that the user will pay the bill in time.

[0143] In an alternative implementation, the overdue payment risk control method may further include:

[0144] Adjusting the first probability threshold based on the target overdue payment risk type corresponding to the target power user to obtain a second probability threshold. The second probability threshold is used to review whether to give an electricity bill collection reminder to the user.

[0145] As mentioned above, the first probability threshold is used to evaluate the level of the credit risk probability value of electricity users. If the credit risk probability value of a user is higher than the first probability threshold, it is considered that the credit risk is relatively high, and the user is identified as a target-risk electricity user; conversely, if the credit risk probability value is less than or equal to the first probability threshold, it is considered that the credit risk does not reach a relatively high level and does not belong to the target-risk electricity user.

[0146] If the target overdue risk type corresponding to the target electricity user is a specific overdue risk type, it is considered that the probability threshold needs to be adjusted, such as increasing or decreasing the probability threshold, and re-determining whether its risk attribute has reached the level that requires collection reminder. Therefore, as mentioned in the above introduction, before the time reaches the overdue risk warning time node, based on the collection reminder strategy, the electricity bills of target-risk electricity users are collected and reminded, specifically including:

[0147] If the credit risk probability value of the target electricity user is higher than the second probability threshold, before the time reaches the overdue risk warning time node, based on the collection reminder strategy, the electricity bills of the target-risk electricity users are collected and reminded.

[0148] If the credit risk probability value is higher than the second probability threshold, it is considered that after review, the risk of this target electricity user is still relatively high and a collection reminder for the electricity bill is required. Through the review means, the problem of overly high reminder frequency for users, which affects the user's SMS sending and receiving experience, is avoided. And through the review means, the overdue risk of users is double-checked, and the collection and control of fees based on this is more theoretically grounded and can more effectively achieve the overdue control of people with a relatively high credit risk level.

[0149] The above technical solution, by studying the characteristic importance representation values of each characteristic in the characteristic data of each user for the credit risk probability value, not only realizes the interpretation of the prediction result, but also can further identify the overdue risk type of the user based on the characteristic importance representation value, guiding the implementation of overdue control measures. This personalized and differentiated electricity bill collection method formulated in combination with the characteristics of users themselves effectively improves the problems faced by the existing technology.

[0150] Figure 8 This is the overall flowchart of an overdue risk control method provided by an embodiment of the present application. In Figure 8 the preprocessing of the characteristic data is shown, such as characteristic data cleaning, screening, etc. In addition, the evaluation of the user's credit risk probability using the LightGBM algorithm is also shown. Furthermore, the prediction of the overdue probability and payment probability of high-risk users is realized through a prediction model based on MDN, the determination of the overdue risk warning time node and the generation of the collection reminder strategy are realized, and finally the process of fee collection is carried out. From Figure 8It can be seen that the feature importance characterization values calculated based on the SHAP algorithm can assist in determining the overdue risk warning time node and / or generating the collection reminder strategy.

[0151] As previously introduced, in this application, the main technical means for realizing the overdue probability distribution prediction and the payment probability distribution prediction is the prediction model based on MDN, that is, the overdue prediction model and the payment prediction model. Next, the loss function of the above prediction model will be introduced.

[0152] Different from the traditional loss function, when training the prediction model in this application, since the actual probability distribution is unknown, only the actual probability can be sampled, and the likelihood value of the predicted probability distribution and the actual sample data is calculated. The likelihood function is shown by the following formula:

[0153]

[0154] The meaning of formula (6) is that the likelihood function of parameter θ given the output X is equal to the probability P(X | θ) of variable X after given parameter θ. N represents the number of samples, and i represents the i-th sample among N samples.

[0155] Due to the randomness of the probability distribution itself, a low likelihood value does not mean that the predicted probability distribution is wrong. In this application, the logarithm of the likelihood value is taken to remove the randomness, and the negative value is taken as the loss. Refer to the following formula, the loss function is expressed as:

[0156]

[0157] In order to minimize this loss, the objective function is set, and the function is expressed as follows:

[0158]

[0159] In the embodiment of this application, the parameters used by MDN are shown in Table 2. Among them, the model parameters include: the number of neural network layers, the number of neurons in each layer, the activation function, and the number of mixture Gaussian distributions. The training parameters include: the number of training batches and the learning rate. The functions of each parameter are shown in the rightmost column of Table 2.

[0160] Table 2

[0161]

[0162]

[0163] Through actual tests, the parameter values are gradually confirmed. During the experiment, various parameter combinations have been tried. Currently, the parameter combination with relatively better performance has the following parameter values: 200 neurons in the first hidden layer, 100 neurons in the second hidden layer, 50 neurons in the third hidden layer, 25 neurons in the fourth hidden layer, and the number of Gaussian mixture distributions is 5.

[0164] In the embodiments of the present application, during the training process of the overdue payment prediction model and the payment prediction model, the range of the negative log-likelihood loss (as shown in formula (7)) is from 0 to +∞, which cannot be converted into a percentage. Therefore, it is not easy to intuitively evaluate and compare this loss. To solve this loss comparison problem, in the embodiments of the present application, for the negative log-likelihood loss of the model, the loss is transformed into the interval [0, 1] with a preset value as the loss upper limit. The value within the [0, 1] interval obtained after transformation is used as an evaluation index for model evaluation, that is, the accuracy index. As an example, in the present application, 20 is used as the preset value because it is found through testing that the loss of the MDN after parameter initialization is approximately 20, so 20 is used as the loss upper limit. The expression of the accuracy index Accuracy is as follows:

[0165]

[0166] In addition, in addition to the accuracy index Accuracy, the present application also proposes to use the mean absolute error index MAE as another evaluation index for the model, and its expression is as follows:

[0167]

[0168] The mean absolute error index is obtained by calculating the mean absolute error between the expected value of the predicted probability distribution and the actual sample label. In formula (10), N represents the number of samples, i represents the i-th sample, y i represents the actual sample label of the i-th sample, represents the expected value of the predicted probability distribution. Taking the overdue payment prediction model as an example, the actual sample label here can refer to the overdue date information. Taking the payment prediction model as an example, the actual sample label here can refer to the payment date information.

[0169] Through the MAE, it is more convenient and intuitive to understand the deviation between the prediction result and the actual result. Using "days" as the deviation unit is more conducive to understanding. That is to say, if the expected value of the probability distribution is determined on a certain day, the time interval with a relatively high probability can be roughly estimated based on the deviation. In practical applications, if the MAE is known, the scheme can be set more effectively when facing the probability distribution.

[0170] Generally speaking, by combining the accuracy metric Accuracy and the mean absolute error metric MAE, in the embodiments of the present application, the usability of the model can be evaluated more intuitively and efficiently, thus facilitating model parameter tuning and result display during the training phase.

[0171] In the technical solution of the present application, through the LightGBM algorithm, a credit evaluation model can be constructed relatively accurately, and based on this, the credit risk probability value of power users can be accurately predicted to indicate the credit risk level (or credit risk grade) of power users. During the model construction process, the historical data from January 2022 to October 2023 was used as the training set, and the historical data from November 2023 to March 2024 was used as the validation set. The prediction results on the validation set are as follows:

[0172] (1) Precision (the proportion of actually high-risk users among the predicted high-risk users): 89.55%.

[0173] (2) Recall (the proportion of correctly predicted high-risk users among all high-risk users): 90.38%.

[0174] Combined with the prediction results on the validation set, it can be seen that the prediction results are relatively good.

[0175] In addition, using the mixture density network MDN, in the present application, the training set and the test set are split at a ratio of 8:2, with 80% of the sample data used for training and 20% of the sample data used for result verification testing. Taking the prediction results in terms of payment as an example, Figure 9 is a schematic diagram of the payment probability distribution for a high-risk power user. Figure 9 represents the probability that the model predicts the user is most likely to make a payment after 12 days, and this probability reaches the highest value. It is also possible on other dates. The expected value y of the final prediction distribution is 12.253264.

[0176] The present application also evaluates the model results from the dimensions of two evaluation metrics (Accuracy and MAE).

[0177] 1) After the features of the training data are input into the model, the output predicted probability density distribution θ is compared with the actual payment / arrears time X. In this process, the present invention transforms the negative log-likelihood loss into the interval [0, 1] for intuitive comparison. The test set results obtained are as follows:

[0178] · Arrears average loss 1.94 Accuracy: 90.30%

[0179] · Payment average loss 1.87 Accuracy: 90.65%

[0180] 2) After the features of the training data are input into the model, the predicted distribution expected value ŷ is compared with the actual payment / arrears time y. Figure 10 This is a schematic diagram of the prediction situation taking the payment probability distribution as an example. The mean absolute error for the payment time in the payment probability distribution is about 2.32 days. And the mean absolute error for the arrears time in the arrears probability distribution is about 2.12 days. The test set results obtained are as follows:

[0181] · Mean absolute error for arrears: 2.13 days

[0182] · Mean absolute error for payment: 2.32 days

[0183] In the embodiments of the present application, starting from the dimensions of two evaluation indicators, good results are obtained through test verification. Accordingly, the arrears risk can be reminded in real time according to the threshold set by the business department, and personalized payment strategies can be formulated based on the probability distribution of payments.

[0184] Generally speaking, in the present application, the classification performance of the LightGBM algorithm and the advantage of the mixture density network MDN that can calculate the probability distribution are effectively utilized. The two algorithms are organically combined to predict the credit risk probability value, the arrears probability distribution, and the payment probability distribution of users in sequence, effectively improving the accuracy of risk assessment. In addition, through the prediction and analysis of the probability distribution of the arrears and payment of electricity users in the time series, this method can not only monitor and warn the arrears risk of users in real time, but also assist in formulating collection reminder strategies based on the payment habits of electricity users.

[0185] In the present application, the shap algorithm is introduced to explain the prediction results, and the contribution degree information of each feature to the prediction results of each electricity user can be obtained, that is, the feature importance characterization value. To a certain extent, it solves the problem that machine learning algorithms are difficult for business personnel to understand. It helps business personnel understand quickly and carry out work efficiently.

[0186] In the present application, a more comprehensive and effective feature design is proposed: the data feature design is more comprehensive, including multiple feature groups such as customer plan feature groups, electricity consumption behavior feature groups, payment behavior feature groups, collection behavior feature groups, and arrears behavior feature groups, with a total of 380 different features, which can describe users more comprehensively and effectively, so as to obtain better model training results.

[0187] The present application also proposes an innovative method to solve the problem of the output range of the MDN negative log-likelihood loss function: The output range of the traditional negative log-likelihood loss function is from 0 to positive infinity, which makes it difficult to intuitively evaluate and compare loss values. To overcome this challenge, based on the experimental results, the present application sets 20 as the upper limit of the loss, and accordingly converts the loss value into a value within the range of [0,1]. This conversion not only standardizes the range of the loss value, but also facilitates intuitive understanding and comparison of the performance of different models. Thus, the present invention can more effectively monitor and optimize the model training process, thereby improving the efficiency and effect of model training.

[0188] Based on the overdue fee risk control method introduced in the foregoing embodiments, correspondingly, the present application also provides an overdue fee risk control device. As Figure 11 is a schematic structural diagram of an overdue fee risk control device provided by an embodiment of the present application. Combining Figure 11 , the overdue fee risk control device includes:

[0189] A risk user determination module 1101, configured to determine a target risk power user;

[0190] A probability distribution prediction module 1102, configured to predict the overdue fee probability distribution and payment probability distribution of the target risk power user based on the characteristic data, account balance, and ladder electricity price usage information of the target risk power user through a mixture density network; wherein, the overdue fee probability distribution reflects the change of the overdue fee probability over time, and the payment probability distribution reflects the change of the payment probability over time; the characteristic data of the target risk power user at least includes: the overdue behavior characteristics and payment behavior characteristics of the target risk power user;

[0191] A time node determination module 1103, configured to determine the overdue fee risk warning time node corresponding to the target risk power user based on the overdue fee probability distribution;

[0192] A policy generation module 1104, configured to generate a collection reminder policy corresponding to the target risk power user based on the payment probability distribution;

[0193] A reminder module 1105, configured to perform an electricity fee collection reminder on the target risk power user based on the collection reminder policy before the time reaches the overdue fee risk warning time node.

[0194] In an alternative implementation, the risk user determination module 1101 includes:

[0195] A risk probability value generation unit, configured to generate a credit risk probability value of the target power user based on the characteristic data of the target power user;

[0196] A risk user determination unit, configured to determine the target power user as a target risk power user if the credit risk probability value is higher than a first probability threshold;

[0197] The overdue payment risk control device further includes:

[0198] An importance characterization value calculation module, configured to, after determining the target power user as a target risk power user, use the SHAP algorithm to calculate the characteristic importance characterization value of each characteristic in the characteristic data of the target power user for the credit risk probability value.

[0199] In an alternative implementation, the overdue payment risk control device further includes: a display module, and at least one of a sorting module and a determination module;

[0200] The sorting module is configured to sort the characteristic data of the target power user according to the characteristic importance characterization value; the display module is configured to display a list of characteristic sorting results; and / or,

[0201] The determination module is configured to determine the top N characteristics with the largest characteristic importance characterization value in the characteristic data of the target power user; the display module is configured to display the determined top N characteristics and the corresponding characteristic importance characterization values; and / or,

[0202] The determination module is configured to determine multiple characteristics in the characteristic data of the target power user whose characteristic importance characterization value is greater than a preset threshold; the display module is configured to display the multiple characteristics and the corresponding characteristic importance characterization values.

[0203] In an alternative implementation, the overdue payment risk control device further includes:

[0204] A category creation module, configured to establish sub-categories of user groups according to the characteristic sorting result lists corresponding to multiple different power users; each sub-category corresponds to an overdue payment risk type;

[0205] A category identification module, configured to identify the sub-category corresponding to the target power user based on the characteristic sorting result list corresponding to the target power user, and determine the target overdue payment risk type corresponding to the target power user.

[0206] In an alternative implementation, the time node determination module 1103 is specifically configured to:

[0207] Based on the overdue payment probability distribution of the target power user and the target overdue payment risk type corresponding to the target power user, determine the overdue payment risk warning time node of the target power user;

[0208] and / or,

[0209] The policy generation module 1104 is specifically configured to:

[0210] Generate a collection reminder policy corresponding to the target power user based on the payment probability distribution and the target overdue risk type corresponding to the target power user.

[0211] In an alternative implementation, the overdue risk control device further includes:

[0212] A threshold adjustment module, configured to adjust the first probability threshold based on the target overdue risk type corresponding to the target power user to obtain a second probability threshold;

[0213] The reminder module 1105 is specifically configured to:

[0214] If the credit risk probability value of the target power user is higher than the second probability threshold, before the time reaches the overdue risk warning time node, perform a power bill collection reminder on the target risk power user based on the collection reminder policy.

[0215] In an alternative implementation, the overdue risk control device further includes:

[0216] A first model construction module, configured to construct a credit evaluation model using the LightGBM algorithm;

[0217] The risk probability value generation unit is specifically configured to:

[0218] Use the feature data of the target power user as the input of the credit evaluation model, analyze the feature data of the target power user through the credit evaluation model, and output the credit risk probability value of the target power user by the credit evaluation model.

[0219] In an alternative implementation, the overdue risk control device further includes:

[0220] A feature screening module, configured to screen out multiple features as historical training features from the historical feature dataset of power users through the SelectKBest method;

[0221] A second model construction module, configured to construct an overdue prediction model based on a mixture density network using the historical training features of power users and overdue date information;

[0222] A third model construction module, configured to construct a payment prediction model based on a mixture density network based on the historical training features of power users and payment date information;

[0223] The probability distribution prediction module 1102 is specifically configured to:

[0224] Input the feature data, account balance, and tiered electricity price usage information of the target risky electricity users screened by the SelectKBest method into the overdue payment prediction model to obtain the overdue payment probability distribution output by the overdue payment prediction model; and input the feature data, account balance, and tiered electricity price usage information of the target risky electricity users screened by the SelectKBest method into the payment prediction model to obtain the payment probability distribution output by the payment prediction model.

[0225] In an alternative implementation, the model evaluation metrics of the overdue payment prediction model and the model evaluation metrics of the payment prediction model both include: an accuracy metric and a mean absolute error metric;

[0226] The accuracy metric is: a value within the range of [0, 1] obtained by converting the negative log-likelihood loss of the model with a preset value as the loss upper limit;

[0227] The mean absolute error metric is obtained by calculating the mean absolute error between the expected value of the predicted probability distribution and the actual sample label.

[0228] Based on the overdue payment risk control method and the overdue payment risk control device introduced in the foregoing embodiments, correspondingly, an embodiment of the present application further proposes an overdue payment risk control device. This device may specifically include a memory and a processor. Among them, the memory is used to store a computer program; and the processor is used to run the above computer program. When the computer program runs, it executes some or all of the steps of any implementation of the overdue payment risk control method introduced in the method embodiment above.

[0229] It should be noted that the embodiments in this specification are all described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiments. The device and device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated. The components described as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.

[0230] As described above, it is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for controlling arrears risk, characterized in that, Including: Determine target risk power users; Based on the characteristic data, account balance and tiered electricity price usage information of the target risk power users, predict the overdue probability distribution and payment probability distribution of the target risk power users through a mixture density network; wherein, the overdue probability distribution reflects the change of the overdue probability over time, and the payment probability distribution reflects the change of the payment probability over time; the characteristic data of the target risk power users at least includes: the overdue behavior characteristics and payment behavior characteristics of the target risk power users; Based on the overdue probability distribution, determine the overdue risk warning time node corresponding to the target risk power users, and generate a collection reminder strategy corresponding to the target risk power users based on the payment probability distribution; Before the time reaches the overdue risk warning time node, give a reminder for electricity bill collection to the target risk power users based on the collection reminder strategy.

2. The method according to claim 1, characterized in that, The determination of the target risk power users includes: Generate a credit risk probability value of the target power users based on the characteristic data of the target power users; If the credit risk probability value is higher than the first probability threshold, determine that the target power users are target risk power users; After determining that the target power users are target risk power users, the method further includes: Use the SHAP algorithm to calculate the characteristic importance representation values of each characteristic in the characteristic data of the target power users for the credit risk probability value.

3. The method according to claim 2, wherein After using the SHAP algorithm to calculate the characteristic importance representation values of each characteristic in the characteristic data of the target power users for the credit risk probability value, the method further includes: Sort the characteristic data of the target power users according to the characteristic importance representation values, and display the list of characteristic sorting results; and / or, Determine the top N characteristics with the largest characteristic importance representation values in the characteristic data of the target power users, and display the determined top N characteristics and the corresponding characteristic importance representation values; and / or, Determine multiple characteristics in the characteristic data of the target power users whose characteristic importance representation values are greater than the preset threshold, and display the multiple characteristics and the corresponding characteristic importance representation values.

4. The method according to claim 3, characterized in that, The method further includes: Establish sub-categories of the user group according to the list of characteristic sorting results corresponding to multiple different power users; each sub-category corresponds to a type of overdue risk; Based on the list of characteristic sorting results corresponding to the target power users, identify the sub-category corresponding to the target power users, and determine the target overdue risk type corresponding to the target power users.

5. The method according to claim 4, wherein The determination of the overdue risk warning time node corresponding to the target risk power users based on the overdue probability distribution includes: Based on the overdue probability distribution of the target power users and the target overdue risk type corresponding to the target power users, determine the overdue risk warning time node of the target power users; and / or, The generation of the collection reminder strategy corresponding to the target risk power users based on the payment probability distribution includes: Generate a collection reminder strategy for the target power user based on the payment probability distribution and the target overdue risk type corresponding to the target power user.

6. The method according to claim 4, wherein The method further includes: Adjust the first probability threshold based on the target overdue risk type corresponding to the target power user to obtain a second probability threshold; Before the time reaches the overdue risk warning time node, based on the collection reminder strategy, giving a collection reminder for the electricity bill to the target risky power user, includes: If the credit risk probability value of the target power user is higher than the second probability threshold, then before the time reaches the overdue risk warning time node, give a collection reminder for the electricity bill to the target risky power user based on the collection reminder strategy.

7. The method according to claim 2, wherein Before generating the credit risk probability value of the target power user based on the feature data of the target power user, the method further includes: Use the LightGBM algorithm to construct a credit evaluation model; Generating the credit risk probability value of the target power user based on the feature data of the target power user includes: Take the feature data of the target power user as the input of the credit evaluation model, analyze the feature data of the target power user through the credit evaluation model, and output the credit risk probability value of the target power user by the credit evaluation model.

8. The method according to claim 1, characterized in that, It also includes: From the historical feature dataset of power users, use the SelectKBest method to screen out multiple features as historical training features; Use the historical training features of power users and overdue date information to construct an overdue prediction model based on a mixture density network; and, use the historical training features of power users and payment date information to construct a payment prediction model based on a mixture density network; Based on the feature data, account balance, and tiered electricity price usage information of the target risky power user, predict the overdue probability distribution and payment probability distribution of the target risky power user through a mixture density network, including: Input the feature data, account balance, and tiered electricity price usage information of the target risky power user screened by the SelectKBest method into the overdue prediction model to obtain the overdue probability distribution output by the overdue prediction model; and input the feature data, account balance, and tiered electricity price usage information of the target risky power user screened by the SelectKBest method into the payment prediction model to obtain the payment probability distribution output by the payment prediction model.

9. The method according to claim 8, wherein The model evaluation indicators of the overdue prediction model and the model evaluation indicators of the payment prediction model both include: an accuracy rate indicator and a mean absolute error indicator; The accuracy rate indicator is: a value in the range of [0,1] obtained by converting the negative log-likelihood loss of the model with a preset value as the loss upper limit; The mean absolute error indicator is obtained by calculating the mean absolute error between the expected value of the predicted probability distribution and the actual sample label.

10. An overdue payment risk control device, characterized in that, It includes: A risk user determination module for determining the target risky power user; A probability distribution prediction module, configured to predict the overdue probability distribution and payment probability distribution of the target risk power user based on the characteristic data, account balance, and ladder electricity price usage information of the target risk power user through a mixture density network; wherein, the overdue probability distribution reflects the change of the overdue probability over time, and the payment probability distribution reflects the change of the payment probability over time; the characteristic data of the target risk power user at least includes: the overdue behavior characteristics and payment behavior characteristics of the target risk power user; A time node determination module, configured to determine the overdue risk warning time node corresponding to the target risk power user based on the overdue probability distribution; A strategy generation module, configured to generate a collection reminder strategy corresponding to the target risk power user based on the payment probability distribution; A reminder module, configured to give an electricity bill collection reminder to the target risk power user based on the collection reminder strategy before the time reaches the overdue risk warning time node.

11. An overdue payment risk control device, characterized in that, including: a memory and a processor; the memory is used to store a computer program; the processor is used to run the computer program, and when the computer program runs, it executes the steps of the method according to any one of claims 1-9.

Citation Information

Cited By

  • Heating defaulting user risk early warning method and system based on user portrait

    CN121639246A