Electricity charge risk assessment method and device, computer equipment and readable storage medium
Through the combination of decision tree and K-means clustering algorithm, the problem of poor overfitting and generalization capabilities in electricity bill risk assessment is solved, and more accurate risk assessment and risk prediction are achieved, supporting the risk prevention and control of power supply enterprises.
Patent Information
- Application Number
- CN202510427178.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-22
AI Technical Summary
The existing electricity bill risk assessment methods have problems with poor overfitting and generalization capabilities, resulting in inaccurate risk assessment results.
The decision tree model is used to process electricity bill data, determine the user's risk score through the user's identification probability result value, and divide the user into different risk levels using the clustering algorithm, combining the decision tree and the K-means clustering algorithm to improve the evaluation accuracy.
It improves the accuracy of electricity bill risk assessment, can better capture complex patterns in the data, provide power supply companies with valuable risk information, help timely discover and warning of electricity bill risks, reduce electricity bill risks, and ensure economic benefits.
Smart Images

Figure CN120355228A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power technologies, and particularly to a method, device, computer device, computer-readable storage medium, and computer program product for electricity bill risk assessment. Background Art
[0002] With the continuous development and complexity of the power market, the potential risks therein are increasing day by day. At present, although there are some electricity bill risk assessment methods, most of them adopt a single algorithm, such as logistic regression. When dealing with complex electricity bill risk data, these methods may have problems such as overfitting and poor generalization ability, resulting in inaccurate risk assessment results. Summary of the Invention
[0003] Based on this, in view of the technical problems that the above-mentioned risk assessment methods may have problems such as overfitting and poor generalization ability, resulting in inaccurate risk assessment results, it is necessary to provide a method, device, computer device, computer-readable storage medium, and computer program product for electricity bill risk assessment.
[0004] In a first aspect, this application provides a method for electricity bill risk assessment. The method includes:
[0005] Obtain electricity bill data of multiple users to be evaluated;
[0006] Process the electricity bill data of each user to be evaluated through a decision tree model to obtain a user identification probability result value for each user to be evaluated;
[0007] Determine the user risk score for each user to be evaluated according to the user identification probability result value;
[0008] Perform clustering processing on each user risk score to obtain multiple clustering clusters;
[0009] Determine the risk level corresponding to each clustering cluster as the risk level of the users to be evaluated within the clustering cluster.
[0010] In one embodiment, before the step of processing the electricity bill data of each user to be evaluated through a decision tree model to obtain a user identification probability result value for each user to be evaluated, the method further includes:
[0011] Obtain sample electricity bill data; the sample electricity bill data includes time series of multiple electricity bill features;
[0012] Calculate the variance for the time series of each electricity bill feature respectively, and screen out target electricity bill data from the sample electricity bill data based on the variance;
[0013] Construct a decision tree model based on the target electricity bill data.
[0014] In one embodiment, screening the target electricity charge data from the sample electricity charge data based on the variance includes:
[0015] Screening out target features with variances greater than a threshold from the multiple electricity charge features;
[0016] Taking the electricity charge data corresponding to the target features in the sample electricity charge data as the target electricity charge data;
[0017] Constructing a decision tree model based on the target electricity charge data includes:
[0018] Using the target features as splitting features to construct a decision tree model based on the target electricity charge data.
[0019] In one embodiment, obtaining the user risk scores of each of the to-be-evaluated users according to the user recognition probability result values includes:
[0020] For each to-be-evaluated user, performing a mapping process on the user recognition probability result value of the to-be-evaluated user to obtain the user risk score of the to-be-evaluated user.
[0021] In one embodiment, clustering the user risk scores to obtain multiple clustering clusters includes:
[0022] Determining the initial number of clustering clusters;
[0023] Based on the initial number of clustering clusters, clustering the user risk scores to obtain multiple initial clustering clusters; respectively obtaining the sum of the squares of the distances from the non-central points in each initial clustering cluster to the clustering center, and adding the sums of squares corresponding to each initial clustering cluster to obtain the total sum of squares;
[0024] Increasing the initial number of clustering clusters to obtain a new number of clustering clusters, and returning to the step of clustering the user risk scores of the to-be-evaluated users based on the initial number of clustering clusters until a preset end condition is reached to obtain the total sum of squares under multiple numbers of clustering clusters;
[0025] According to the change trend of the total sum of squares under each number of clustering clusters, determining the target number of clustering clusters, and taking the multiple clustering clusters obtained by clustering the target number of clustering clusters as the clustering results of each of the to-be-evaluated users.
[0026] In one embodiment, determining the risk level corresponding to each clustering cluster includes:
[0027] Determining the clustering center of each clustering cluster;
[0028] Obtain the mean value of the distances from the non-central points in each of the clustering clusters to the cluster center respectively, sort the mean values corresponding to each of the clustering clusters, and determine the risk levels corresponding to each of the clustering clusters based on the sorting result.
[0029] In a second aspect, the present application also provides an electricity fee risk assessment device. The device includes:
[0030] A data acquisition module, configured to acquire electricity fee data of multiple users to be evaluated;
[0031] A decision tree processing module, configured to process the electricity fee data of each of the users to be evaluated through a decision tree model respectively to obtain the user identification probability result values of each of the users to be evaluated;
[0032] A score determination module, configured to determine the user risk scores of each of the users to be evaluated according to the user identification probability result values;
[0033] A score clustering module, configured to perform clustering processing on each of the user risk scores to obtain multiple clustering clusters;
[0034] A level determination module, configured to determine the risk level corresponding to each of the clustering clusters as the risk level of the users to be evaluated within the clustering cluster.
[0035] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0036] Acquire electricity fee data of multiple users to be evaluated;
[0037] Process the electricity fee data of each of the users to be evaluated through a decision tree model respectively to obtain the user identification probability result values of each of the users to be evaluated;
[0038] Determine the user risk scores of each of the users to be evaluated according to the user identification probability result values;
[0039] Perform clustering processing on each of the user risk scores to obtain multiple clustering clusters;
[0040] Determine the risk level corresponding to each of the clustering clusters as the risk level of the users to be evaluated within the clustering cluster.
[0041] In a fourth aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0042] Acquire electricity fee data of multiple users to be evaluated;
[0043] Process the electricity bill data of each of the to-be-evaluated users through a decision tree model to obtain the user identification probability result values of each of the to-be-evaluated users;
[0044] Determine the user risk scores of each of the to-be-evaluated users according to the user identification probability result values;
[0045] Perform clustering processing on each of the user risk scores to obtain multiple clustering clusters;
[0046] Determine the risk level corresponding to each of the clustering clusters as the risk level of the to-be-evaluated users within the clustering cluster.
[0047] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0048] Obtain the electricity bill data of multiple to-be-evaluated users;
[0049] Process the electricity bill data of each of the to-be-evaluated users through a decision tree model to obtain the user identification probability result values of each of the to-be-evaluated users;
[0050] Determine the user risk scores of each of the to-be-evaluated users according to the user identification probability result values;
[0051] Perform clustering processing on each of the user risk scores to obtain multiple clustering clusters;
[0052] Determine the risk level corresponding to each of the clustering clusters as the risk level of the to-be-evaluated users within the clustering cluster.
[0053] The above electricity bill risk assessment method, device, computer equipment, storage medium and computer program product obtain the electricity bill data of multiple users to be evaluated; process the electricity bill data through a decision tree model respectively to obtain the user identification probability result values of each user to be evaluated; determine the user risk scores of each user to be evaluated according to the user identification probability result values; perform clustering processing on each user risk score to obtain multiple clustering clusters; determine the risk level corresponding to each clustering cluster as the risk level of the users to be evaluated within the clustering cluster. By combining the use of a decision tree and a clustering algorithm, this method can make full use of the advantages of the two algorithms, overcome the limitations of a single algorithm, improve the accuracy of electricity bill risk assessment, effectively solve the complex multi-scenario problems of the power grid, and also predict the electricity bill risk status and potential risks of the corresponding users. Moreover, the clustering algorithm can perform further clustering analysis on the output results of the decision tree model, discover the potential patterns and structures in the data, and provide more valuable risk information for power supply enterprises; thus overcoming the certain limitations of traditional algorithms in electricity bill risk assessment and the defect that it is difficult to accurately capture the complex patterns in the data. In addition, this method can provide strong support for the electricity bill risk prevention and control of power supply enterprises, help power supply enterprises discover and warn of electricity bill risks in a timely manner, take corresponding risk prevention and control measures, reduce electricity bill risks, and safeguard the economic benefits of power supply enterprises. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is a schematic flowchart of the electricity bill risk assessment method in an embodiment;
[0055] Figure 2 It is a schematic flowchart of the decision tree model construction method in an embodiment;
[0056] Figure 3 It is a schematic diagram of the decision tree model construction principle in an embodiment;
[0057] Figure 4 It is a schematic flowchart of the user risk score clustering step in an embodiment;
[0058] Figure 5 It is a schematic flowchart of the electricity bill risk assessment method in another embodiment;
[0059] Figure 6 It is a structural block diagram of the electricity bill risk assessment device in an embodiment;
[0060] Figure 7 It is an internal structure diagram of the computer equipment in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] In order to make the objectives, technical solutions, and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0062] In one embodiment, as Figure 1 shown, a method for evaluating electricity bill risks is provided. In this embodiment, an example is given where this method is applied to a terminal. It can be understood that this method can also be applied to a server, or to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and Internet of Things devices, etc. The server can be implemented by an independent server or a server cluster composed of multiple servers. In this embodiment, the method includes the following steps:
[0063] Step S110: Obtain the electricity bill data of multiple users to be evaluated.
[0064] In specific implementation, data in multiple dimensions such as the user's electricity consumption behavior characteristics, payment behavior characteristics, user attribute characteristics, and historical risk characteristics can be obtained as the user's electricity bill data. Among them, the electricity consumption behavior characteristics can be monthly electricity consumption, peak electricity consumption, valley electricity consumption, and electricity consumption volatility, etc.; the payment behavior characteristics can be the number of payment delays, the amount of arrears, the duration of arrears, and the number of payments within a short period of time, etc.; the user attribute characteristics can be user type, region, season, and climate data, etc.; the historical risk characteristics can be historical arrears records, historical fee behaviors, etc.
[0065] Step S120: Process the electricity bill data of each user to be evaluated through a decision tree model to obtain the user identification probability result value of each user to be evaluated.
[0066] In specific implementation, for the electricity bill data of any user to be evaluated, the electricity bill data of this user can be input into a pre-constructed decision tree model, and traversal starts from the root node of the decision tree model. At each internal node, the data is judged according to the splitting condition of the node. For example, if the splitting condition is "the number of payments within a short period of time > 10", then compare the value of feature X in the data with 10. According to the judgment result, select the corresponding branch (left branch or right branch) to enter the next node. Continue this process until reaching the leaf node to obtain the user identification probability result value of this user. The above process is executed for the electricity bill data of each user to be evaluated to obtain the user identification probability result value of each user to be evaluated. The user identification probability result value is in the range of [0, 1].
[0067] In some embodiments, after obtaining the electricity bill data of each user to be evaluated, preprocessing can be performed first, such as cleaning the data, handling missing values and outliers, and processing the preprocessed electricity bill data through a decision tree model respectively to obtain the user identification probability result values of each user to be evaluated.
[0068] Step S130, according to the user identification probability result values, determine the user risk scores of each user to be evaluated.
[0069] In a specific implementation, the user identification probability result values are in the interval [0, 1]. According to the user identification probability result values, determining the user risk scores of each user to be evaluated can be achieved by performing a mapping process on the user identification probability result values, and the interval of the user risk scores is [0, 100].
[0070] Step S140, perform clustering processing on each user risk score to obtain multiple clustering clusters.
[0071] In a specific implementation, the K-means clustering algorithm can be used to perform clustering processing on each user risk score to obtain multiple clustering clusters. The clustering process includes:
[0072] (1) Determine the number of clustering clusters K, that is, divide the users into several risk levels. For example, five categories, and then select K = 5.
[0073] (2) Initialize the center points. Randomly select K data points as the initial clustering centers, and these center points will be used as the representatives of each cluster.
[0074] (3) Assign data points. Traverse all the data points and assign each data point to the nearest clustering center. Specifically, the Euclidean distance can be used as the distance metric to assign each data point. The Euclidean distance formula is as follows:
[0075] (1)
[0076] Among them, is the coordinate of the clustering center, is the coordinate of the data point.
[0077] (4) Update the clustering centers. Recalculate the clustering centers for all the data points in each cluster. The new clustering center is the mean of all the data points in the cluster, which is expressed by the formula:
[0078] (2)
[0079] Among them, N is the number of data points in the cluster, is the data point in the cluster.
[0080] (5) Repeat steps (3) and (4): Continue to assign data points to the new cluster centers and update the cluster centers until the cluster centers no longer change (converge) or the preset number of iterations is reached.
[0081] (6) Result output: Output the final clustering result, including the cluster to which each data point belongs and the cluster centers of each cluster.
[0082] In step S150, determine the risk level corresponding to each clustering cluster as the risk level of the users to be evaluated within the clustering cluster.
[0083] In a specific implementation, the risk level of each clustering cluster can be determined according to the distribution of data points in each clustering cluster. The distribution is positively correlated with the risk level, that is, the more dispersed the data points are, the higher the risk level. The more compact the data points are, the lower the risk level. After determining the risk level corresponding to each clustering cluster, all users within the clustering cluster are classified into the corresponding risk level.
[0084] In the above electricity bill risk assessment method, electricity bill data of multiple users to be evaluated are obtained; the electricity bill data are processed respectively through a decision tree model to obtain the user identification probability result values of each user to be evaluated; according to the user identification probability result values, the user risk scores of each user to be evaluated are determined; clustering processing is performed on each user risk score to obtain multiple clustering clusters; the risk level corresponding to each clustering cluster is determined as the risk level of the users to be evaluated within the clustering cluster. By combining the use of a decision tree and a clustering algorithm, this method can make full use of the advantages of the two algorithms, overcome the limitations of a single algorithm, improve the accuracy of electricity bill risk assessment, effectively solve the complex multi-scenario problems of the power grid, and also predict the electricity bill risk status and potential risks of the corresponding users. Moreover, the clustering algorithm can perform further clustering analysis on the output results of the decision tree model, discover potential patterns and structures in the data, and provide more valuable risk information for power supply enterprises; thus overcoming the limitations of traditional algorithms in electricity bill risk assessment, which are difficult to accurately capture complex patterns in the data. In addition, this method can provide strong support for the electricity bill risk prevention and control of power supply enterprises, help power supply enterprises discover and warn of electricity bill risks in a timely manner, take corresponding risk prevention and control measures, reduce electricity bill risks, and safeguard the economic benefits of power supply enterprises.
[0085] In an exemplary embodiment, as Figure 2 shown, before the above step S120 processes the electricity bill data of each user to be evaluated respectively through a decision tree model to obtain the user identification probability result values of each user to be evaluated, it further includes:
[0086] Step S210, obtain sample electricity bill data; the sample electricity bill data includes time series of multiple electricity bill features;
[0087] Step S220: Calculate the variance for the time series of each electricity charge feature respectively, and filter out the target electricity charge data from the sample electricity charge data based on the variance.
[0088] Step S230: Construct a decision tree model based on the target electricity charge data.
[0089] Among them, the variance is used to measure the degree of dispersion of data. The larger the variance, the more obvious the change of the data.
[0090] In specific implementation, the variance can be calculated for the time series of each electricity charge feature in the sample electricity charge data respectively. Based on the variance, the data with more obvious changes is filtered out from the sample electricity charge data as the target electricity charge data. Further, a decision tree model is constructed based on the target electricity charge data.
[0091] In some embodiments, in step S220, filtering out the target electricity charge data from the sample electricity charge data based on the variance includes: filtering out the target features with variances greater than the threshold from multiple electricity charge features; and taking the electricity charge data corresponding to the target features in the sample electricity charge data as the target electricity charge data.
[0092] Step S230 constructing a decision tree model based on the target electricity charge data includes: using the target feature as the splitting feature and constructing a decision tree model based on the target electricity charge data.
[0093] Specifically, a minimum value of the variance can be set. Features below this threshold are considered to have too little change and have limited contribution to the identification of users' electricity charge risks, so they are removed, and features with variances greater than the threshold are retained. Among them, the variances of different electricity charge features are different to adapt to the characteristics of different electricity charge features.
[0094] In specific implementation, the first step of constructing a decision tree is to select features for data splitting. As Figure 3 shown, for each node, the decision tree algorithm will evaluate all possible features and select the best feature for splitting through the selected feature evaluation criterion. After the preliminary construction of the decision tree is completed, pruning can be further performed to remove the child nodes of the decision nodes to prevent overfitting.
[0095] In the related art, information gain and information gain ratio are often used for feature selection. However, the information gain criterion has a preference for attributes with more values. That is to say, using information gain as the judgment method will tend to select attributes with more values. And the information gain ratio criterion generates a multi-way tree, and the information gain ratio criterion can only be used for classification tasks. There are a large number of logarithmic operations in the entropy model, and the operation is very time-consuming. In this embodiment, the binary tree algorithm is adopted. Therefore, the Gini coefficient is used to replace the information gain ratio. The Gini coefficient represents the impurity of the model. The smaller the Gini coefficient, the lower the impurity and the better the feature.
[0096] Assume K categories, and the probability of the k-th category is , and the expression of the Gini coefficient of the probability distribution is:
[0097] (3)
[0098] For the sample D with the number |D|, according to a certain value a of the feature A, D is divided into and , then under the condition of the feature A, the expression of the Gini coefficient of the sample D is:
[0099] (4)
[0100] Based on the test effect on the validation set, a binary tree is selected as the final selected algorithm model to construct a decision tree. Among them, the test effect on the validation set can be evaluated by verification metrics such as accuracy, precision, recall, F1-score, and confusion matrix.
[0101] After completing the training of the decision tree model with the target electricity fee data, save the node information, feature list, prediction value, topological structure (such as the child nodes of each node, the branch logic corresponding to each splitting condition), etc. during the training process to generate a model file. When applying, parse this model file, reconstruct the decision tree model generated during the training process, and process the user's electricity fee data.
[0102] In this embodiment, by calculating the variance of each feature in the historical data, the features with variances greater than the set threshold are screened out, so as to retain the features that contribute to the model, thereby simplifying the decision tree model and improving the generation efficiency of the decision tree model.
[0103] In an exemplary embodiment, the above step S130 obtains the user risk scores of each user to be evaluated according to the user recognition probability result value, including: for each user to be evaluated, performing a mapping process on the user recognition probability result value of the user to be evaluated to obtain the user risk score of the user to be evaluated.
[0104] In a specific implementation, map the user recognition probability result value [0,1], and the obtained user risk score is in the interval [0,100], that is, the user recognition probability result value can be multiplied by 100 to obtain the user risk score of the user to be evaluated.
[0105] In this embodiment, after mapping the probability to a wider numerical range, similar data can be distinguished more finely. Moreover, when the probability values are concentrated at both ends of [0,1] (such as most sample probabilities are close to 0 or 1), the direct clustering effect may not be good. After mapping to [0,100], the data distribution is more dispersed, which is conducive to the clustering algorithm to capture differences, can alleviate data sparsity, and optimize the clustering effect.
[0106] In an exemplary embodiment, as Figure 4 shown, the above step S140 performs clustering processing on the user risk scores of the users to be evaluated, obtaining multiple clustering clusters, including:
[0107] Step S141, determining the initial number of clustering clusters;
[0108] Step S142, based on the initial number of clustering clusters, performing clustering processing on the user risk scores of the users to be evaluated, obtaining multiple initial clustering clusters; respectively obtaining the sum of the squares of the distances from the non-central points in each initial clustering cluster to the clustering center, adding the sums of the squares corresponding to each initial clustering cluster, obtaining the total sum of squares;
[0109] Step S143, increasing the initial number of clustering clusters to obtain a new number of clustering clusters, returning to the step of performing clustering processing on the user risk scores of the users to be evaluated based on the initial number of clustering clusters until a preset end condition is reached, obtaining the total sum of squares under multiple numbers of clustering clusters;
[0110] Step S144, according to the change trend of the total sum of squares under each number of clustering clusters, determining the target number of clustering clusters, and using the multiple clustering clusters obtained by clustering the target number of clustering clusters as the clustering results for each user to be evaluated.
[0111] In specific implementation, an initial number of clustering clusters can be determined first, such as 2, and then starting from K = 2, gradually increasing the number of clusters, and calculating the total sum of squares corresponding to each K. Among them, the determination process of the total sum of squares corresponding to each number of clustering clusters includes: performing clustering processing on the user risk scores of the users to be evaluated according to the number of clustering clusters, obtaining multiple clustering clusters. Respectively obtaining the sum of the squares of the distances from the non-central points in each initial clustering cluster to the clustering center, adding the sums of the squares corresponding to each initial clustering cluster, obtaining the total sum of squares, which is expressed by the formula:
[0112] (5)
[0113] where K is the number of clustering clusters, represents the i-th cluster, represents the clustering center of the i-th cluster, and x is the data point.
[0114] Taking each number of clustering clusters as the abscissa and the total sum of squares under each number of clustering clusters as the ordinate to draw a curve, determining the inflection point of the curve, that is, the point where the decline rate of the total sum of squares significantly slows down, as the optimal target number of clustering clusters.
[0115] In some embodiments, the determined target number of clustering clusters can also be adjusted according to actual needs, and the user risk scores are clustered using the adjusted target number of clustering clusters.
[0116] In this embodiment, based on the relationship curve between the number of clustering clusters and the total sum of squared distances within the clusters, the inflection point can be visually observed to determine the target number of clustering clusters, and the optimal number of clustering clusters can be selected, reducing the subjective influence of manual presetting and facilitating the provision of an objective reference basis when the business requirements are unclear.
[0117] In one exemplary embodiment, step S150 above for determining the risk level corresponding to each clustering cluster includes: determining the clustering center of each clustering cluster; respectively obtaining the mean value of the distances from the non-central points in each clustering cluster to the clustering center, sorting the mean values corresponding to each clustering cluster, and determining the risk level corresponding to each clustering cluster based on the sorting result.
[0118] In specific implementation, after obtaining multiple clustering clusters, the clustering center of each cluster can be obtained simultaneously. For any clustering cluster, calculate the distances from the non-central points in this clustering cluster to the clustering center, and calculate the mean value of the obtained distances. Thus, such a mean value can be obtained for each clustering cluster. Further, sort the mean values corresponding to each clustering cluster in descending order to obtain a clustering cluster sequence. It can be understood that the larger the mean distance, the more dispersed the points within the cluster, and there may be potential anomalies or high-risk behaviors (such as large fluctuations in electricity consumption and frequent payment delays). Therefore, the risk level corresponding to each clustering cluster can be determined in the order from high to low risk level, and thus user groups with different risk levels can be obtained.
[0119] In this embodiment, the risk is quantified by the mean distance between the clustering center and the member points, which can achieve the rapid identification of user groups with different risks and improve the identification efficiency while ensuring the accuracy of the level.
[0120] In one embodiment, to more clearly illustrate the embodiments of the present application, specific examples in combination with the accompanying drawings will be described below. Refer to Figure 5 , which shows a specific process schematic diagram of an electricity bill risk assessment method. To solve the problem of complex power grid scenarios, this solution proposes a combined algorithm for electricity bill risk scoring (decision tree algorithm + K-means clustering algorithm), and predicts the electricity bill risk status and potential risks of corresponding users according to user payment characteristics. To improve the accuracy and effectiveness of the model algorithm, the blacklist user data pushed by the risk monitoring center is used as the benchmark truth set. It includes two processes: training and prediction.
[0121] During the training process:
[0122] S1. Input the training samples (the sample data has been pre-processed in the data factory);
[0123] S2. Perform decision tree algorithm processing, complete training to generate a model file and save it. The model file stores information such as node information, feature list, prediction value, topological structure (e.g., child nodes of each node, branch logic corresponding to each splitting condition), etc.
[0124] During the prediction process:
[0125] S3. Parse the model file generated by training and reconstruct the decision tree model generated during the training process.
[0126] S4. Input the processed electricity bill data, and through the trained decision tree model, analyze and calculate the corresponding risk score of the user;
[0127] S5. Perform clustering through the K-means clustering algorithm to obtain user groups with different risk levels.
[0128] This application proposes a combined algorithm for electricity bill risk scoring to solve the complex multi-scenarios of the Southern Power Grid. By constructing a decision tree, for each node, the decision tree algorithm will evaluate all possible features and select the best feature for splitting through the Gini impurity as the feature selection method. The decision tree algorithm selects the CART tree (Classification And Regression Tree, a binary tree) to map the user recognition probability result value [0, 1] to the range of 0 to 100, forming the corresponding risk score. According to the risk score generated by the decision tree, then perform clustering through the K-means clustering algorithm, output the final clustering results, including the cluster to which each data point belongs and the clustering center of each cluster, and finally obtain the user groups at different levels by taking the mean in descending order of the cluster center points, so as to fully understand the potential risks of each user.
[0129] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily have to be sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.
[0130] Based on the same inventive concept, an embodiment of the present application further provides a power consumption cost risk assessment device for implementing the power consumption cost risk assessment method involved above. The solution for solving the problem provided by this device is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the power consumption cost risk assessment device provided below can refer to the limitations on the power consumption cost risk assessment method in the above text, and will not be repeated here.
[0131] In one embodiment, as Figure 6 shown, a power consumption cost risk assessment device is provided, including:
[0132] A data acquisition module 610, configured to acquire power consumption cost data of multiple users to be evaluated;
[0133] A decision tree processing module 620, configured to process the power consumption cost data of each user to be evaluated through a decision tree model respectively, and obtain the user identification probability result value of each user to be evaluated;
[0134] A score determination module 630, configured to determine the user risk score of each user to be evaluated according to the user identification probability result value;
[0135] A score clustering module 640, configured to perform clustering processing on each user risk score to obtain multiple clustering clusters;
[0136] A level determination module 650, configured to determine the risk level corresponding to each clustering cluster as the risk level of the users to be evaluated within the clustering cluster.
[0137] In one of the embodiments, the device further includes a decision tree construction module, configured to acquire sample power consumption cost data; the sample power consumption cost data includes time series of multiple power consumption cost features; calculate variances for the time series of each power consumption cost feature respectively, and screen out target power consumption cost data from the sample power consumption cost data based on the variances; construct a decision tree model based on the target power consumption cost data.
[0138] In one of the embodiments, the decision tree construction module is further configured to screen out target features with variances greater than a threshold from multiple power consumption cost features; use the power consumption cost data corresponding to the target features in the sample power consumption cost data as the target power consumption cost data; construct a decision tree model based on the target power consumption cost data with the target features as splitting features.
[0139] In one of the embodiments, the score determination module 630 is further configured to perform mapping processing on the user identification probability result value of each user to be evaluated to obtain the user risk score of the user to be evaluated.
[0140] In one embodiment, the fractional clustering module 640 is further configured to determine the number of initial clustering clusters; cluster the user risk scores based on the number of initial clustering clusters to obtain a plurality of initial clustering clusters; respectively obtain the sum of the squares of the distances from the non-central points in each initial clustering cluster to the clustering center, add up the sums of squares corresponding to each initial clustering cluster to obtain the total sum of squares; increase the number of initial clustering clusters to obtain a new number of clustering clusters, and return to the step of clustering the user risk scores of the user to be evaluated based on the number of initial clustering clusters until a preset end condition is reached, to obtain the total sum of squares under multiple numbers of clustering clusters; determine the target number of clustering clusters according to the change trend of the total sum of squares under each number of clustering clusters, and use the multiple clustering clusters obtained by clustering the target number of clustering clusters as the clustering results for each user to be evaluated.
[0141] In one embodiment, the level determination module 650 is further configured to determine the clustering center of each clustering cluster; respectively obtain the mean value of the distances from the non-central points in each clustering cluster to the clustering center, sort the mean values corresponding to each clustering cluster, and determine the risk level corresponding to each clustering cluster based on the sorting result.
[0142] Each module in the above electricity fee risk assessment device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0143] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 7 shown. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. The computer program, when executed by the processor, implements an electricity fee risk assessment method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.
[0144] Those skilled in the art can understand,Figure 7 The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0145] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0146] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0147] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0148] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0149] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0150] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0151] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for evaluating electricity fee risks, characterized in that, The method includes: Obtaining electricity bill data of multiple users to be evaluated; Processing the electricity bill data of each user to be evaluated through a decision tree model respectively to obtain the user identification probability result value of each user to be evaluated; Determining the user risk score of each user to be evaluated according to the user identification probability result value; Performing clustering processing on each of the user risk scores to obtain multiple clustering clusters; Determining the risk level corresponding to each clustering cluster as the risk level of the users to be evaluated within the clustering cluster.
2. The method according to claim 1, wherein Before the step of processing the electricity bill data of each user to be evaluated through a decision tree model respectively to obtain the user identification probability result value of each user to be evaluated, it further includes: Obtaining sample electricity bill data; the sample electricity bill data includes time series of multiple electricity bill features; Calculating the variance of the time series of each electricity bill feature respectively, and screening out target electricity bill data from the sample electricity bill data based on the variance; Constructing a decision tree model based on the target electricity bill data.
3. The method according to claim 2, characterized in that, The screening out of the target electricity bill data from the sample electricity bill data based on the variance includes: Screening out target features with a variance greater than a threshold from the multiple electricity bill features; Taking the electricity bill data corresponding to the target features in the sample electricity bill data as the target electricity bill data; The constructing of a decision tree model based on the target electricity bill data includes: Taking the target feature as the splitting feature and constructing a decision tree model based on the target electricity bill data.
4. The method according to claim 1, wherein The obtaining of the user risk score of each user to be evaluated according to the user identification probability result value includes: For each user to be evaluated, performing mapping processing on the user identification probability result value of the user to be evaluated to obtain the user risk score of the user to be evaluated.
5. The method according to claim 1, characterized in that, The performing of clustering processing on each of the user risk scores to obtain multiple clustering clusters includes: Determining the initial number of clustering clusters; Based on the initial number of clustering clusters, performing clustering processing on each of the user risk scores to obtain multiple initial clustering clusters; respectively obtaining the sum of the squares of the distances from the non-central points in each initial clustering cluster to the clustering center, and adding the sums of the squares corresponding to each initial clustering cluster to obtain the total sum of squares; Increasing the initial number of clustering clusters to obtain a new number of clustering clusters, and returning to the step of performing clustering processing on the user risk scores of the users to be evaluated based on the initial number of clustering clusters until a preset end condition is reached, to obtain the total sum of squares under multiple numbers of clustering clusters; According to the change trend of the total sum of squares under each number of clustering clusters, determining the target number of clustering clusters, and taking the multiple clustering clusters obtained by clustering the target number of clustering clusters as the clustering result of each user to be evaluated.
6. The method according to claim 1, wherein The determining of the risk level corresponding to each clustering cluster includes: Determining the clustering center of each clustering cluster; Respectively obtaining the mean value of the distances from the non-central points in each clustering cluster to the clustering center, sorting the mean values corresponding to each clustering cluster, and determining the risk level corresponding to each clustering cluster based on the sorting result.
7. An electricity charge risk assessment device, characterized in that, The device includes: A data acquisition module, configured to obtain electricity bill data of multiple users to be evaluated; A decision tree processing module, configured to process the electricity bill data of each of the to-be-evaluated users through a decision tree model, and obtain the user identification probability result value of each of the to-be-evaluated users; A score determination module, configured to determine the user risk score of each of the to-be-evaluated users according to the user identification probability result value; A score clustering module, configured to perform clustering processing on each of the user risk scores to obtain a plurality of clustering clusters; A level determination module, configured to determine the risk level corresponding to each of the clustering clusters as the risk level of the to-be-evaluated users within the clustering cluster.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the electricity bill risk assessment method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the electricity bill risk assessment method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the electricity bill risk assessment method according to any one of claims 1 to 6.