Electric charge payment reminding grading method and device, electronic equipment and storage medium
By classifying and automatically labeling the characteristic information of users in arrears, and constructing and correcting decision trees, the problems of single level classification and manual labeling bias in traditional electricity bill collection methods are solved. This enables accurate user segmentation and personalized collection strategies, improving collection efficiency and credit matching.
Patent Information
- Application Number
- CN202510842198.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-11-18
AI Technical Summary
Traditional electricity bill collection methods suffer from problems such as a single collection level classification, large deviations in manual labeling, a single dimension of user segmentation, and a mismatch between collection strategies and user credit, making it difficult to adapt to the needs of large-scale and dynamically changing user data processing.
By identifying the arrears characteristics of users, users are classified based on these characteristics, generating the first cluster center and performing automated labeling. A risk classification decision tree for arrears users is then constructed, and finally, the decision tree is corrected to generate an accurate classification decision tree.
It accurately reflects the dynamic changes in user behavior, eliminates the bias of manual annotation, efficiently processes high-dimensional data, improves the matching degree between collection strategies and user credit, and enhances collection efficiency and user satisfaction.
Smart Images

Figure CN120975776A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of power system data analysis, and particularly relates to a method and device for electricity payment collection grading, an electronic device and a storage medium. BACKGROUND
[0002] This section is intended to provide background or context to the embodiments of the disclosure recited in the claims. The description herein does not constitute admission that the prior art is prior art nor does it constitute an admission of any description in this section as prior art to an application.
[0003] With the continuous progress of the power industry, the number of users is growing, and the electricity consumption behavior is increasingly complex and diverse. The traditional electricity payment collection method gradually exposes many limitations. The collection means mainly relies on manual statistics and simple threshold judgment, which is difficult to adapt to large-scale, dynamic user data processing needs.
[0004] However, in the related art, there are problems such as single collection level division, large manual annotation deviation, single user grouping dimension, and mismatch between collection strategy and user credit. SUMMARY
[0005] Therefore, the purpose of the present disclosure is to provide a method and device for electricity payment collection grading, an electronic device and a storage medium, which at least partially solve one of the technical problems in the related art.
[0006] To achieve the above purpose, the first aspect of the exemplary embodiments of the present disclosure provides a method for electricity payment collection grading, which comprises:
[0007] determining the arrears feature information of the arrears user, classifying the arrears user based on the arrears feature information, and obtaining a first clustering center;
[0008] annotating the arrears user in the first clustering center to obtain arrears user annotation information;
[0009] constructing a decision structure based on the arrears user annotation information to obtain an arrears user risk grading decision tree;
[0010] correcting the arrears user risk grading decision tree to obtain a corrected grading decision tree.
[0011] Based on the same inventive concept, the second aspect of the exemplary embodiments of the present disclosure provides a device for electricity payment collection grading, which comprises:
[0012] a clustering center determination module configured to determine the arrears feature information of the arrears user, classify the arrears user based on the arrears feature information, and obtain a first clustering center;
[0013] The label information determination module is configured to label the delinquent users in the first clustering center to obtain delinquent user label information.
[0014] The decision tree determination module is configured to construct a decision structure based on the delinquent user label information to obtain a delinquent user risk classification decision tree.
[0015] The decision tree correction module is configured to correct the delinquent user risk classification decision tree to obtain a corrected classification decision tree.
[0016] Based on the same inventive concept, a third aspect of the exemplary embodiments of the present disclosure provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method of the first aspect when executing the program.
[0017] Based on the same inventive concept, a fourth aspect of the exemplary embodiments of the present disclosure provides a non-transitory computer-readable storage medium, which stores computer instructions for causing a computer to execute the method of the first aspect.
[0018] Based on the same inventive concept, a fifth aspect of the exemplary embodiments of the present disclosure provides a computer program product, comprising computer program instructions, which, when executed on a computer, cause the computer to execute the method of the first aspect.
[0019] As can be seen from the above, the electricity payment collection classification method, device, electronic device, and storage medium provided by the embodiments of the present disclosure, the method comprises: determining delinquent user characteristic information, classifying the delinquent users based on the delinquent user characteristic information to obtain a first clustering center; labeling the delinquent users in the first clustering center to obtain delinquent user label information; constructing a decision structure based on the delinquent user label information to obtain a delinquent user risk classification decision tree; and correcting the delinquent user risk classification decision tree to obtain a corrected classification decision tree. The present disclosure can accurately reflect the dynamic changes of user behavior, eliminate artificial labeling bias, efficiently process high-dimensional data, and realize multi-dimensional user grouping to improve the matching degree of collection strategies and user credit. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the present disclosure or the related art, the following will briefly introduce the drawings needed to be used in the embodiments or related art descriptions. Obviously, the drawings in the following description are only embodiments of the present disclosure, and for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.
[0021] Figure 1 An application scenario schematic diagram of the electricity payment collection grading method provided for the exemplary embodiments of the present disclosure is provided.
[0022] Figure 2 A flow schematic diagram of the electricity payment collection grading method provided for the exemplary embodiments of the present disclosure is provided.
[0023] Figure 3 A structure schematic diagram of the electricity payment collection grading device provided for the exemplary embodiments of the present disclosure is provided.
[0024] Figure 4 A structure schematic diagram of the electronic device hardware provided for the exemplary embodiments of the present disclosure is provided. DETAILED DESCRIPTION
[0025] It can be understood that, before using the technical solutions disclosed in the embodiments of the present application, the type, use range, use scenario, etc. of the personal information involved in the present application should be informed to the user and the authorization of the user should be obtained in a proper manner according to relevant laws and regulations.
[0026] For example, in response to receiving the active request of the user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will need to obtain and use the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the software or hardware such as the electronic device, application program, server or storage medium, etc. that performs the operation of the technical solutions of the present application according to the prompt information.
[0027] As an optional but non-limiting implementation manner, in response to receiving the active request of the user, the manner of sending the prompt information to the user may, for example, be the manner of a pop-up window, and the prompt information may, for example, be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide the personal information to the electronic device.
[0028] It can be understood that the above notification and user authorization process is only illustrative, and does not limit the implementation manner of the present application, and other manners that meet the relevant laws and regulations can also be applied to the implementation manner of the present application.
[0029] It can be understood that the data (including but not limited to the data itself, the acquisition or use of the data) involved in the present technical solutions should comply with the requirements of the relevant laws and regulations and the relevant provisions.
[0030] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the principles and spirits of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are only given to enable those skilled in the art to better understand and implement the present disclosure, and in no way limit the scope of the present disclosure. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to enable the scope of the present disclosure to be fully conveyed to those skilled in the art.
[0031] In this document, it should be understood that any quantity of elements in the drawings is used for illustration only and not limitation, and any naming is only for differentiation and does not have any limiting meaning.
[0032] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should be understood as the common meanings understood by those skilled in the art to which the present disclosure belongs. The terms "first", "second", and similar terms used in the embodiments of the present disclosure do not represent any order, number, or importance, but are only used to distinguish different components. The terms "include", "contain", and similar terms mean that the elements or objects before the terms encompass the elements or objects listed after the terms and their equivalents, and do not exclude other elements or objects. The terms "connect" or "connected" and similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms "up", "down", "left", "right", and the like are only used to represent relative positional relationships, and when the absolute positions of the described objects change, the relative positional relationships may also change accordingly. The article "a" or "an" before an element does not exclude the existence of multiple such elements.
[0033] The principles and spirits of the present disclosure will be described in detail below with reference to several representative embodiments of the present disclosure.
[0034] As described in the background, in the related art, there are problems such as single division of arrears level, large deviation of manual annotation, single dimension of user grouping, and mismatch between arrears strategy and user credit. Specifically, in the electricity arrears collection scene, although traditional machine learning algorithms and time series analysis can classify and predict user payment behavior, existing technologies have limitations when processing high-dimensional and dynamically changing data. For example, some existing technologies only divide users into arrears levels according to a single dimension of arrears amount, and use a fixed arrears amount threshold to divide users into light arrears and heavy arrears levels. This method is simple and direct, but completely ignores the dynamic change characteristics of user behavior, cannot accurately reflect the real credit status of users, and leads to lack of pertinence of arrears collection strategy.
[0035] In addition, the prior art is deficient in feature engineering and lacks effective fusion of multi-modal data. Although some technologies introduce machine learning classification methods, they only simply input a few features such as basic information of users and overdue amount into the model, and do not fully mine potential information in user behavior data. At the same time, although time series analysis can analyze the evolution law of user payment behavior over time, it is deficient in user grouping and arrears collection strategy formulation in combination with time series features, and does not take the fluctuation pattern of payment behavior into account in grouping.
[0036] These problems lead to rough division of arrears collection grades, large deviation of manual annotation, single dimension of user grouping, and mismatch between arrears collection strategy and user credit, etc. For example, some users may occasionally have high overdue fees, but have good payment records in general. Single threshold division cannot accurately reflect the real credit status of such users. Manual annotation is greatly influenced by subjective factors, and different annotators may have different annotations for the same user. When facing high-dimensional user behavior data, manual annotation is inefficient and difficult to ensure accuracy. The prior art does not analyze the change law of user payment behavior in different time periods and the influence of fluctuation of payment amount on user credit evaluation when grouping users. Finally, due to inaccurate division of arrears collection grades and unreasonable grouping of users, the arrears collection strategy is seriously out of line with the real credit status of users. Excessive arrears collection measures are taken for users with good credit but occasional arrears, while the arrears collection effort may be insufficient for users with long-term malicious arrears, affecting the arrears collection efficiency and user satisfaction.
[0037] To solve the above problems, the present disclosure provides an electricity arrears collection grading method, device, electronic equipment and storage medium scheme, which comprises:
[0038] The arrears feature information of the arrears user is determined, the arrears user is classified based on the arrears feature information, and a first clustering center is obtained; the arrears user in the first clustering center is labeled, and arrears user labeling information is obtained; a decision structure is constructed based on the arrears user labeling information, and an arrears user risk grading decision tree is obtained; the arrears user risk grading decision tree is corrected, and a corrected grading decision tree is obtained. The disclosure overcomes the limitations of traditional single threshold division by determining the arrears feature information of the arrears user, accurately captures the dynamic changes of user behavior, and combines multi-dimensional information such as time series features and payment behavior fluctuation patterns to classify the arrears user and obtain the first clustering center. On this basis, an automatic labeling mechanism is used to label the arrears user in the first clustering center, eliminating the bias of manual labeling, efficiently processing high-dimensional data, and obtaining accurate arrears user labeling information. Based on these labeling results, a decision structure is constructed, an arrears user risk grading decision tree is generated, and the decision tree is optimized through a correction mechanism, and finally a corrected grading decision tree is obtained, realizing a personalized collection strategy closely matched with the real credit status of the user, improving the collection efficiency, and reducing the arrears risk.
[0039] After introducing the basic principles of the disclosure, the various non-limiting embodiments of the disclosure will be specifically introduced below.
[0040] Reference Figure 1 , which is a schematic diagram of an application scenario of the electricity fee collection grading method provided by an exemplary embodiment of the disclosure.
[0041] In this application scenario, a terminal device 101 and a server 102 are included. The terminal device 101 and the server 102 can be connected through a wired or wireless communication network to realize data interaction.
[0042] The terminal device 101 can be an electronic device close to the user side with data transmission, multimedia input / output functions, including but not limited to desktop computers, mobile phones, mobile computers, tablet computers, media players, smart wearable devices, personal digital assistants (PDA), or other electronic devices capable of realizing the above functions. The electronic device can include a processor and a display screen with touch input function, the display screen is used to present a graphical user interface, the graphical user interface can display an application interface, the processor is used to process application data, generate a graphical user interface, and control the display of the graphical user interface on the display screen.
[0043] The server 102 can be a stand-alone physical server, a server cluster or a distributed system composed of multiple physical servers, a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms.
[0044] In some example embodiments, the electricity payment collection grading method can run on the terminal device 101 or the server 102.
[0045] When the electricity payment collection grading method runs on the server 102, the server 102 is configured to provide the electricity payment collection grading service to the user of the terminal device 101.
[0046] The server 102 determines the arrears characteristic information of the arrears user, classifies the arrears user based on the arrears characteristic information, and obtains a first clustering center.
[0047] The server 102 labels the arrears user in the first clustering center and obtains arrears user labeling information.
[0048] The server 102 constructs a decision structure based on the arrears user labeling information and obtains an arrears user risk grading decision tree.
[0049] The server 102 corrects the arrears user risk grading decision tree and obtains a corrected grading decision tree. The server 102 visualizes the decision path atlas based on the corrected grading decision tree, and then transmits the decision path atlas to the terminal device 101.
[0050] It should be noted that the above application scenarios are only for the convenience of understanding the spirit and principles of the present disclosure, and the embodiments of the present disclosure are not limited in this respect. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.
[0051] Reference Figure 2 An electricity payment collection grading method, the method comprising the following steps:
[0052] Step S210, determining the arrears characteristic information of the arrears user, classifying the arrears user based on the arrears characteristic information, and obtaining a first clustering center.
[0053] In this step, by determining the arrears feature information of the arrears user, such as the arrears amount, the arrears frequency, the arrears duration and the like, the dynamic change of the user behavior can be accurately captured, and the limitation of the traditional single threshold division is overcome. Based on these feature information, the arrears user is classified to obtain the first clustering center, and the preliminary fine division of the user group is realized. This process not only improves the accuracy of user grouping, but also provides a solid foundation for subsequent labeling and decision tree construction, which can effectively support the development of differentiated collection strategies and improve the pertinence and efficiency of the collection work.
[0054] In some embodiments, the arrears feature information includes: time decay type feature, arrears frequency feature, arrears amount proportion feature, arrears duration feature and arrears behavior feature.
[0055] In specific implementation, the time decay type feature refers to:
[0056] An exponential decay function is used to calculate the historical late fee influence factor. As time goes by, the influence of early late fees on the current user credit status gradually weakens. Assuming that there are n late fees, the i-th late fee is generated at t i , the late fee amount is P i , and the current time is t, then the arrears late fee influence factor formula is Where λ is the decay coefficient, (t-t i ) represents the time interval from the generation of the i-th late fee to the present. This formula gives different degrees of attenuation according to the time distance from the generation of each late fee to the present, so that the late fees generated in the near future have a greater impact on the current credit status. This feature can more accurately reflect the dynamic influence of historical late fees on the current credit of the user.
[0057] In specific implementation, the arrears frequency feature refers to:
[0058] The ratio of the number of arrears to the total number of payments of the arrears user in a certain period of time (such as one year) is calculated. The higher the ratio, the more frequent the user's arrears, and the greater the arrears risk. For example, if a user has 10 payments in the past 12 months, 4 of which are arrears, the arrears frequency is 4 ÷ 10 = 0.4. Through this feature, the frequency of occurrence of the user's arrears behavior can be directly measured, providing a key basis for evaluating the credit status and developing collection strategies.
[0059] In specific implementation, the arrears amount proportion feature refers to:
[0060] The proportion of the amount of each overdue fee to the total amount of electricity fee of the time is calculated. This proportion reflects the severity of each overdue fee. For example, if the total amount of electricity fee is 200 yuan and the actual overdue fee is 50 yuan, the proportion of the overdue fee is 50 ÷ 200 = 0.25. Statistical analysis (such as calculating the mean, maximum value, etc.) of these proportions over a period of time can more accurately grasp the severity and regularity of the overdue behavior of the overdue user, and help to accurately locate the high-risk overdue user.
[0061] In specific implementation, the overdue behavior feature refers to:
[0062] The change trend of the payment behavior of the overdue user. A "normal payment- overdue- late fee" three-state transition probability matrix is established for the overdue user to quantify the state transition mode. By analyzing the transition probability between different payment states of the overdue user, the change trend of the payment behavior of the overdue user can be understood in depth, and the future payment state of the overdue user can be predicted.
[0063] In some embodiments, the classification of the overdue user based on the overdue feature information obtains a first clustering center, including:
[0064] The time decay feature, the overdue frequency feature, the overdue amount proportion feature, the overdue time feature and the overdue behavior feature are quantified to obtain a multi-dimensional feature vector;
[0065] The overdue user data points of the multi-dimensional feature vector are determined, and coarse-grained clustering is performed based on the overdue user data points to obtain a second clustering center;
[0066] A similarity graph is constructed based on the second clustering center to obtain a Laplacian matrix;
[0067] The Laplacian matrix is subjected to eigenvalue decomposition to obtain an eigenvector matrix;
[0068] The eigenvector matrix is normalized to obtain a normalized eigenvector matrix;
[0069] The normalized eigenvector matrix is clustered to obtain the first clustering center.
[0070] In specific implementation, the time decay feature, the overdue frequency feature, the overdue amount proportion feature, the overdue time feature and the overdue behavior feature are quantified to obtain a multi-dimensional feature vector in the following manner:
[0071] By quantifying the time decay feature, the arrears frequency feature, the arrears amount proportion feature, the arrears duration feature and the arrears behavior feature, the features are converted into specific numerical values to form a multi-dimensional feature vector (for example, the time decay feature, the arrears frequency feature, the arrears amount proportion feature, the arrears duration feature and the arrears behavior feature are quantified by using an exponential decay function, proportional calculation, statistical analysis and transition probability matrix, etc.). This process can comprehensively and accurately reflect the payment behavior patterns and credit status of the arrears users, provide rich data support for subsequent user classification, risk assessment and arrears collection strategy formulation, and enhance the model's ability to capture dynamic changes in user behavior and classification accuracy.
[0072] In specific implementation, the manner of determining the arrears user data points of the multi-dimensional feature vector, performing coarse-grained clustering based on the arrears user data points, and obtaining the second clustering center is:
[0073] The improved K-Means++ algorithm is adopted. This algorithm overcomes the problem of sensitivity to initial center and easy to fall into local optimum of the traditional K-Means algorithm by optimizing the selection of initial clustering center.
[0074] In specific implementation, the manner of constructing a similarity graph based on the second clustering center and obtaining a Laplacian matrix is:
[0075] The sample similarity graph is constructed, each second clustering center is regarded as a node of the graph, and the weight of the edge is determined according to the similarity degree of the samples in the multi-dimensional features such as arrears frequency, arrears amount proportion and arrears duration, for example, the similarity is measured by using the reciprocal of Euclidean distance or cosine similarity, and the higher the similarity, the greater the weight of the edge. Then, the Laplacian matrix is generated according to the constructed similarity graph, and the Laplacian matrix reflects the structure information of the graph.
[0076] In specific implementation, the manner of performing eigenvalue decomposition on the Laplacian matrix to obtain an eigenvector matrix is:
[0077] The eigenvalue decomposition is performed on the Laplacian matrix to calculate its eigenvalues and eigenvectors. Generally, the eigenvectors corresponding to the smallest k non-zero eigenvalues are selected, and these eigenvectors are combined to form a new matrix, i.e., the eigenvector matrix.
[0078] In specific implementation, the manner of normalizing the eigenvector matrix to obtain a normalized eigenvector matrix is:
[0079] Each row of the eigenvector matrix is normalized, and the normalized row vector is regarded as a new eigenvector, i.e., the normalized eigenvector matrix, for example, the normalization can be performed by using the minimum-maximum normalization, Z-score normalization (standardization), decimal scaling normalization, etc., which is not specifically limited herein.
[0080] In particular implementation, the normalized feature vector matrix is clustered to obtain the first clustering center in the following manner:
[0081] The normalized feature vector matrix is clustered using a traditional clustering algorithm (such as the K-Means algorithm), so as to achieve more accurate division of the boundary arrear user sample, effectively solve the problem of cluster overlap, and make the clustering result more accurate and reasonable.
[0082] In the above example embodiment, the manner of obtaining the first clustering center is introduced, and the manner of obtaining the second clustering center is introduced as follows:
[0083] In some embodiments, the arrear user data point includes a plurality of arrear user sub-data points.
[0084] The coarse-grained clustering based on the arrear user data point to obtain the second clustering center includes:
[0085] A target arrear user sub-data point is determined from the plurality of arrear user sub-data points, and the target arrear user sub-data point is taken as a third clustering center.
[0086] The Mahalanobis distance between the third clustering center and other arrear user sub-data points is determined, and the arrear user sub-data point with the farthest Mahalanobis distance is taken as a fourth clustering center.
[0087] The plurality of arrear user sub-data points are assigned to the third clustering center and the fourth clustering center based on a minimum distance algorithm.
[0088] The mean of the arrear user sub-data points of the third clustering center and the fourth clustering center is calculated to obtain the second clustering center.
[0089] In particular implementation, the target arrear user sub-data point is determined from the plurality of arrear user sub-data points, and the target arrear user sub-data point is taken as the third clustering center in the following manner:
[0090] An arrear user sub-data point is randomly selected as a first initial clustering center, i.e., a third clustering center.
[0091] In particular implementation, the Mahalanobis distance between the third clustering center and other arrear user sub-data points is determined, and the arrear user sub-data point with the farthest Mahalanobis distance is taken as the fourth clustering center in the following manner:
[0092] The Mahalanobis distance of each arrears user sub-data point and the selected third cluster center is calculated, and the Mahalanobis distance fully considers the correlation and scale difference between features. Based on this, the arrears user sub-data point farthest from the currently selected third cluster center is selected as the next fourth cluster center. This process is repeatedly repeated until a preset number of fourth cluster centers are selected, and the final number of fourth cluster centers is consistent with the preset number.
[0093] In specific implementation, the manner of assigning the plurality of arrears user sub-data points to the third cluster center and the fourth cluster center based on the minimum distance algorithm, and calculating the mean of the arrears user sub-data points of the third cluster center and the fourth cluster center to obtain the second cluster center is as follows:
[0094] According to the minimum distance algorithm, all arrears user sub-data points are assigned to the nearest cluster center to form a preliminary cluster. Then, the mean of the arrears user sub-data points in each cluster is calculated, and the mean is taken as a new cluster center. Then, the data points are re-assigned based on the new cluster center. This process is iteratively repeated until the cluster center no longer changes significantly, and finally a stable second cluster center is generated. The minimum distance algorithm refers to: for each arrears user sub-data point, the distance (such as Euclidean distance or Mahalanobis distance) between it and each cluster center is calculated, and the data point is assigned to the cluster to which the nearest cluster center belongs.
[0095] In step S220, the arrears users in the first cluster center are labeled to obtain arrears user labeling information.
[0096] In this step, for arrears users, the sampling ratio is automatically calculated according to the variance of the samples in the cluster. A cluster with a larger variance indicates that the arrears user samples in the cluster have larger differences, and more samples need to be extracted for labeling to ensure the accuracy and representativeness of the labeling. A cluster with a smaller variance appropriately reduces the sampling ratio, which can realize fine classification and labeling of user arrears behavior and provide accurate basis for subsequent decision tree construction and collection strategy formulation. This step not only improves the usability and accuracy of the data, but also effectively eliminates the subjective bias of manual labeling.
[0097] In some embodiments, the labeling of the arrears users in the first cluster center to obtain arrears user labeling information comprises:
[0098] determining the difference of the arrears user data points in the first cluster center, and determining a sampling ratio based on the difference;
[0099] sampling and labeling the arrears user data points in the first cluster center based on the sampling ratio to obtain a labeled data set;
[0100] determine a labeling condition of the labeling data set, assign a weight to the labeling condition to obtain a set of weighted labeling conditions;
[0101] construct a membership function for the set of weighted labeling conditions to obtain a comprehensive membership degree;
[0102] classify the labeling data set based on the comprehensive membership degree to obtain the overdue user labeling information.
[0103] In specific implementation, the difference of the overdue user data points in the first clustering center is determined, and the sampling ratio is determined based on the difference in the following manner:
[0104] First, the overdue users are divided into different clusters according to the first clustering center. For each cluster, the characteristics of overdue frequency, amount ratio, and duration are comprehensively considered to determine the difference degree of the samples in the cluster. For example, the distribution range of these characteristics is observed directly. If the overdue frequency of some users is extremely low and that of some users is extremely high, the distribution range is large, which means that the sample difference is large, and the variance is large. On the contrary, if the values of these characteristics are relatively concentrated, the sample difference is small, and the variance is small. Then, a basic sampling ratio is set, which is assumed to be 20%. At the same time, two variance thresholds are set, which are a high threshold and a low threshold. When the sample variance in the cluster is higher than the high threshold, it means that the sample difference of the overdue users in the cluster is extremely large. In order to ensure the accuracy and representativeness of labeling, the sampling ratio is greatly increased based on the basic ratio. For example, the sampling ratio is additionally increased by 5% for each time the threshold is exceeded by a certain degree. When the variance is lower than the low threshold, it means that the samples in the cluster are similar, and the sampling ratio is appropriately reduced to improve the labeling efficiency. For example, the sampling ratio is reduced by 3% for each time the threshold is lowered by a certain degree. When the variance is between the two thresholds, the sampling ratio is maintained at the basic ratio.
[0105] In specific implementation, the overdue user data points in the first clustering center are sampled and labeled based on the sampling ratio in the following manner:
[0106] According to the sampling ratio determined in the above steps, samples are randomly selected in each cluster. For example, if a cluster has 100 overdue users and the sampling ratio is 30%, 30 user samples are selected (the actual operation is rounded up). The selected samples are labeled in detail to generate a labeling data set, and the labeling content includes key information such as overdue type, risk level, and fee collection priority, which provides a reliable basis for subsequent data analysis and decision-making.
[0107] In specific implementation, the labeling condition of the labeling data set is determined in the following manner:
[0108] The labeling conditions that may be involved in all the labeling data sets are combed, such as high arrearage frequency, large arrearage amount proportion, long arrearage duration, recent abnormal payment fluctuation, etc. For each labeling condition, its corresponding value range or judgment standard is determined. For example, "high arrearage frequency" is defined as the number of arrearage times reaching 3 times and above in the past half year; "large arrearage amount proportion" is set as the arrearage amount accounting for 30% and above of the total electricity payment, etc.
[0109] In specific implementation, the labeling conditions are assigned weights to obtain a set of weighted labeling conditions in the following manner:
[0110] Experts in the power industry and data analysts can jointly discuss or use existing relevant data analysis algorithms, which are not specifically limited here; according to the importance of each labeling condition to the user arrearage risk assessment, each labeling condition is given a corresponding weight. For example, it is assessed that "large arrearage amount proportion" has the most critical impact on arrearage risk, and it is given a weight of 0.4; "long arrearage duration" is second, and it is given a weight of 0.3; "high arrearage frequency" is given a weight of 0.2; "recent abnormal payment fluctuation" is given a weight of 0.1, etc., and the sum of all weights is 1.
[0111] In specific implementation, the set of weighted labeling conditions is subjected to membership function construction to obtain a comprehensive membership degree in the following manner:
[0112] For each labeling condition in the set of weighted labeling conditions, a suitable membership degree function is constructed to describe the degree to which the user belongs to a certain labeling category under this condition. Taking "high arrearage frequency" as an example, if the number of arrearage times is divided into low, medium and high three levels, the following membership degree function can be constructed: when the number of arrearage times is less than 2, the membership degree of "low arrearage frequency" is 1, and the membership degrees of "medium" and "high" are 0; when the number of arrearage times is between 2 and 4, the membership degree of "medium arrearage frequency" gradually increases from 0 to 1, and the membership degrees of "low" and "high" change accordingly; when the number of arrearage times is greater than 4, the membership degree of "high arrearage frequency" is 1, and the membership degrees of "low" and "medium" are 0. Other labeling conditions are constructed in the same way.
[0113] When a certain arrears user triggers two or more labeling conditions, according to its specific circumstances, the membership degree value under each condition is calculated according to the membership function. For example, the user has been in arrears for 4 times in the past half year, and the membership degree of belonging to "high arrears frequency" is calculated to be 0.8 through the membership function of "high arrears frequency". Then, the membership degree values under each condition are multiplied by the corresponding weight and accumulated to obtain the comprehensive membership degree. Assuming that the membership degrees of the user under the conditions of "arrears amount ratio is large", "arrears duration is too long", and "recently there is abnormal fluctuation in payment" are 0.6, 0.7, and 0.4 respectively, then the comprehensive membership degree = 0.2x0.8+0.4x0.6+0.3x0.7+0.1x0.4=0.63.
[0114] In specific implementation, the manner of classifying the labeling data set based on the comprehensive membership degree to obtain the arrears user labeling information is:
[0115] The arrears user labeling information corresponding to different comprehensive membership degree intervals is pre-set. For example, the comprehensive membership degree greater than 0.8 is labeled as "high risk urgent collection"; between 0.6 and 0.8 is labeled as "medium risk focus attention"; and less than 0.6 is labeled as "low risk regular reminder". According to the calculated comprehensive membership degree, the final labeling result of the arrears user is determined, so as to realize reasonable determination of the arrears user labeling and ensure the consistency of the arrears user labeling information.
[0116] Step S230, constructing a decision structure based on the arrears user labeling information to obtain an arrears user risk classification decision tree.
[0117] In this step, the arrears user risk classification decision tree is constructed based on the labeling result, which can realize accurate classification and hierarchical division of the arrears user risk. Through the decision tree, the system can automatically judge the risk level of the user according to the user characteristics, provide scientific and visual collection strategy basis for the power company, thereby improving the collection efficiency and reducing the arrears risk.
[0118] In this step, the manner of constructing a decision structure based on the arrears user labeling information to obtain an arrears user risk classification decision tree includes:
[0119] Determining the split feature of the arrears user, classifying the arrears user labeling information based on the split feature to obtain a risk level classification of the arrears user;
[0120] Constructing a decision structure for the risk level classification to obtain the arrears user risk classification decision tree.
[0121] In specific implementation, the manner of determining the split feature of the arrears user, classifying the arrears user labeling information based on the split feature to obtain a risk level classification of the arrears user includes:
[0122] By analyzing the delinquent user label information, determine the feature that maximizes the "purity" of the delinquent user label information as the split feature, and then classify the delinquent user label information to obtain the risk level classification. For example, in the decision tree construction, compare the delinquent amount and delinquent frequency and other features, select the feature (such as delinquent amount) that can concentrate high-risk users in the same sub-node as the split feature, and maximize the risk category distinction of the sub-node samples by testing different split points (such as delinquent amount 500 yuan), thereby realizing accurate risk level division.
[0123] In specific implementation, the risk level classification is constructed into a decision structure to obtain the delinquent user risk grading decision tree in the following manner:
[0124] The decision tree is gradually constructed based on the risk level classification, and each leaf node is marked with the corresponding risk category (such as high risk, low risk). Then, pruning optimization is performed, starting from the leaf node, trying to cut off some sub-trees and replacing them with leaf nodes representing the majority of sample categories, verifying the pruning effect through the test set, and if the classification accuracy rate does not decrease or increases, the pruning operation is retained, and the optimization is continued until the decision tree is simple and adaptable. Finally, the optimized decision tree is applied to new delinquent user data, and from the root node, the user feature value is judged according to the rules layer by layer until the leaf node is reached to determine the risk category of the user.
[0125] Step S240, correcting the delinquent user risk grading decision tree to obtain a corrected grading decision tree.
[0126] In this step, the delinquent user risk grading decision tree is corrected, which can further optimize the classification accuracy and adaptability of the decision tree. By introducing additional evidence or rules (such as special case correction based on D-S evidence theory), the classification results of the decision tree are adjusted, thereby more accurately reflecting the actual risk situation of the delinquent user. This correction process makes up for the limitations of a single decision tree and improves the system's ability to handle complex situations.
[0127] In some embodiments, the delinquent user risk grading decision tree is corrected to obtain a corrected grading decision tree, including:
[0128] Determine the evidence source of the delinquent user, statistically analyze the evidence source, and obtain a probability distribution function;
[0129] Fuse the evidence source and the probability distribution function to obtain a probability distribution result;
[0130] Based on the probability distribution result, the delinquent user risk grading decision tree is corrected to obtain the corrected grading decision tree.
[0131] In implementation, the manner of determining the evidence source of the delinquent user is:
[0132] All kinds of information related to the delinquent user are comprehensively collected. The payment history details of the user are covered, such as whether the payment is on time in the past year, whether there is continuous delinquency, the trend of the delinquent amount, whether the delinquent amount gradually increases, decreases or fluctuates greatly in the past few months, the economic status related information of the user, and the communication feedback situation of the user with the power company, the response attitude to the payment notice, whether the user actively explains the reason for delinquency, etc. These information will be used as important evidence source for subsequent analysis.
[0133] In implementation, the manner of statistically analyzing the evidence source to obtain the probability distribution function is:
[0134] For each determined evidence source, combined with the experience of power industry experts and a large number of historical data statistical analysis, the basic probability distribution function is set for different classification results (high-risk delinquent user, medium-risk delinquent user, low-risk delinquent user). For example, for the evidence source of "payment history", if a user has been paying on time most of the time in the past year, only occasionally a small amount of delinquency and can be paid in time, it is analyzed that the basic probability distribution value of the low-risk delinquent user is 0.7, the medium-risk is 0.2, and the high-risk is 0.1. Each evidence source is assigned a corresponding basic probability distribution value for different classification results according to its own characteristics and the degree of association with delinquency risk in a similar manner, and the sum of all basic probability distribution values is 1.
[0135] In implementation, the manner of modifying the delinquent user risk classification decision tree based on the probability distribution result is:
[0136] The basic probability distribution functions of multiple evidence sources are fused by using the combination rule of D-S evidence theory. Taking two evidence sources as an example, assuming that the basic probability distribution values of evidence source A for high-risk, medium-risk and low-risk delinquent users are m1(H), m1(M) and m1(L) respectively, and the basic probability distribution values of evidence source B are m2(H), m2(M) and m2(L) respectively. According to the combination rule, the information of the two evidence sources is integrated to calculate the basic probability distribution values of the fused user belonging to different risk categories. For example, for high-risk delinquent users, the fused basic probability distribution value m(H) is obtained by a specific calculation method (although a complex formula is not used here, the actual calculation follows the D-S combination rule). After all the evidence sources are sequentially combined, the basic probability distribution result after synthesizing multiple evidences is obtained.
[0137] According to the basic probability assignment result obtained after evidence combination, the final classification of the special overdue user is determined. The basic probability assignment values corresponding to different risk categories are compared, and the user is classified into the category with the maximum basic probability assignment value. For example, if the basic probability assignment value of the user after fusion belongs to the high-risk overdue user is the maximum, even if the decision tree originally classifies the user as a medium-risk or low-risk, it will be corrected to a high-risk overdue user. In this way, various complex factors are considered to effectively correct the classification result of the decision tree, and the accuracy of the classification of the special overdue user is improved.
[0138] In the above example embodiment, the way to obtain the corrected hierarchical decision tree is introduced, and the way to obtain the decision path atlas after the corrected hierarchical decision tree is obtained is introduced as follows:
[0139] In some embodiments, after the corrected hierarchical decision tree is obtained, the method further includes:
[0140] Based on the corrected hierarchical decision tree, visualization is performed to obtain a decision path atlas.
[0141] In specific implementation, the way to obtain the decision path atlas based on the corrected hierarchical decision tree is:
[0142] The key feature influence factors and the decision process of the corrected hierarchical decision tree are presented in the form of an atlas through visualization technology, and the action path of each key feature on the decision result is intuitively displayed, wherein the visualization technology can use decision tree visualization tools (such as Graphviz), data visualization libraries (such as Matplotlib, Seaborn), or interactive visualization platforms (such as Plotly, D3.js) to generate and display the corrected hierarchical decision tree and the decision path atlas. The power company staff can clearly understand the decision basis of the model, which is convenient for manual review and adjustment, thereby significantly enhancing the explainability and credibility of the model.
[0143] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of the embodiments of the present disclosure can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In this distributed scenario, one of the multiple devices can only execute one or more steps in the method of the embodiments of the present disclosure, and the multiple devices can interact with each other to complete the method.
[0144] It is to be understood that the foregoing description is directed to example embodiments of the disclosure. Other embodiments fall within the scope of the following claims. In some cases, the actions or steps recited in the claims can be performed in a different order and still achieve desirable results. Additionally, processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve desirable results. In certain implementations, multitasking and parallel processing can be advantageous.
[0145] Based on the same inventive concept, the disclosure also provides an electricity payment collection grading device corresponding to any of the above-mentioned embodiment methods.
[0146] Reference Figure 3 The electricity payment collection grading device comprises:
[0147] The cluster center determination module 310 is configured to determine the overdue feature information of the overdue user, classify the overdue user based on the overdue feature information, and obtain a first cluster center;
[0148] The label information determination module 320 is configured to label the overdue user in the first cluster center, and obtain overdue user label information;
[0149] The decision tree determination module 330 is configured to construct a decision structure based on the labeling result, and obtain an overdue user risk grading decision tree;
[0150] The decision tree correction module 340 is configured to correct the overdue user risk grading decision tree, and obtain a corrected grading decision tree.
[0151] In the present exemplary embodiment, the cluster center determination module 310 is specifically configured to:
[0152] determining arrears feature information of the arrears user, wherein the arrears feature information comprises a time decay type feature, an arrears frequency feature, an arrears amount proportion feature, an arrears time frequency feature and an arrears behavior feature; quantifying the time decay type feature, the arrears frequency feature, the arrears amount proportion feature, the arrears time frequency feature and the arrears behavior feature to obtain a multi-dimensional feature vector; determining an arrears user data point of the multi-dimensional feature vector, wherein the arrears user data point comprises a plurality of arrears user sub-data points; determining a target arrears user sub-data point from the plurality of arrears user sub-data points and taking the target arrears user sub-data point as a third clustering center; determining Mahalanobis distances of the third clustering center and other arrears user sub-data points, and taking the arrears user sub-data point with the farthest Mahalanobis distance as a fourth clustering center; assigning the plurality of arrears user sub-data points to the third clustering center and the fourth clustering center based on a minimum distance algorithm; calculating a mean of the arrears user sub-data points of the third clustering center and the fourth clustering center to obtain a second clustering center; constructing a similarity graph based on the second clustering center to obtain a Laplacian matrix; performing eigenvalue decomposition on the Laplacian matrix to obtain an eigenvector matrix; normalizing the eigenvector matrix to obtain a normalized eigenvector matrix; and clustering the normalized eigenvector matrix to obtain the first clustering center.
[0153] In the example embodiment, the labeling information determination module 320 is specifically configured to:
[0154] determine a difference of the arrears user data points in the first clustering center, determine a sampling ratio based on the difference, sample and label the arrears user data points in the first clustering center based on the sampling ratio to obtain a labeled data set, determine a labeling condition of the labeled data set, assign weights to the labeling condition to obtain a weighted labeling condition set, construct a membership function for the weighted labeling condition set to obtain a comprehensive membership degree, and classify the labeled data set based on the comprehensive membership degree to obtain the arrears user labeling information.
[0155] In the example embodiment, the decision tree determination module 330 is specifically configured to:
[0156] determine a split feature of the arrears user, classify the labeling result based on the split feature to obtain a risk level classification of the arrears user, and construct a decision structure for the risk level classification to obtain a decision tree for risk grading of the arrears user.
[0157] In the example embodiment, the decision tree correction module 340 is specifically configured to:
[0158] Determine the evidence source of the delinquent user, statistically analyze the evidence source to obtain a probability distribution function, fuse the evidence source and the probability distribution function to obtain a probability distribution result, and correct the risk classification decision tree of the delinquent user based on the probability distribution result to obtain the corrected classification decision tree.
[0159] For the convenience of description, the above apparatus is described in various modules in terms of functions. Of course, the functions of the modules can be implemented in one or more software and / or hardware when implementing the present disclosure.
[0160] The apparatus of the above embodiments is used to implement the corresponding electricity charge collection classification method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be described here.
[0161] Based on the same inventive concept, the present disclosure also provides an electronic device corresponding to the method of any of the above embodiments, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the electricity charge collection classification method of any of the above embodiments when executing the program.
[0162] Figure 4 A more specific hardware structure of an electronic device provided by the present embodiment is shown, which can include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 for communication within the device.
[0163] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits, etc., for executing related programs to implement the technical solutions provided by the present embodiment.
[0164] The memory 1020 can be implemented by a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs, and when the technical solutions provided by the present embodiment are implemented by software or firmware, the related program codes are stored in the memory 1020 and executed by the processor 1010.
[0165] The input / output interface 1030 is configured to connect an input / output module to realize information input and output. The input / output module can be configured in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. The input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.
[0166] The communication interface 1040 is configured to connect a communication module (not shown in the figure) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired manner (such as a USB, a network cable, etc.) or a wireless manner (such as a mobile network, WIFI, Bluetooth, etc.).
[0167] The bus 1050 includes a channel to transmit information between various components (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040) of the device.
[0168] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device can also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device can also only include components necessary for implementing the embodiments of the present disclosure, and does not necessarily include all the components shown in the figure.
[0169] The electronic device of the above embodiments is used to implement the corresponding electricity payment collection classification method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which are not described here.
[0170] Based on the same inventive concept, the disclosure also provides a non-transitory computer readable storage medium storing computer instructions for causing the computer to execute the electricity payment collection classification method according to any of the above embodiments.
[0171] The computer readable medium of the present embodiments includes permanent and non-permanent, removable and non-removable media, which can be implemented by any method or technology to store information. The information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0172] The above non-transitory computer readable storage medium can be any available medium or data storage device that can be accessed by a computer, including but not limited to magnetic memory (such as floppy disk, hard disk, magnetic tape, magneto-optical disk (MO) and the like), optical memory (such as CD, DVD, BD, HVD and the like), and semiconductor memory (such as ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid state disk (SSD)) and the like.
[0173] The storage medium of the above embodiments stores computer instructions for causing the computer to perform the electricity payment collection grading method as described in any of the above exemplary method embodiments, and has the beneficial effects of the corresponding method embodiments, which are not repeated here.
[0174] Based on the same inventive concept, the disclosure also provides a computer program product comprising computer program instructions corresponding to the electricity payment collection grading method described in any of the above embodiments. In some embodiments, the computer program instructions can be executed by one or more processors of a computer to cause the computer and / or the processor to perform the electricity payment collection grading method. The processor performing the corresponding step can belong to the corresponding execution subject corresponding to each step in each embodiment of the electricity payment collection grading method.
[0175] The computer program product of the above embodiments is used to cause the computer and / or the processor to perform the electricity payment collection grading method as described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which are not repeated here.
[0176] Those skilled in the art will appreciate that embodiments of the disclosure can be devised for a system, method, or computer program product. Accordingly, the disclosure can be embodied in hardware and / or in software (including firmware, resident software, micro-code, etc.) that runs on a processor such as a computerized platform. Furthermore, the disclosure can be embodied as computer-readable code carried on a computer readable medium of a computer program product. Such program code can be supplied to or downloaded into the computerized platform by, for example, such computer-readable media, a manufacturer of the computerized platform, or an owner of software.
[0177] Any combination of one or more computer readable medium can be utilized. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer readable storage medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0178] A computer readable signal medium can include a propagated data signal with computer readable program code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal can take any of a variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer readable signal medium can be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0179] Program code embodied on a computer readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire line, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0180] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0181] It should be understood that each block of the flowchart and / or block diagram illustrations, and combinations of blocks in the flowchart and / or block diagram illustrations, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0182] These computer program instructions can also be stored in a computer- readable medium that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture including instructions which implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0183] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0184] Further, while operations of the present disclosure are described in a particular order in the figures, it is not necessary to perform the operations in the particular order shown or in all of the described operations are required to produce a desired result. Rather, the steps depicted in the flowcharts can be altered in order of execution. Additionally or alternatively, certain steps can be omitted, combined, performed in a different order, and / or performed in parallel.
[0185] The computer program product of the present application can be a computer program embodied on a non-transitory computer readable medium. The body of computer program instructions can be a source file, object file, executable file, or any other tangible form of computer program instructions. The computer program product can be supplied on a non-transitory computer readable medium such as a floppy disk, a hard disk, a CD ROM, a flash memory, a USB memory, a RAM, a ROM, or any other non-transitory computer readable medium. The computer program product can also be supplied as a digital download, such as from an Internet site, or via the air interface, or via any other suitable digital transmission mechanism.
[0186] It should be noted that although several modules or units of the device for action execution are mentioned in the foregoing detailed description, this division is not mandatory. Indeed, according to the embodiments of the application, the features and functionalities of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functionalities of one module or unit described above can be further divided into embodied by several modules or units.
[0187] It should be understood by those of ordinary skill in the art that the above discussion of any embodiment is merely exemplary in nature and is not intended to imply limitations on the scope of the application, including the claims. Indeed, variations on the techniques described herein can be made that fall within the scope of the application. Additionally, those of ordinary skill in the art will recognize that the steps of the disclosed embodiments can be altered, modified, or combined in any number of ways. For the sake of brevity, only certain aspects of the embodiments have been described in detail. Those skilled in the art will recognize that the techniques described herein can be adapted for use with a number of other application, and that the spheres of use discussed are merely exemplary. Accordingly, this application is intended to embrace all such alterations, modifications, and combinations of the embodiments discussed and other applications that fall within the scope of the appended claims.
[0188] In addition, for the sake of brevity and clarity, well-known power supply and ground connections can not be shown in the drawings and detailed descriptions thereof can not be provided. Further, devices can be shown in block diagram form in order to avoid obscuring the concepts of the present embodiments. It will be appreciated that the embodiments can be implemented in a variety of platforms, including integrated circuit (IC) chips and other components, and that the details provided herein are merely exemplary. It will be appreciated that the embodiments described herein can be implemented in hardware, software, or a combination thereof. It will be appreciated that the details provided herein are merely exemplary and that the scope of the application is not limited to any particular details.
[0189] While the present application has been described in connection with certain embodiments thereof, many modifications, substitutions, and alterations, thereof, will be apparent to those of ordinary skill in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) can use the embodiments discussed.
[0190] It is intended to encompass all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Accordingly, any one or more features of any element in the drawings, or the specification, can be combined to create modifications, equivalents, or alternatives.
[0191] While the principles of the disclosure have been described above in connection with specific embodiments, it is to be understood that this disclosure is not limited to the disclosed embodiments, but is instead applicable to various modifications and equivalent arrangements. The scope of the disclosure encompasses various modifications and equivalent arrangements. The scope of the claims appended hereto is accorded the broadest interpretation so as to encompass all such modifications and equivalent structures and functions.
Claims
1. A tiered method for electricity bill collection, characterized in that, include: Determine the arrears characteristic information of users in arrears, classify the users in arrears based on the arrears characteristic information, and obtain the first cluster center; The users in arrears within the first cluster center are labeled to obtain the labeling information of the users in arrears. Based on the information of the users in arrears, a decision structure is constructed to obtain a risk classification decision tree for users in arrears. The risk classification decision tree for the users in arrears is modified to obtain the modified classification decision tree.
2. The method according to claim 1, characterized in that, The overdue payment characteristics include: time decay characteristics, overdue payment frequency characteristics, overdue payment amount ratio characteristics, overdue payment duration characteristics, and overdue payment behavior characteristics. The step of classifying the delinquent users based on the delinquency feature information to obtain the first cluster center includes: The time decay feature, the overdue payment frequency feature, the overdue payment amount ratio feature, the overdue payment duration feature, and the overdue payment behavior feature are quantified to obtain a multi-dimensional feature vector; Determine the arrears user data points of the multidimensional feature vector, and perform coarse-grained clustering based on the arrears user data points to obtain the second cluster center; A similarity graph is constructed based on the second cluster centers, and the Laplace matrix is obtained. The Laplacian matrix is subjected to eigenvalue decomposition to obtain the eigenvector matrix; The eigenvector matrix is normalized to obtain a normalized eigenvector matrix; Clustering is performed on the normalized feature vector matrix to obtain the first cluster center.
3. The method according to claim 2, characterized in that, The data points of users in arrears include: several sub-data points of users in arrears; The step of performing coarse-grained clustering based on the data points of users with outstanding payments to obtain the second cluster center includes: From a plurality of the aforementioned delinquent user sub-data points, a target delinquent user sub-data point is determined, and the target delinquent user sub-data point is used as the third cluster center; Determine the Mahalanobis distance between the third cluster center and the other delinquent user sub-data points, and take the delinquent user sub-data point with the farthest Mahalanobis distance as the fourth cluster center; The several data points of users with outstanding payments are assigned to the third cluster center and the fourth cluster center based on the minimum distance algorithm; The second cluster center is obtained by calculating the mean value of the arrears user sub-data points of the third cluster center and the fourth cluster center.
4. The method according to claim 2, characterized in that, The step of labeling the delinquent users within the first cluster center to obtain delinquent user labeling information includes: Determine the differences among the data points of the users in arrears within the first cluster center, and determine the sampling ratio based on the differences; Based on the sampling ratio, the data points of the users in arrears within the first cluster center are sampled and labeled to obtain a labeled dataset; The labeling conditions of the labeled dataset are determined, and weights are assigned to the labeling conditions to obtain a weighted labeling condition set; Membership functions are constructed on the weighted labeling condition set to obtain the comprehensive membership degree; The labeled dataset is classified based on the comprehensive membership degree to obtain the labeled information of the users in arrears.
5. The method according to claim 1, characterized in that, The step of constructing a decision structure based on the delinquent user annotation information to obtain a risk classification decision tree for delinquent users includes: Determine the splitting characteristics of the users in arrears, classify the labeling information of the users in arrears based on the splitting characteristics, and obtain the risk level classification of the users in arrears. A decision structure is constructed based on the risk level classification to obtain the risk classification decision tree for the users in arrears.
6. The method according to claim 1, characterized in that, The step of revising the risk classification decision tree for the delinquent users to obtain the revised classification decision tree includes: Identify the evidence sources for the users in arrears, perform statistical analysis on the evidence sources, and obtain a probability allocation function; The evidence source and the probability allocation function are fused to obtain the probability allocation result; The risk classification decision tree for the arrears user is modified based on the probability allocation result to obtain the modified classification decision tree.
7. The method according to claim 1, characterized in that, The method further includes: The decision path map is obtained by visualizing the modified hierarchical decision tree.
8. A tiered electricity bill collection device, characterized in that, include: The cluster center determination module is configured to determine the arrears feature information of users in arrears, classify the users in arrears based on the arrears feature information, and obtain the first cluster center; The labeling information determination module is configured to label the users in arrears within the first cluster center to obtain the labeling information of the users in arrears. The decision tree determination module is configured to construct a decision structure based on the overdue user annotation information to obtain an overdue user risk classification decision tree. The decision tree correction module is configured to correct the risk classification decision tree of the arrears users to obtain the corrected classification decision tree.
9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions for causing the computer to perform the method of any one of claims 1 to 7.