Urban village power consumer portrait construction method and system based on three-channel grouping
By employing a three-channel clustering method and an optimal text feature prediction model, the problem of identifying electricity user types in urban villages was solved, resulting in more accurate electricity user profiles and improving the accuracy of identifying electricity-sensitive users.
Patent Information
- Application Number
- CN202511094136.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies are insufficient to accurately identify the types of electricity users in urban villages, resulting in inaccurate electricity user profiles, especially in complex electricity usage environments where it is difficult to identify electricity bill-sensitive users.
A three-channel clustering approach is adopted to collect user behavior characteristics from power service data sources, classify users into low, medium and high activity levels according to their activity, and predict electricity price sensitivity through optimal text features and a three-channel prediction model to construct an accurate power user profile.
It improves the accuracy of identifying the electricity cost sensitivity of urban village electricity users, builds a more accurate electricity user profile, and supports targeted optimization of electricity services.
Smart Images

Figure CN120995171A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of power user behavior analysis, specifically involving a method and system for constructing power user profiles in urban villages based on three-channel clustering. Background Technology
[0002] Electricity users are generally divided into electricity bill-sensitive users and non-sensitive users. Electricity bill-sensitive users are those who proactively submit work orders through customer service hotlines or online service platforms to file complaints, inquiries, or raise objections in situations such as abnormal electricity bills, billing disputes, or electricity price adjustments. Non-sensitive users, on the other hand, are those who are not proactive in providing feedback in situations involving electricity bill fluctuations or disputes. In complex electricity usage environments such as urban villages, the diverse electricity consumption behaviors and unique billing models of the user group make the creation of electricity user profiles more complex than in ordinary communities. The difficulty in identifying electricity bill-sensitive users is significantly increased due to data sparsity and scenario complexity, making it difficult to optimize corresponding electricity service strategies for electricity users in urban villages.
[0003] Existing methods for constructing electricity user profiles in urban villages rely too heavily on manual statistical analysis based on experience. Furthermore, these methods are unfamiliar with the unique power grid structure of urban villages, making it difficult to predict the high density of transient populations and frequent changes in electricity demand. Additionally, the complex and diverse structure of electricity users in urban villages means that using a single model cannot accurately identify users who are sensitive to electricity charges but not those who are not, resulting in electricity user profiles that deviate significantly from reality. Summary of the Invention
[0004] This application proposes a method and system for constructing a profile of urban village electricity users based on three-channel clustering, which can solve the problem in the prior art that it is difficult to accurately predict the composition of electricity user types in urban villages, resulting in an inaccurate profile of urban village electricity users.
[0005] The first aspect of this application provides a method for constructing a user profile of urban village electricity users based on three-channel clustering, the method comprising:
[0006] Collect several user behavior characteristics related to electricity users in urban villages from electricity service data sources;
[0007] According to the preset user group segmentation rules, urban village electricity users are divided according to user activity, and corresponding user feature groups are constructed based on the user behavior characteristics; wherein, the segmentation results include low-activity users, medium-activity users, and high-activity users;
[0008] By calculating feature importance, the features of the user feature group are filtered to obtain the optimal text features;
[0009] Based on a pre-defined three-channel prediction model, the sensitivity of urban village electricity users to electricity fees is predicted, and a profile of urban village electricity users is constructed; wherein, the three-channel prediction model is optimized through the optimal text features.
[0010] The above scheme first collects multiple user behavior features representing user concern about electricity costs from various data sources related to electricity bill payment and electricity demand services. Then, based on preset user group segmentation rules, urban village electricity users are divided into low-activity, medium-activity, and high-activity users according to their activity levels, and corresponding user feature groups are constructed. These user feature groups can centrally reflect whether users are sensitive or insensitive to electricity costs. The user feature groups are then filtered, selecting features with higher importance to achieve data dimensionality reduction and improve the efficiency of subsequent user prediction. The model is optimized based on the optimal text features to ensure that each channel has a high sensitivity to users with different activity levels, focusing on predicting different types of users. The trained three-channel prediction model is used to make multiple predictions on the sensitivity of urban village electricity users to electricity costs, identifying electricity cost-sensitive users and constructing a more accurate profile of urban village electricity users.
[0011] In one possible implementation of the first aspect, urban village electricity users are divided according to user activity based on a preset user group segmentation rule, and a corresponding user feature group is constructed based on the user behavior characteristics of the segmentation results, specifically as follows:
[0012] According to the user group segmentation rules, the user activity of urban village power users is assessed by detecting whether users have engaged in follow-up or supervision behavior and by counting the number of times they make calls to request electricity services. Urban village power users are divided into low-activity users, medium-activity users, and high-activity users.
[0013] Based on the electricity request service tickets of each type of active user, the user feature groups corresponding to low-activity users, medium-activity users, and high-activity users are constructed by integrating the user behavior characteristics from three aspects: ticket duration, electricity cost, and business sensitivity.
[0014] The user characteristic groups corresponding to the medium-activity users and the high-activity users are also associated with reminder and supervision information.
[0015] The above scheme measures user activity by detecting whether users engage in follow-up or supervision activities and the frequency of making calls to request electricity services, thereby classifying users into low-activity, medium-activity, and high-activity users. Then, based on user behavior characteristics, corresponding user feature groups are constructed for different activity levels to more accurately represent the characteristics of different users.
[0016] In one possible implementation of the first aspect, the features of the user feature group are filtered by calculating feature importance to obtain the optimal text features, specifically:
[0017] The sensitive text information of the user feature group is subjected to feature extraction and vectorization to obtain text features;
[0018] The importance scores of the text features and the statistical features of the user feature group are calculated using an XGBoost tree to obtain the importance scores of each text feature.
[0019] Based on the importance score, the text features and the statistical features are sorted in descending order, and then the top first threshold features are selected as the optimal text features.
[0020] The above scheme evaluates the user feature group data through a tree structure, selects the optimal text features that better reflect the user's sensitivity to electricity costs, and achieves data dimensionality reduction and selection of the most representative features from multi-dimensional features.
[0021] In one possible implementation of the first aspect, the sensitive text information of the user feature group is subjected to feature extraction and vectorization to obtain text features, specifically:
[0022] Simultaneously considering both single words and two consecutive words, a logarithmic transformation is performed on the word frequencies of the sensitive text information;
[0023] Based on a pre-defined standard vocabulary, the logarithmic transformation results are standardized and features are extracted. Redundant and useless features are removed during the feature extraction process to obtain text features.
[0024] The above scheme considers both single words and consecutive words, enriching the feature representation; it mitigates the influence of high-frequency words through logarithmic transformation to balance the data distribution; and it standardizes the data through a standard vocabulary to ensure the consistency of feature names.
[0025] In one possible implementation of the first aspect, the sensitive text information specifically refers to:
[0026] The electricity request service order and the follow-up information are processed by word segmentation and regular expression replacement to obtain a number of work order terms;
[0027] The number of words and average word length of the work order vocabulary are counted, and the sensitive text information is constructed by combining the work order vocabulary.
[0028] In one possible implementation of the first aspect, based on a pre-defined three-channel prediction model, the sensitivity of urban village electricity users to electricity fees is predicted, and a profile of urban village electricity users is constructed, specifically as follows:
[0029] The electricity user information of urban villages is input into the three-channel prediction model. The electricity cost sensitivity of low-activity users, medium-activity users and high-activity users is predicted by each channel and the judgment threshold of each channel, so as to obtain the prediction label of each electricity user in urban villages; wherein, the prediction label is an electricity cost sensitive label or an electricity cost insensitive label.
[0030] Based on the predicted labels, a profile of electricity users in urban villages is constructed.
[0031] In one possible implementation of the first aspect, the three-channel prediction model is optimized using the optimal text features, specifically as follows:
[0032] A model training set is constructed based on the optimal text features, and the preset prediction model is trained using the model training set.
[0033] During model training, each combination of model parameters is traversed using a grid search method to determine the optimal combination of hyperparameters.
[0034] The three channels of the prediction model are tuned according to the optimal hyperparameter combination to obtain the trained three-channel prediction model; wherein the channels are used to predict the sensitivity of low-activity users, medium-activity users and high-activity users to electricity costs.
[0035] The above approach trains different channels of the model using optimal text features, ensuring that each channel focuses on predicting the corresponding active users, thus greatly improving the accuracy of data prediction.
[0036] In one possible implementation of the first aspect, the electricity cost sensitivity of low-activity users, medium-activity users, and high-activity users is predicted by using each channel and the judgment threshold for each channel, specifically as follows:
[0037] By using each channel and the judgment threshold of each channel, a second threshold prediction is performed on low-activity users, medium-activity users, and high-activity users respectively, and the second threshold prediction results are obtained.
[0038] The prediction results of the second threshold are integrated, and the labels of the integrated results are voted on using a voting method to obtain the prediction label corresponding to each urban village power user.
[0039] The above scheme predicts the electricity cost sensitivity of users with different activity levels through different channels in the model and assigns corresponding prediction labels to them. Furthermore, it integrates multiple prediction results to find the most likely category for each user, accurately identifying users who are sensitive to electricity costs.
[0040] In one possible implementation of the first aspect, several user behavior characteristics related to urban village electricity users are collected from electricity service data sources, specifically:
[0041] Obtain service order information and electricity consumption information of electricity users in urban villages from electricity service data sources;
[0042] Extract the user behavior features related to electricity service request information and follow-up information from the service work order information;
[0043] Extract user behavior features related to electricity billing information and the user's historical electricity usage from the electricity consumption information.
[0044] The above scheme constructs user behavior characteristics by obtaining user service work order information, electricity consumption information, and follow-up and supervision information related to electricity services, which can more accurately describe the user's sensitivity to changes in electricity prices.
[0045] The second aspect of this application provides a system for constructing user profiles of urban villages based on three-channel clustering, the system comprising: a feature extraction module, a user feature cluster construction module, a feature filtering module, and a user profile construction module;
[0046] Among them, the feature extraction module is used to collect several user behavior features related to urban village power users from power service data sources;
[0047] The user feature group construction module is used to classify urban village power users according to user activity based on preset three-channel user group classification rules, and construct corresponding user feature groups based on the user behavior characteristics of the classification results; wherein, the classification results include low-activity users, medium-activity users, and high-activity users;
[0048] The feature filtering module is used to filter the features of the user feature group by calculating the feature importance, so as to obtain the optimal text features;
[0049] The user profile building module is used to predict the sensitivity of urban village electricity users to electricity fees based on a preset three-channel prediction model, and to build a user profile of urban village electricity users; wherein, the three-channel prediction model is optimized by the optimal text features. Attached Figure Description
[0050] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0051] Figure 1 This is a schematic diagram of a specific process for constructing a user profile of urban village electricity users based on three-channel clustering, provided in a certain embodiment of this application;
[0052] Figure 2 This is an indicator curve variation diagram of a method for constructing a profile of urban village power users based on three-channel clustering provided in a certain embodiment of this application;
[0053] Figure 3 This is a graph showing the change of an index curve based on a logistic regression algorithm for a method of constructing a profile of urban village electricity users based on three-channel clustering, provided in a certain embodiment of this application.
[0054] Figure 4 This is a graph showing the change of index curves based on the random forest algorithm for a method of constructing a profile of urban village electricity users based on three-channel clustering, provided in a certain embodiment of this application.
[0055] Figure 5 This is a structural diagram of a system for constructing power user profiles in urban villages based on three-channel clustering, provided in a certain embodiment of this application. Detailed Implementation
[0056] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0057] It should be understood that the step numbers used in the text are for ease of description only and are not intended to limit the order in which the steps are performed.
[0058] First Embodiment
[0059] Electricity user profiles are constructed using data tags. By analyzing massive amounts of data, the data is abstracted into tags, which are then used to concretize the user profile, ultimately forming the user persona. However, in urban village environments, the proportion of electricity bill-sensitive users is significantly higher than in ordinary communities. The difficulty of identification is greatly increased due to the more complex urban village setting. Therefore, traditional user profile construction methods are insufficient to accurately identify the electricity bill sensitivity of urban village users. This application proposes a novel identification method that integrates user behavior and textual semantics to accurately locate potential electricity bill-sensitive users.
[0060] like Figure 1As shown, to address the problem in existing technologies that make it difficult to accurately predict the composition of electricity user types in urban villages, resulting in inaccurate electricity user profiles for urban villages, the first embodiment of this application provides a detailed flowchart of a method for constructing an urban village electricity user profile based on three-channel clustering. This embodiment's method for constructing an urban village electricity user profile based on three-channel clustering includes steps S1 to S4, detailed below:
[0061] Step S1: Collect several user behavior characteristics related to urban village electricity users from the electricity service data source.
[0062] In this embodiment, service order information and electricity consumption information of urban village electricity users are collected from various power service data sources, such as the State Grid system, electricity fee information collection system, call center system, customer relationship management system, marketing business system, and tag library. Feature fields are then selected as user behavior features, as follows:
[0063] Extract feature fields such as customer code, electricity category code, power supply unit code, business type code, urban / rural category code, work order acceptance time, and work order content from the obtained service work order information.
[0064] Extract feature fields such as service request start time and service request end time from the electricity service request information.
[0065] The system extracts characteristic fields such as electricity status code, load attribute number, contract capacity, property ownership type, and historical power outage frequency from the user's historical electricity consumption status; among them, property ownership type and historical power outage frequency are data specific to urban villages.
[0066] Extract feature fields such as total electricity consumption, amount due, amount actually received, late payment penalties, late payment penalties actually received, and electricity bill amount from the electricity bill information.
[0067] Extract feature fields such as the content of the follow-up and supervision and the type of follow-up and supervision business from the follow-up and supervision information.
[0068] Step S2: According to the preset three-channel user group division rules, urban village power users are divided according to user activity, and corresponding user feature groups are constructed based on the user behavior characteristics.
[0069] Based on preset user group segmentation rules, the system detects whether urban village electricity users engage in follow-up or supervision activities and counts the number of calls made requesting electricity services. This categorizes electricity users into low-activity, medium-activity, and high-activity users. Because these users are divided into three different activity levels, urban village electricity users are thus divided into three channels.
[0070] Optionally, Table 1 below shows the user group segmentation results obtained according to the user group segmentation rules in this application embodiment:
[0071] Table 1. User segmentation based on activity level across three channels.
[0072]
[0073] The ratio of negative to positive samples is the ratio of electricity-insensitive users to electricity-sensitive users based on historical data.
[0074] Subsequently, stratified sampling was used to divide users with different activity levels into training and test sets at a ratio of 9:1, providing data support for the subsequent construction of user feature groups. Furthermore, during the data partitioning process, it was ensured that the ratio of positive to negative samples was consistent between the training and test sets.
[0075] Because the user behavior characteristics of the three types of active users are inconsistent, it is necessary to integrate the user behavior characteristics so that the constructed user feature group can more accurately represent the characteristics of different types of active users.
[0076] Before feature integration, for all users, any work order content containing the field "This work order is a branch center test work order" was directly deleted, along with all corresponding content, to ensure data authenticity. Duplicate data and records with abnormal times were also deleted.
[0077] In addition, for medium- and high-activity users, the number of work order records corresponding to customer codes is counted, and data with more than 10 work order records are deleted to ensure the accuracy of the model built later.
[0078] Then, based on the electricity request service work orders of each type of active user, the user feature groups corresponding to low-activity users, medium-activity users, and high-activity users are constructed by integrating the user behavior characteristics from three aspects: work order duration, electricity cost, and business sensitivity.
[0079] For inactive users, the user feature group focuses on lightweight modeling of users with only one work order record. Therefore, users with only one work order record are selected from the electricity request service work orders, and the following user feature group is constructed based on the user behavior characteristics:
[0080] (1) Basic statistical characteristics: call duration is normalized and service type is one-hot encoded;
[0081] (2) Time characteristics: The work order processing time is divided into month, day and hour, and the work order recording time is marked as early / mid / late ten days of the month;
[0082] (3) Electricity consumption characteristics: classification of the first digit of the electricity consumption type code and mapping of positive examples; historical power outage frequency, Z-Score standardization; time window statistics, calculating the power outage frequency in the past 3 months, 6 months and 12 months, capturing short-term and long-term impacts;
[0083] Because there are many types of electricity use, including residential electricity use (type number 200), rural residential electricity use (type number 201), and urban residential electricity use (type number 202), all of which fall under the broad category of residential electricity use, a new classification is performed using the first digit of the electricity use type number to reduce feature dimensions and redundancy. The proportional mapping refers to calculating the proportion of electricity-sensitive users for each type of electricity use number.
[0084] (4) Characteristics of electricity bill information: The range of the amount is from single digits to millions or even tens of millions. Therefore, logarithmic transformation of the amount receivable / receivable and calculation of the difference after logarithmic transformation are required to enhance the interpretability of the data.
[0085] (5) Sensitive text information: Sensitive information (such as account number and mobile phone number) in the electricity request service order is processed by word segmentation and regular expression replacement.
[0086] For moderately active users, deep time series analysis and business sensitivity scoring were introduced. Therefore, users with 2 to 10 service requests were selected from the electricity request service orders, and after removing reminder and follow-up information, a corresponding user characteristic group was constructed:
[0087] (1) Time series analysis: Calculate the minimum / maximum / mean of customer call intervals and the standard deviation of the date distribution;
[0088] (2) Business sensitivity characteristics: positive sample ratio mapping based on the topic of telephone request service work order;
[0089] (3) Analysis of power supply unit coding: classification of coding length and its distribution ratio;
[0090] (4) In-depth characteristics of accounts receivable electricity bill information: difference in late payment penalties and average number of monthly records;
[0091] (5) Sensitive text information: Sensitive information (such as account number and mobile phone number) in the electricity request service order is processed by word segmentation and regular expression replacement.
[0092] For highly active users, the system integrates reminder and supervision information and complex business logic. It filters users with 2 to 10 service requests from the electricity demand service orders, and after associating these with reminder and supervision information, constructs corresponding user characteristic groups.
[0093] (1) Derivative features of expedited and supervised information: Expand the types of supervised business into multiple columns and classify customer names (such as institutions and merchants) to generate unique hot codes;
[0094] (2) In-depth time analysis: median and standard deviation of customer call intervals, maximum number of calls per month;
[0095] (3) Enhanced topic sensitivity: Calculate multi-topic ratio scores based on training set data and fill missing values based on the global median;
[0096] Specifically, for example, a work order might read, "[One household without electricity] A non-residential customer reported one household without electricity. After guiding the customer to inspect, the equipment malfunction and its ownership cannot be determined. Please investigate on-site." In this case, "one household without electricity" in the brackets represents the specific topic. The training set contains all work order records and the total number of different topics, as well as the electricity cost sensitivity ratio for each topic. It's important to note that the same customer may have multiple work order records, potentially corresponding to different topics; the same topic appearing multiple times may correspond to multiple users, and these users may correspond to different electricity cost sensitivity tags. Using the median of the entire dataset for imputation refers to using the median of the proportion scores of all topics in the entire dataset to impute users with missing proportion scores for multiple topics.
[0097] (4) Sensitive text information: Sensitive information (such as account number and mobile phone number) in electricity request service work orders and reminder and supervision information is processed by word segmentation and regular expression replacement to obtain the corresponding work order vocabulary, and the number of words and average word length of work order vocabulary are counted.
[0098] Step S3: By calculating the feature importance, the features of the user feature group are filtered to obtain the optimal text features.
[0099] First, the sensitive text information of the user feature group is extracted and vectorized using the TF-IDF algorithm to obtain text features of more than 10,000 dimensions.
[0100] The specific parameter settings of the TF-IDF algorithm in this embodiment are as follows: both single words and two consecutive words are considered to enrich the feature expression; the term frequency (TF) is used to represent the sensitive text information, but the inverse document frequency (IDF) is not applied; a logarithmic transformation is performed on the term frequency to reduce the influence of high-frequency words; a preset standard vocabulary is forced to be used to standardize the logarithmic transformation result and extract features during feature extraction, and redundant and useless features are removed during the feature extraction process to reduce dimensionality and obtain text features.
[0101] Among them, the TF-IDF algorithm is a commonly used weighting technique for information retrieval and data mining. TF stands for Term Frequency, and IDF stands for Inverse Document Frequency.
[0102] In addition, feature extraction is based on user feature groups corresponding to low-activity users, medium-activity users, and high-activity users.
[0103] After obtaining the text features, xgb.DMatrix is used to convert the text features and statistical features into the XGBoost internal format, specifying the feature column names for subsequent analysis of feature importance. The statistical features are those features within the user feature group excluding the text features.
[0104] Then, the feature importance of the text features and the statistical features of the user feature group are calculated using an XGBoost tree to obtain the importance score of each text feature.
[0105] Specifically, for the parameter settings of the XGBoost tree, a strong L2 regularization is set to prevent overfitting, while the sampling ratio of urban village power user samples and the text features, as well as the depth of the XGBoost tree and the weight of child nodes, are set.
[0106] For XGBoost trees, the number of iterations is set to 6666, but early stopping is enabled. The error rates on both the training and validation sets are monitored to prevent overfitting. Training information is output every 200 rounds to facilitate timely adjustments to training parameters.
[0107] The statistical features of the text features and the user feature group are input into the XGBoost tree-based model to evaluate the feature importance. The importance score of each feature is obtained by setting the parameters mentioned above.
[0108] Then, according to the importance score, the text features and the statistical features are sorted in descending order, and the top 5% of features are selected as the optimal text features.
[0109] For example, the top 10 characteristics of highly active users selected in this application embodiment are shown in Table 2 below:
[0110] Table 2 Importance Scoring Table of Characteristics of Highly Active Users
[0111]
[0112] Here, `counts_jobinfo_jianqu_comm` represents the integer obtained by subtracting the number of times the same customer appears in the electricity request work order record from the number of times they appear in the call record; "Caller ID" is one dimension of the text feature obtained through TF-IDF conversion; `counts_of_09flow` represents the statistical count of the electricity bill record corresponding to the customer's unique code; `sum_len_of_content` represents the total number of text words in all work order records of a single customer; `mean_topic_ratio` represents the average positive sample ratio of all work order topics associated with each customer; `has_biao09` indicates whether the customer has an electricity bill record (1 represents yes, 0 represents no); `std_holding_time_seconds` represents the standard deviation of the call record duration of a single customer in seconds; `max_date_ratio` represents the maximum positive sample ratio of the date of all work order records associated with the customer; `mean_date_ratio` represents the average positive sample ratio of the date of all work order records associated with the customer; and `sum_date_ratio` represents the sum of the positive sample ratios of the date of all work order records associated with the customer.
[0113] Furthermore, Table 3 provides dimensional information on the characteristics of the three types of active users:
[0114] Table 3 Three-channel feature dimensions
[0115]
[0116] Step S4: Based on the preset three-channel prediction model, predict the sensitivity of urban village electricity users to electricity fees and construct a profile of urban village electricity users.
[0117] The electricity user information of urban villages is input into a trained three-channel prediction model. Then, the electricity cost sensitivity of low-activity users, medium-activity users and high-activity users is predicted by the judgment threshold of each channel in the model. The predicted label of each urban village electricity user is output.
[0118] Specifically, the three-channel prediction model is trained based on the XGBoost model. First, a model training set is constructed based on the obtained optimal text features. The XGBoost model is then trained using this training set. During training, a grid search method is used to iterate through each combination of model parameters to determine the optimal hyperparameter combination. Table 4 below shows the parameter tuning table for the XGBoost model:
[0119] Table 4. Parameter tuning table for the XGBoost model
[0120]
[0121]
[0122] In determining the optimal combination of hyperparameters, the number of training epochs is dynamically controlled using early stopping to prevent overfitting. In this embodiment, the maximum number of training epochs is set to 50.
[0123] Then, the three channels of the XGBoost model are tuned according to the optimal hyperparameter combination to obtain the trained three-channel prediction model. These three channels are used to predict the sensitivity of low-activity, medium-activity, and high-activity users to electricity costs, respectively.
[0124] In addition, the XGBoost model is trained independently five times, with the seed parameter (seed=i) randomly set each time to generate diverse base learners. The training data is converted into DMatrix format, and the predicted probability output is optimized using a binary classification objective function (binary:logistic).
[0125] The trained XGBoost model performs five independent predictions for low-activity users, medium-activity users, and high-activity users through its three channels, resulting in five independent predictions.
[0126] The five prediction results were integrated using Bagging, and then the final prediction label was generated using a simple voting method.
[0127] The model prediction process also uses a threshold value corresponding to each channel to classify users, predicting those with values greater than or equal to the threshold as electricity cost-sensitive users. This threshold value is derived from test data trained on a model with the goal of maximizing the F1 score.
[0128] The F1 score is a metric used to comprehensively evaluate the performance of a classification model. It combines precision and recall, and its expression is as follows:
[0129]
[0130] In the formula, precision is the proportion of the number of samples correctly predicted as positive to the total number of samples predicted as positive by the classifier; recall is the proportion of the number of samples correctly predicted as positive to the total number of samples that were actually positive.
[0131] Specifically, the prediction process of the three-channel prediction model is actually to obtain a binary classification result, labeling urban village electricity users with either a price-sensitive label or a price-insensitive label. For a single channel of the three-channel prediction model, its final prediction probability is:
[0132]
[0133] In the formula, p i To predict the probability that a sample belongs to the positive class, this application embodiment determines the probability of predicting a sample as a positive or negative sample by setting a judgment threshold. t (xi) represents the prediction contribution of the t-th decision tree to sample xi, where t is the decision tree index in the XGBoost iteration process.
[0134] Furthermore, embodiments of this application provide some indicator curve changes related to highly active users. The graph shows the changes in the PR curve, where P is the precision rate mentioned above and R is the recall rate mentioned above. Figure 2 The graph shows the PR curve using a single XGBoost algorithm, where the horizontal axis represents recall and the vertical axis represents precision. The solid line is the PR curve and the dashed line is the baseline (representing the horizontal part of the curve). The graph shows that the F1 score reaches its optimal value of 0.97 when both P and R are close to 1.
[0135] Figure 3 The graph shows the changes in the index curve based on the logistic regression algorithm. The horizontal axis represents recall, the vertical axis represents precision, the solid line represents the PR curve, and the dashed line represents the baseline. The optimal F1 score is 0.64.
[0136] Figure 4 The graph shows the changes in the index curves based on the random forest algorithm. The horizontal axis represents recall, the vertical axis represents precision, the solid line represents the PR curve, and the dashed line represents the baseline. The optimal F1 score is 0.85.
[0137] Table 5 below provides a comprehensive performance evaluation of the models under the three algorithms mentioned above:
[0138] Table 5. Comprehensive Performance Evaluation of the Model
[0139]
[0140] Therefore, the three-channel prediction model is evaluated using the F1 score, precision, and recall. The performance of the three-channel and single-channel models under these performance evaluation metrics after integration is shown in Table 6.
[0141] Table 6 Performance Comparison
[0142]
[0143] The table above shows that the prediction accuracy of the model based on three channels is higher than that of the single-channel model.
[0144] Finally, based on the final predicted labels output by the model, the output is the electricity users in urban villages who are sensitive to electricity costs. Based on this, a corresponding profile of electricity users in urban villages is constructed to provide data support for the governance of electricity costs in urban villages.
[0145] Implementing the embodiments of this application has the following beneficial effects:
[0146] This application first collects multiple user behavior features representing user concern about electricity costs from various data sources related to electricity payment and electricity demand services. Then, based on preset user group segmentation rules, urban village electricity users are divided into low-activity, medium-activity, and high-activity users according to their activity levels, and corresponding user feature groups are constructed. These user feature groups can centrally reflect whether users are sensitive or insensitive to electricity costs. The user feature groups are then filtered, selecting features with higher importance to achieve data dimensionality reduction and improve the efficiency of subsequent user prediction. The model is optimized based on the optimal text features to ensure that each channel has a high sensitivity to users with different activity levels, focusing on predicting different types of users. The trained three-channel prediction model is used to make multiple predictions on the sensitivity of urban village electricity users to electricity costs, identifying electricity cost-sensitive users and constructing a more accurate profile of urban village electricity users.
[0147] Second Embodiment
[0148] Furthermore, in order to implement the urban village power user profile construction system based on three-channel clustering corresponding to the above method embodiments, and to achieve the corresponding functions and technical effects, Figure 5 A structural diagram of a system for constructing power user profiles in urban villages based on three-channel clustering is provided. For ease of explanation, only the parts relevant to this embodiment are shown. The system for constructing power user profiles in urban villages based on three-channel clustering provided in this application embodiment includes:
[0149] The feature extraction module 201 is used to collect several user behavior features related to urban village power users from power service data sources.
[0150] In this embodiment, service order information and electricity consumption information of urban village electricity users are collected from various power service data sources, such as the State Grid system, electricity fee information collection system, call center system, customer relationship management system, marketing business system, and tag library. Feature fields are then selected as user behavior features, as follows:
[0151] Extract feature fields such as customer code, electricity category code, power supply unit code, business type code, urban / rural category code, work order acceptance time, and work order content from the obtained service work order information.
[0152] Extract feature fields such as service request start time and service request end time from the electricity service request information.
[0153] The system extracts characteristic fields such as electricity status code, load attribute number, contract capacity, property ownership type, and historical power outage frequency from the user's historical electricity consumption status; among them, property ownership type and historical power outage frequency are data specific to urban villages.
[0154] Extract feature fields such as total electricity consumption, amount due, amount actually received, late payment penalties, late payment penalties actually received, and electricity bill amount from the electricity bill information.
[0155] Extract feature fields such as the content of the follow-up and supervision and the type of follow-up and supervision business from the follow-up and supervision information.
[0156] The user feature group construction module 202 is used to classify urban village power users according to user activity based on the preset three-channel user group division rules, and construct corresponding user feature groups based on the user behavior characteristics of the division results; wherein, the division results include low-activity users, medium-activity users and high-activity users.
[0157] In this embodiment of the application, according to the user group division rules, the user activity of urban village power users is assessed by detecting whether users have engaged in follow-up or supervision behavior and by counting the number of times they make calls to request electricity services. Urban village power users are divided into low-activity users, medium-activity users, and high-activity users.
[0158] Based on the electricity request service tickets of each type of active user, the user feature groups corresponding to low-activity users, medium-activity users, and high-activity users are constructed by integrating the user behavior characteristics from three aspects: ticket duration, electricity cost, and business sensitivity.
[0159] The user characteristic groups corresponding to the medium-activity users and the high-activity users are also associated with reminder and supervision information.
[0160] The feature filtering module 203 is used to filter the features of the user feature group by calculating the feature importance to obtain the optimal text features.
[0161] In this embodiment of the application, the sensitive text information of the user feature group is subjected to feature extraction and vectorization to obtain text features;
[0162] The importance scores of the text features and the statistical features of the user feature group are calculated using an XGBoost tree to obtain the importance scores of each text feature.
[0163] Based on the importance score, the text features and the statistical features are sorted in descending order, and then the top first threshold features are selected as the optimal text features.
[0164] User profile building module 204 is used to predict the sensitivity of urban village electricity users to electricity fees based on a preset three-channel prediction model and build a user profile of urban village electricity users; wherein, the three-channel prediction model is optimized by the optimal text features.
[0165] In this embodiment of the application, the electricity user information of urban villages is input into the three-channel prediction model. The electricity cost sensitivity of low-activity users, medium-activity users and high-activity users is predicted by each channel and the judgment threshold of each channel, so as to obtain the prediction label of each urban village electricity user; wherein, the prediction label is an electricity cost sensitive label or an electricity cost insensitive label.
[0166] Based on the predicted labels, a profile of electricity users in urban villages is constructed.
[0167] In some embodiments, the user feature group construction module 202 specifically comprises:
[0168] Based on preset user group segmentation rules, the system detects whether urban village electricity users engage in follow-up or supervision activities and counts the number of calls made requesting electricity services. This categorizes electricity users into low-activity, medium-activity, and high-activity users. Because these users are divided into three different activity levels, urban village electricity users are thus divided into three channels.
[0169] Subsequently, stratified sampling was used to divide users with different activity levels into training and test sets at a ratio of 9:1, providing data support for the subsequent construction of user feature groups. Furthermore, during the data partitioning process, it was ensured that the ratio of positive to negative samples was consistent between the training and test sets.
[0170] Because the user behavior characteristics of the three types of active users are inconsistent, it is necessary to integrate the user behavior characteristics so that the constructed user feature group can more accurately represent the characteristics of different types of active users.
[0171] Before feature integration, for all users, any work order content containing the field "This work order is a branch center test work order" was directly deleted, along with all corresponding content, to ensure data authenticity. Duplicate data and records with abnormal times were also deleted.
[0172] In addition, for medium- and high-activity users, the number of work order records corresponding to customer codes is counted, and data with more than 10 work order records are deleted to ensure the accuracy of the model built later.
[0173] Then, based on the electricity request service work orders of each type of active user, the user feature groups corresponding to low-activity users, medium-activity users, and high-activity users are constructed by integrating the user behavior characteristics from three aspects: work order duration, electricity cost, and business sensitivity.
[0174] For inactive users, the user feature group focuses on lightweight modeling of users with only one work order record. Therefore, users with only one work order record are selected from the electricity request service work orders, and the following user feature group is constructed based on the user behavior characteristics:
[0175] (1) Basic statistical characteristics: call duration is normalized and service type is one-hot encoded;
[0176] (2) Time characteristics: The work order processing time is divided into month, day and hour, and the work order recording time is marked as early / mid / late ten days of the month;
[0177] (3) Electricity consumption characteristics: classification of the first digit of the electricity consumption type code and mapping of positive examples; historical power outage frequency, Z-Score standardization; time window statistics, calculating the power outage frequency in the past 3 months, 6 months and 12 months, capturing short-term and long-term impacts;
[0178] Because there are many types of electricity use, including residential electricity use (type number 200), rural residential electricity use (type number 201), and urban residential electricity use (type number 202), all of which fall under the broad category of residential electricity use, a new classification is performed using the first digit of the electricity use type number to reduce feature dimensions and redundancy. The proportional mapping refers to calculating the proportion of electricity-sensitive users for each type of electricity use number.
[0179] (1) Characteristics of electricity bill information: The range of the amount is from single digits to millions or even tens of millions. Therefore, logarithmic transformation of the amount receivable / receivable and calculation of the difference after logarithmic transformation are necessary to enhance the interpretability of the data.
[0180] (2) Sensitive text information: Sensitive information (such as account number and mobile phone number) in the electricity request service order is processed by word segmentation and regular expression replacement.
[0181] For moderately active users, deep time series analysis and business sensitivity scoring were introduced. Therefore, users with 2 to 10 service requests were selected from the electricity request service orders, and after removing reminder and follow-up information, a corresponding user characteristic group was constructed:
[0182] (1) Time series analysis: Calculate the minimum / maximum / mean of customer call intervals and the standard deviation of the date distribution;
[0183] (2) Business sensitivity characteristics: positive sample ratio mapping based on the topic of telephone request service work order;
[0184] (3) Analysis of power supply unit coding: classification of coding length and its distribution ratio;
[0185] (4) In-depth characteristics of accounts receivable electricity bill information: difference in late payment penalties and average number of monthly records;
[0186] (5) Sensitive text information: Sensitive information (such as account number and mobile phone number) in the electricity request service order is processed by word segmentation and regular expression replacement.
[0187] For highly active users, the system integrates reminder and supervision information and complex business logic. It filters users with 2 to 10 service requests from the electricity demand service orders, and after associating these with reminder and supervision information, constructs corresponding user characteristic groups.
[0188] (1) Derivative features of expedited and supervised information: Expand the types of supervised business into multiple columns and classify customer names (such as institutions and merchants) to generate unique hot codes;
[0189] (2) In-depth time analysis: median and standard deviation of customer call intervals, maximum number of calls per month;
[0190] (3) Enhanced topic sensitivity: Calculate multi-topic ratio scores based on training set data and fill missing values based on the global median;
[0191] Specifically, for example, a work order might read, "[One household without electricity] A non-residential customer reported one household without electricity. After guiding the customer to inspect, the equipment malfunction and its ownership cannot be determined. Please investigate on-site." In this case, "one household without electricity" in the brackets represents the specific topic. The training set contains all work order records and the total number of different topics, as well as the electricity cost sensitivity ratio for each topic. It's important to note that the same customer may have multiple work order records, potentially corresponding to different topics; the same topic appearing multiple times may correspond to multiple users, and these users may correspond to different electricity cost sensitivity tags. Using the median of the entire dataset for imputation refers to using the median of the proportion scores of all topics in the entire dataset to impute users with missing proportion scores for multiple topics.
[0192] Sensitive text information: Sensitive information (such as account number and mobile phone number) in electricity request service orders and follow-up information is segmented and replaced with regular expressions to obtain the corresponding work order vocabulary, and the number of words and average word length of work order vocabulary are counted.
[0193] In some embodiments, the feature filtering module 203 specifically comprises:
[0194] First, the sensitive text information of the user feature group is extracted and vectorized using the TF-IDF algorithm to obtain text features of more than 10,000 dimensions.
[0195] The specific parameter settings of the TF-IDF algorithm in this embodiment are as follows: both single words and two consecutive words are considered to enrich the feature expression; the term frequency (TF) is used to represent the sensitive text information, but the inverse document frequency (IDF) is not applied; a logarithmic transformation is performed on the term frequency to reduce the influence of high-frequency words; a preset standard vocabulary is forced to be used to standardize the logarithmic transformation result and extract features during feature extraction, and redundant and useless features are removed during the feature extraction process to reduce dimensionality and obtain text features.
[0196] Among them, the TF-IDF algorithm is a commonly used weighting technique for information retrieval and data mining. TF stands for Term Frequency, and IDF stands for Inverse Document Frequency.
[0197] In addition, feature extraction is based on user feature groups corresponding to low-activity users, medium-activity users, and high-activity users.
[0198] After obtaining the text features, xgb.DMatrix is used to convert the text features and statistical features into the XGBoost internal format, specifying the feature column names for subsequent analysis of feature importance. The statistical features are those features within the user feature group excluding the text features.
[0199] Then, the feature importance of the text features and the statistical features of the user feature group are calculated using an XGBoost tree to obtain the importance score of each text feature.
[0200] Specifically, for the parameter settings of the XGBoost tree, a strong L2 regularization is set to prevent overfitting, while the sampling ratio of urban village power user samples and the text features, as well as the depth of the XGBoost tree and the weight of child nodes, are set.
[0201] For XGBoost trees, the number of iterations is set to 6666, but early stopping is enabled. The error rates on both the training and validation sets are monitored to prevent overfitting. Training information is output every 200 rounds to facilitate timely adjustments to training parameters.
[0202] The statistical features of the text features and the user feature group are input into the XGBoost tree-based model to evaluate the feature importance. The importance score of each feature is obtained by setting the parameters mentioned above.
[0203] Then, according to the importance score, the text features and the statistical features are sorted in descending order, and the top 5% of features are selected as the optimal text features.
[0204] In some embodiments, the user profile construction module 204 specifically comprises:
[0205] The electricity user information of urban villages is input into a trained three-channel prediction model. Then, the electricity cost sensitivity of low-activity users, medium-activity users and high-activity users is predicted by the judgment threshold of each channel in the model. The predicted label of each urban village electricity user is output.
[0206] Specifically, the three-channel prediction model is trained based on the XGBoost model. First, a model training set is constructed based on the obtained optimal text features, and the XGBoost model is trained using the model training set. During the model training process, a grid search method is used to traverse each combination of model parameters to determine the optimal combination of hyperparameters.
[0207] In determining the optimal combination of hyperparameters, the number of training epochs is dynamically controlled using early stopping to prevent overfitting. In this embodiment, the maximum number of training epochs is set to 50.
[0208] Then, the three channels of the XGBoost model are tuned according to the optimal hyperparameter combination to obtain the trained three-channel prediction model. These three channels are used to predict the sensitivity of low-activity, medium-activity, and high-activity users to electricity costs, respectively.
[0209] In addition, the XGBoost model is trained independently five times, with the seed parameter (seed=i) randomly set each time to generate diverse base learners. The training data is converted into DMatrix format, and the predicted probability output is optimized using a binary classification objective function (binary:logistic).
[0210] The trained XGBoost model performs five independent predictions for low-activity users, medium-activity users, and high-activity users through its three channels, resulting in five independent predictions.
[0211] The five prediction results were integrated using Bagging, and then the final prediction label was generated using a simple voting method.
[0212] The model prediction process also uses a threshold value corresponding to each channel to classify users, predicting those with values greater than or equal to the threshold as electricity cost-sensitive users. This threshold value is derived from test data trained on a model with the goal of maximizing the F1 score.
[0213] The F1 score is a metric used to comprehensively evaluate the performance of a classification model. It combines precision and recall, and its expression is as follows:
[0214]
[0215] In the formula, precision is the proportion of the number of samples correctly predicted as positive to the total number of samples predicted as positive by the classifier; recall is the proportion of the number of samples correctly predicted as positive to the total number of samples that were actually positive.
[0216] Specifically, the prediction process of the three-channel prediction model is actually to obtain a binary classification result, labeling urban village electricity users with either a price-sensitive label or a price-insensitive label. For a single channel of the three-channel prediction model, its final prediction probability is:
[0217]
[0218] In the formula, p i To predict the probability that a sample belongs to the positive class, this application embodiment determines the probability of predicting a sample as a positive or negative sample by setting a judgment threshold. t (xi) represents the prediction contribution of the t-th decision tree to sample xi, where t is the decision tree index in the XGBoost iteration process.
[0219] Furthermore, embodiments of this application provide some indicator curve changes related to highly active users. The graph shows the changes in the PR curve, where P is the precision rate mentioned above and R is the recall rate mentioned above. Figure 2 The graph shows the PR curve using a single XGBoost algorithm, where the horizontal axis represents recall and the vertical axis represents precision. The solid line is the PR curve and the dashed line is the baseline (representing the horizontal part of the curve). The graph shows that the F1 score reaches its optimal value of 0.97 when both P and R are close to 1.
[0220] Finally, based on the final predicted labels output by the model, the output is the electricity users in urban villages who are sensitive to electricity costs. Based on this, a corresponding profile of electricity users in urban villages is constructed to provide data support for the governance of electricity costs in urban villages.
[0221] Implementing the embodiments of this application has the following beneficial effects:
[0222] This application first collects multiple user behavior features representing user concern about electricity costs from various data sources related to electricity payment and electricity demand services. Then, based on preset user group segmentation rules, urban village electricity users are divided into low-activity, medium-activity, and high-activity users according to their activity levels, and corresponding user feature groups are constructed. These user feature groups can centrally reflect whether users are sensitive or insensitive to electricity costs. The user feature groups are then filtered, selecting features with higher importance to achieve data dimensionality reduction and improve the efficiency of subsequent user prediction. The model is optimized based on the optimal text features to ensure that each channel has a high sensitivity to users with different activity levels, focusing on predicting different types of users. The trained three-channel prediction model is used to make multiple predictions on the sensitivity of urban village electricity users to electricity costs, identifying electricity cost-sensitive users and constructing a more accurate profile of urban village electricity users.
[0223] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. In particular, it should be noted that any modifications, equivalent substitutions, or improvements made by those skilled in the art within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for constructing user profiles of urban village electricity users based on three-channel clustering, characterized in that, include: Collect several user behavior characteristics related to electricity users in urban villages from electricity service data sources; According to the preset user group segmentation rules, urban village electricity users are divided according to user activity, and corresponding user feature groups are constructed based on the user behavior characteristics; wherein, the segmentation results include low-activity users, medium-activity users, and high-activity users; By calculating feature importance, the features of the user feature group are filtered to obtain the optimal text features; Based on a pre-defined three-channel prediction model, the sensitivity of urban village electricity users to electricity fees is predicted, and a profile of urban village electricity users is constructed; wherein, the three-channel prediction model is optimized through the optimal text features.
2. The method for constructing urban village power user profiles based on three-channel clustering according to claim 1, characterized in that, The process involves classifying urban village electricity users according to preset user group segmentation rules based on user activity, and constructing corresponding user feature groups based on the user behavior characteristics. Specifically: According to the user group segmentation rules, the user activity of urban village power users is assessed by detecting whether users have engaged in follow-up or supervision behavior and by counting the number of times they make calls to request electricity services. Urban village power users are divided into low-activity users, medium-activity users, and high-activity users. Based on the electricity request service tickets of each type of active user, the user feature groups corresponding to low-activity users, medium-activity users, and high-activity users are constructed by integrating the user behavior characteristics from three aspects: ticket duration, electricity cost, and business sensitivity. The user characteristic groups corresponding to the medium-activity users and the high-activity users are also associated with reminder and supervision information.
3. The method for constructing urban village power user profiles based on three-channel clustering according to claim 1, characterized in that, The process of calculating feature importance and filtering the features of the user feature group to obtain the optimal text features is as follows: The sensitive text information of the user feature group is subjected to feature extraction and vectorization to obtain text features; The importance scores of the text features and the statistical features of the user feature group are calculated using an XGBoost tree to obtain the importance scores of each text feature. Based on the importance score, the text features and the statistical features are sorted in descending order, and then the top first threshold features are selected as the optimal text features.
4. The method for constructing urban village power user profiles based on three-channel clustering according to claim 3, characterized in that, The process of extracting and vectorizing the sensitive text information of the user feature group to obtain text features is as follows: Simultaneously considering both single words and two consecutive words, a logarithmic transformation is performed on the word frequencies of the sensitive text information; Based on a pre-defined standard vocabulary, the logarithmic transformation results are standardized and features are extracted. Redundant and useless features are removed during the feature extraction process to obtain text features.
5. The method for constructing urban village power user profiles based on three-channel clustering according to claim 3, characterized in that, The sensitive text information specifically includes: The electricity request service order and the follow-up information are processed by word segmentation and regular expression replacement to obtain a number of work order terms; The number of words and average word length of the work order vocabulary are counted, and the sensitive text information is constructed by combining the work order vocabulary.
6. The method for constructing urban village power user profiles based on three-channel clustering according to claim 1, characterized in that, The process involves predicting the sensitivity of urban village electricity users to electricity costs based on a pre-set three-channel prediction model, and constructing a user profile for urban village electricity users. Specifically: The electricity user information of urban villages is input into the three-channel prediction model. The electricity cost sensitivity of low-activity users, medium-activity users and high-activity users is predicted by each channel and the judgment threshold of each channel, so as to obtain the prediction label of each electricity user in urban villages; wherein, the prediction label is an electricity cost sensitive label or an electricity cost insensitive label. Based on the predicted labels, a profile of electricity users in urban villages is constructed.
7. The method for constructing urban village power user profiles based on three-channel clustering according to claim 6, characterized in that, The three-channel prediction model is optimized using the optimal text features, specifically as follows: A model training set is constructed based on the optimal text features, and the preset prediction model is trained using the model training set. During model training, each combination of model parameters is traversed using a grid search method to determine the optimal combination of hyperparameters. The three channels of the prediction model are tuned according to the optimal hyperparameter combination to obtain the trained three-channel prediction model; wherein the channels are used to predict the sensitivity of low-activity users, medium-activity users and high-activity users to electricity costs.
8. The method for constructing urban village power user profiles based on three-channel clustering according to claim 6, characterized in that, The method of predicting the electricity cost sensitivity of low-activity, medium-activity, and high-activity users by using each channel and the judgment threshold for each channel is as follows: By using each channel and the judgment threshold of each channel, a second threshold prediction is performed on low-activity users, medium-activity users, and high-activity users respectively, and the second threshold prediction results are obtained. The prediction results of the second threshold are integrated, and the labels of the integrated results are voted on using a voting method to obtain the prediction label corresponding to each urban village power user.
9. The method for constructing urban village power user profiles based on three-channel clustering according to claim 1, characterized in that, The collection of several user behavior characteristics related to urban village electricity users from electricity service data sources specifically includes: Obtain service order information and electricity consumption information of electricity users in urban villages from electricity service data sources; Extract the user behavior features related to electricity service request information and follow-up information from the service work order information; Extract user behavior features related to electricity billing information and the user's historical electricity usage from the electricity consumption information.
10. A system for constructing user profiles of urban village electricity users based on three-channel clustering, characterized in that, include: The module includes a feature extraction module, a user feature cluster construction module, a feature filtering module, and a user profile construction module. Among them, the feature extraction module is used to collect several user behavior features related to urban village power users from power service data sources; The user feature group construction module is used to classify urban village electricity users according to user activity based on preset user group classification rules, and construct corresponding user feature groups based on the user behavior characteristics of the classification results; wherein, the classification results include low-activity users, medium-activity users, and high-activity users; The feature filtering module is used to filter the features of the user feature group by calculating the feature importance, so as to obtain the optimal text features; The user profile building module is used to predict the sensitivity of urban village electricity users to electricity fees based on a preset three-channel prediction model, and to build a user profile of urban village electricity users; wherein, the three-channel prediction model is optimized by the optimal text features.