User electricity charge abnormity identification method based on weighted residual depth forest model

By constructing a three-dimensional dataset and a weighted residual deep forest model, and dynamically generating feature weights, the problems of high false positive rate and poor adaptability in electricity bill anomaly detection are solved, and efficient electricity bill anomaly identification is achieved.

CN120910709APending Publication Date: 2025-11-07STATE GRID ZHEJIANG ELECTRIC POWER CO LTD NINGHAI COUNTY POWER SUPPLY CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510792186.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing methods for detecting abnormal electricity charges have a high false alarm rate, lack self-learning capabilities, have low detection efficiency, cannot adapt to seasonal load fluctuations and changes in user electricity consumption patterns, and lack adaptability to scenarios involving the integration of new energy equipment.

Method used

A three-dimensional dataset containing time series, user attributes, and environmental features is constructed. A feature importance matrix is ​​generated by training a weighted residual deep forest model, and a feature weight vector is dynamically generated. The threshold is dynamically adjusted by combining historical prediction deviation data to achieve multi-dimensional data fusion and intelligent adaptation of electricity consumption patterns.

Benefits of technology

It significantly improves the accuracy and timeliness of electricity bill anomaly detection, reduces the false positive rate, and is suitable for electricity bill anomaly detection scenarios for different types of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910709A_ABST
    Figure CN120910709A_ABST
Patent Text Reader

Abstract

The invention discloses a user electricity charge anomaly identification method based on a weighted residual depth forest model. The method comprises the steps of 1, constructing a three-dimensional data set including time sequence features, user attribute features and environment correlation features; 2, dividing the users into a plurality of feature groups with similar power consumption modes by adopting a clustering algorithm; step 3, training and generating a feature importance matrix of the feature group through a weighted residual depth forest model, and dynamically generating a feature weight vector through an attention mechanism based on the matrix; and step 4, for new user data, judging whether the user electricity charge is abnormal or not. According to the method, intelligent adaptation of multi-dimensional data fusion and power utilization modes is realized, the accuracy and timeliness of anomaly detection are remarkably improved, the problems of high misjudgment rate, poor adaptability and the like of a traditional method are effectively solved, and the method is suitable for electric charge anomaly detection scenes of multiple types of users.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of machine learning, in particular to a user electricity fee anomaly identification method based on a weighted residual deep forest model. BACKGROUND

[0002] Electricity fee anomaly detection is a key technology to guarantee the efficient operation of the power grid and the safety of user electricity consumption. Existing methods mainly rely on manually set static rules and single-dimensional data analysis, which have significant defects. First, static threshold rules cannot adapt to dynamic electricity consumption patterns such as seasonal load fluctuations and production cycle adjustments. The misjudgment rate is high when the residential electricity heating load suddenly increases in winter. The traversal screening of hundreds of rules leads to low efficiency of grassroots review. Second, single-dimensional data analysis only focuses on time series data of electricity consumption, ignoring the correlation between user attributes and external environmental factors. Reasonable load growth caused by industrial user capacity expansion is often misjudged. Finally, traditional models are not trained for user electricity consumption patterns. The "one-size-fits-all" detection results in a high false negative rate in complex scenarios. Moreover, the models lack self-learning ability for new scenarios such as the connection of new energy equipment, and the threshold adjustment is lagging behind. SUMMARY

[0003] The present application aims to overcome the defects of high misjudgment rate, lack of self-learning ability, and low detection efficiency in the prior art, and provides a user electricity fee anomaly identification method based on a weighted residual deep forest model.

[0004] The present application is achieved by the following technical solutions: The user electricity fee anomaly identification method based on the weighted residual deep forest model comprises the following steps: Step 1: Collect multi-dimensional data including user historical electricity consumption, user portrait, and external influencing factors, and construct a three-dimensional data set including time series features, user attribute features, and environmental correlation features. Step 2: Based on the three-dimensional data set, use a clustering algorithm to divide users into several feature groups with similar electricity consumption patterns. Step 3: For each feature group, generate a feature importance matrix for the feature group through a weighted residual deep forest model training, and dynamically generate a feature weight vector based on the feature importance matrix through an attention mechanism to adapt to the electricity consumption pattern characteristics of the feature group. Step 4: For new user data, first determine the corresponding feature group through a clustering algorithm, and then use the corresponding weighted residual deep forest model of the feature group to output a predicted value. Calculate the deviation rate of the actual value and the predicted value. At the same time, calculate a dynamic adjustment threshold based on the historical prediction deviation data of the feature group. If the deviation rate is greater than or equal to the dynamic adjustment threshold, it is determined to be an electricity fee anomaly, otherwise it is determined to be normal.

[0005] As preferred, in the step 1, the time series features include hourly or daily load curves and periodic indicators, the user attribute features include power consumption types, load characteristics and historical anomaly records, and the environment-related features include weather-sensitive coefficients and electricity price impact factors.

[0006] As preferred, in the step 2, the clustering algorithm is a combination of dynamic time warping and improved K-means algorithm, specifically: calculating the dynamic time warping distance of the user's hourly power consumption sequence to measure the time sequence similarity of the load curve; using the improved K-means algorithm to initialize the clustering center, and integrating the time sequence similarity and user attribute difference through the weighted distance measurement formula.

[0007] As preferred, the feature importance matrix of the feature group is generated by training the weighted residual deep forest model, including: scanning the data of the feature group using different time windows to generate multi-scale time sequence features; constructing a deep forest containing multiple residual blocks, each residual block learns the residual between the predicted value and the actual value through identity mapping; through Gini index attenuation or gradient flow analysis of the decision tree node, the contribution of each feature to anomaly detection is calculated to generate a feature importance matrix.

[0008] As preferred, the different time windows include 1 hour, 1 day or 1 week.

[0009] As preferred, the feature weight vector is dynamically generated through the attention mechanism, specifically: normalizing the feature importance matrix to obtain a feature weight vector; introducing an attention mechanism in the residual block to perform channel weighting on the input features according to the feature weight vector.

[0010] As preferred, in the step 4, the calculation method of the dynamically adjusted threshold is: collecting historical prediction bias data of the feature group, calculating the mean and standard deviation of the bias data, and generating a dynamically adjusted threshold according to the mean and standard deviation.

[0011] As preferred, the user electricity fee anomaly identification method based on the weighted residual deep forest model also optimizes the weighted residual deep forest model, specifically: when the false positive rate of a certain feature group continuously exceeds the set threshold within a set time, the weighted residual deep forest model of the group is incrementally trained; by dynamically updating the feature importance matrix, new key features are identified, and the feature weight vector is automatically adjusted.

[0012] The beneficial effects of the present application are: the present application constructs a three-dimensional data set containing time series, user attributes and environment correlation characteristics, divides the feature groups of similar power consumption modes by using a clustering algorithm, trains a dedicated model for each group and dynamically generates an adaptive feature weight vector, and calculates a dynamically adjusted threshold combined with historical prediction bias data. The scheme breaks through the limitations of static rules, realizes multi-dimensional data fusion and intelligent adaptation of power consumption modes, significantly improves the accuracy and timeliness of anomaly detection, effectively solves the problems of high misjudgment rate and poor adaptability of traditional methods, and is suitable for electricity fee anomaly detection scenarios of multiple types of users. BRIEF DESCRIPTION OF DRAWINGS

[0013] Figure 1 is a flowchart of the present application. DETAILED DESCRIPTION

[0014] Example implementations will now be described more fully with reference to the accompanying drawings. Example implementations may, however, be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these implementations are provided so that this disclosure will be thorough and complete, and will fully convey the inventive aspects of the example implementations to those skilled in the art.

[0015] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of the application. One skilled in the relevant art will recognize, however, that the application can be practiced without one or more of the specific details, or with other methods, components, devices, steps, etc. In other instances, well-known structures, devices, implementations, or operations are not shown or described in detail in order to avoid obscuring aspects of the application.

[0016] The flowchart shown in the accompanying drawings is only an exemplary illustration, and does not necessarily include all contents and operations / steps, nor does it necessarily have to be executed in the order described. For example, some operations / steps can be further broken down, and some operations / steps can be combined or partially combined, so the actual execution order can be changed according to the actual situation.

[0017] Embodiment: A user electricity fee anomaly identification method based on a weighted residual depth forest model, as shown in Figure 1 includes the following steps: Step 1, collect multi-dimensional data containing user historical power consumption, user portrait and external influencing factors, and construct a three-dimensional data set containing time series features, user attribute features and environment correlation features; Step 2, based on the three-dimensional data set, the user is divided into several feature groups with similar power consumption modes by using a clustering algorithm; Step 3, for each feature group, generate a feature importance matrix of the feature group by training a weighted residual deep forest model, and dynamically generate a feature weight vector based on the feature importance matrix by an attention mechanism to adapt to the power consumption mode features of the feature group; Step 4, for new user data, first determine the corresponding feature group by a clustering algorithm, and then calculate the deviation rate of the actual value and the predicted value by using the predicted value output by the weighted residual deep forest model corresponding to the feature group; at the same time, calculate a dynamic adjustment threshold based on the historical prediction deviation data of the feature group, if the deviation rate is greater than or equal to the dynamic adjustment threshold, it is determined as an electricity fee anomaly, otherwise it is determined as normal.

[0018] In the step 1, the time series features include hourly or daily load curves and periodic indicators, the user attribute features include power consumption types, load characteristics and historical anomaly records, and the environment-related features include meteorological sensitivity coefficients and electricity price influence factors.

[0019] In the step 2, the clustering algorithm is a combination of dynamic time warping and improved K-means algorithm, specifically: calculating the dynamic time warping distance of the user's time power consumption sequence to measure the time sequence similarity of the load curve; adopting an improved K-means algorithm to initialize the clustering center, and integrating the time sequence similarity and user attribute difference through a weighted distance measurement formula: wherein, a is the time sequence feature weight, w k is the importance weight of the user attribute feature, which is calculated by Gini coefficient, and DTW is the dynamic time warping representation.

[0020] The feature importance matrix of the feature group is generated by training a weighted residual deep forest model, including: scanning the data of the feature group with different time windows to generate multi-scale time sequence features; constructing a deep forest containing multiple residual blocks, each residual block learning the residual between the predicted value and the actual value through an identity mapping; calculating the contribution of each feature to anomaly detection through Gini index attenuation or gradient flow analysis of decision tree nodes, and generating an n x m feature importance matrix, n is the number of features, and m is the cluster group number.

[0021] The different time windows include 1 hour, 1 day or 1 week.

[0022] The feature weight vector is dynamically generated by an attention mechanism, specifically: normalize the feature importance matrix to obtain a feature weight vector wk = [mk,1, mk,2, …, mk,n], wherein mk,i represents the importance score of feature i in group k; The attention mechanism is introduced in the residual block, and the input features are channel-weighted according to the feature weight vector: x weighted =x⊙w k , wherein ⊙ is an element-wise multiplication, and the influence of key features is strengthened.

[0023] In step 4, the calculation method of the dynamically adjusted threshold value is: The historical prediction deviation data of the feature group is collected, the mean and standard deviation of the deviation data are calculated, and the dynamically adjusted threshold value is generated according to the mean and standard deviation.

[0024] The user electricity fee anomaly identification method based on the weighted residual depth forest model also optimizes the weighted residual depth forest model, specifically: When the false alarm rate of a certain feature group continuously exceeds the set threshold value within a set time, the incremental training of the weighted residual depth forest model of the group is performed; By dynamically updating the feature importance matrix, new key features are identified, and the feature weight vector is automatically adjusted.

[0025] In this embodiment, the user electricity fee anomaly identification method based on the weighted residual depth forest model is specifically used as follows: through the clustering algorithm in step 2, four core groups are finally formed: Class A (stable load type): government agencies, hospitals, load fluctuation <5% / day, low weather influence; Class B (seasonal fluctuation type): residential users, summer / winter load surge is obvious, weather sensitive coefficient >0.7; Class C (production cycle type): manufacturing industry, load presents strict periodicity with production shifts (three shifts / single shift); D class (abnormal risk type): historical electricity stealing users, high failure rate users in the area, load curve has high frequency burr.

[0026] Example 1: Resident user winter heating load surge scenario: A resident user's daily electricity consumption increased from 8kWh in autumn to 15kWh in winter due to the use of electric heating equipment in December, lasting for one week. The historical static threshold value is set as "baseline electricity consumption ±30% fluctuation rate".

[0027] The static baseline uses the data of the last 30 days, without distinguishing seasonal differences. The December data includes low load in autumn and high load in winter, resulting in a low baseline value, and without correlating with weather data. When the temperature is ≤0℃, the electric heating load belongs to normal fluctuation, but the traditional rule directly triggers the anomaly (15kWh>13kWh), with high false alarm rate.

[0028] Through the method of this embodiment, the user is classified as type B (seasonal fluctuation type), historical data shows that the user's electricity consumption is strongly negatively correlated with air temperature, and the air temperature factor is automatically weighted when calculating the benchmark electricity consumption. Dynamic window switching is also performed, and the benchmark is automatically switched to "winter 90-day window" calculation on February 1 to avoid interference from autumn data.

[0029] It should be noted that the state of the electric meter, payment records, and the possibility of equipment failure or electricity theft also need to be detected synchronously to exclude normal fluctuations.

[0030] Example two, production cycle adjustment scenario of industrial and commercial users: A manufacturing user adjusts from single-shift to double-shift due to an increase in orders starting in July, and the daily electricity consumption increases from 200 kWh to 350 kWh, lasting for two weeks. The traditional rule sets "continuous 3-day electricity consumption increase > 50% triggers an anomaly."

[0031] The static rule does not identify the production cycle change, only judges by the absolute increase, ignores the business logic of "shift increase → reasonable load increase", and does not associate with user attributes (the enterprise belongs to "production cycle type" C), and the misjudgment rate of historical similar users due to shift adjustment is high.

[0032] Through the method of this embodiment, the user is classified as type C (production cycle type), the historical load curve is strictly aligned with the production shift (DTW distance < 0.5), and the "production plan calendar" and "equipment account" are input. The model identifies it as "reasonable fluctuation due to capacity expansion", generates a new benchmark of 340 kWh based on the historical data of double-shift, and the fluctuation rate threshold is ± 15%.

[0033] Example three: equipment failure scenario of abnormal risk type user: The current data of a user in a transformer area abnormally fluctuates due to the voltage sampling module failure of the smart electric meter. The traditional rule only relies on a single dimension of electricity consumption and misjudges it as electricity theft.

[0034] In the traditional method, single-dimensional detection cannot distinguish between data collection failure and real anomaly, and does not associate with device state data, lacking anomaly cause identification capability.

[0035] Through the method of this embodiment, the user is classified as type D (abnormal risk type), and the model focuses on device state features. After the dynamic threshold is triggered, the secondary verification retrieves device data: the voltage standard deviation reaches 5V, and the waveform distortion rate is 18%, which is determined as "meter failure" rather than electricity theft, generating a device failure work order rather than an electricity theft warning, avoiding misjudgment by the grassroots.

[0036] In summary, the method of the embodiment divides users into independent groups according to power consumption modes, avoids static threshold pollution of stable load type users on dynamic model of seasonal fluctuation type users, forces embedding of features strongly related to the mode in each group model, blocks misjudgment of other dimension features when single dimension is abnormal, allows threshold to change gradually in a reasonable range, avoids black and white judgment of traditional rules, and reduces misjudgment rate of load gradual change scene.

[0037] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the embodiments disclosed herein. It is intended that the specification and examples be considered as exemplary only, with a true scope of the application being indicated by the following claims.

[0038] It is to be understood that the application is not limited to the precise details of construction and the exact arrangements of the components described above and illustrated in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the application being limited only by the claims appended hereto.

Claims

1. A user electricity bill anomaly identification method based on a weighted residual depth forest model, characterized in that, The method comprises the following steps: Step 1, collecting multi-dimensional data including user historical electricity consumption, user portrait and external influencing factors, and constructing a three-dimensional data set including time series features, user attribute features and environment-related features; Step 2, based on the three-dimensional data set, using a clustering algorithm to divide users into several feature groups with similar electricity consumption patterns; Step 3, for each feature group, generating a feature importance matrix of the feature group through a weighted residual deep forest model training, and dynamically generating a feature weight vector based on the feature importance matrix through an attention mechanism to adapt to the electricity consumption pattern characteristics of the feature group; Step 4, for new user data, first determine the corresponding feature group through a clustering algorithm, and then use the corresponding weighted residual deep forest model of the feature group to output a predicted value, and calculate the deviation rate of the actual value and the predicted value; at the same time, a dynamic adjustment threshold is calculated based on the historical prediction deviation data of the feature group, if the deviation rate is greater than or equal to the dynamic adjustment threshold, it is determined as an electricity fee anomaly, otherwise it is determined as normal. 2.The user electricity bill anomaly identification method based on the weighted residual depth forest model according to claim 1, characterized in that, In step 1, the time series features include hourly or daily load curves and periodic indicators, the user attribute features include electricity consumption types, load characteristics and historical anomaly records, and the environment-related features include weather-sensitive coefficients and electricity price influencing factors. 3.The user electricity bill anomaly identification method based on the weighted residual depth forest model according to claim 1, characterized in that, In step 2, the clustering algorithm is a combination of dynamic time warping and improved K-means algorithm, specifically: Calculate the dynamic time warping distance of the user's time electricity consumption sequence to measure the time sequence similarity of the load curve; Use the improved K-means algorithm to initialize the clustering center, and integrate the time sequence similarity and user attribute difference through a weighted distance measurement formula. 4.The user electricity bill anomaly identification method based on the weighted residual depth forest model according to claim 2, characterized in that, The feature importance matrix of the feature group is generated by the weighted residual deep forest model training, which comprises: Scanning the data of the feature group with different time windows to generate multi-scale time sequence features; Construct a deep forest containing multiple residual blocks, and learn the residual between the predicted value and the actual value through an identity mapping in each residual block; Through Gini index attenuation or gradient flow analysis of the decision tree node, the contribution of each feature to anomaly detection is calculated to generate a feature importance matrix.

5. The user electricity bill anomaly identification method based on the weighted residual depth forest model according to claim 4, characterized in that, The different time windows include 1 hour, 1 day or 1 week.

6. The user electricity bill anomaly identification method based on the weighted residual depth forest model according to claim 4 or 5, characterized in that, The feature weight vector is dynamically generated through the attention mechanism, specifically: Normalize the feature importance matrix to obtain a feature weight vector; Introduce an attention mechanism in the residual block to weight the input features according to the feature weight vector. 7.The user electricity bill anomaly identification method based on the weighted residual depth forest model according to claim 1, characterized in that, In step 4, the dynamic adjustment threshold is calculated by: Collecting historical prediction deviation data of the feature group, calculating the mean and standard deviation of the deviation data, and generating a dynamic adjustment threshold according to the mean and standard deviation. 8.The user electricity bill anomaly identification method based on the weighted residual depth forest model according to claim 1, characterized in that, The weighted residual deep forest model is also optimized, specifically: When the false positive rate of a certain feature group continuously exceeds a set threshold within a set time, the weighted residual deep forest model of the group is incrementally trained; Through dynamic updating of the feature importance matrix, new key features are identified, and the feature weight vector is automatically adjusted.