An edge cooperative electricity stealing analysis method for electric energy meter

By constructing a denoising autoencoder network within a collaborative framework between the electricity meter and terminal computing devices, and extracting features using kurtosis coefficient and sliding window methods, the real-time performance and communication bandwidth issues of electricity theft detection are resolved. This enables efficient identification and feedback of electricity theft behavior, thereby improving the intelligence level of the electricity meter.

CN121051393BActive Publication Date: 2026-02-24LIYANG HUAPENG ELECTRIC POWER METER
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511563250.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-02-24
Estimated Expiration
2045-10-30

AI Technical Summary

Technical Problem

Existing methods for detecting electricity theft rely on a central cloud platform, which puts a strain on communication bandwidth and results in delayed detection response, failing to fully utilize the value of smart meters as edge nodes.

Method used

A denoising autoencoder network based on contrastive learning is constructed. Data preprocessing, feature extraction, and training are performed collaboratively by the electricity meter and terminal computing device to achieve accurate detection of electricity theft. The kurtosis features of the data are extracted using the kurtosis coefficient and sliding window method. A joint loss function is constructed for training by combining principal component analysis and oversampling technology.

Benefits of technology

It significantly reduces the communication bandwidth requirements of the central cloud platform, improves the real-time performance and response speed of electricity theft detection, enhances the accuracy and robustness of detection, and enables electricity meters to have sensing, analysis, feedback, and execution capabilities, thereby improving the intelligence level of the anti-electricity theft system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121051393B_ABST
    Figure CN121051393B_ABST
Patent Text Reader

Abstract

The application discloses an edge cooperative electricity stealing analysis method for an electric energy meter, and the method comprises a model construction method, and the steps of the model construction method comprise the following steps: S05: constructing a denoising autoencoder network based on contrast learning, training the denoising autoencoder network by using a user electricity consumption feature dataset, and obtaining a final denoising autoencoder network; the step S05 comprises the following steps: for any input sample of the user electricity consumption feature data, two different random noises are added respectively to obtain two enhanced samples; a denoising autoencoder is taken as a backbone to construct a feature extraction model to extract latent space low-dimensional features and a reconstruction sample corresponding to the two enhanced samples; a reconstruction loss is constructed based on the reconstruction sample, and a contrast loss is constructed based on the latent space low-dimensional features. Through the method, a model for recognizing electricity stealing behaviors can be obtained well, and accurate capture of user electricity stealing features is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an edge-coordinated electricity theft analysis method for electricity meters, belonging to the field of electricity meter and power grid technology. Background Technology

[0002] Currently, electricity plays a vital role in economic development and people's lives. Driven by profit, the phenomenon of electricity theft through various technical means is also on the rise. Electricity theft, a widespread problem plaguing power grids worldwide, not only damages the economic interests of power grid systems but also poses serious safety hazards, threatening users' electricity safety.

[0003] Traditional electricity theft detection methods primarily rely on manual verification of suspicious users. However, this method consumes significant human and time resources and is gradually being phased out in the context of informatization and intelligentization. Currently, Advanced Metering Infrastructure (AMI) based on the Energy Internet is being established, and with the large-scale deployment of smart meters, large amounts of user electricity consumption data can be quickly and easily acquired. Against this backdrop, data-driven machine learning methods are increasingly being introduced into electricity theft detection. However, most existing data-driven methods rely on uploading all raw data to a central cloud platform for centralized processing. This not only puts pressure on communication bandwidth but may also lead to delays in detection response and fails to fully leverage the potential value of smart meters as edge nodes.

[0004] Autoencoders, as a typical unsupervised learning model, can compress high-dimensional electricity consumption data into a low-dimensional latent space through nonlinear mapping, thereby extracting key features from the data and preserving the original information to the greatest extent during reconstruction. In recent years, the development of deep learning has enabled improved models such as multi-layer stacked autoencoders and variational autoencoders to exhibit higher performance in anomaly detection and pattern recognition, providing a new technical path for electricity theft detection. Combining the massive electricity consumption data collected by the AMI system with advanced methods such as autoencoders, contrastive learning, and graph neural networks, the generalization ability and robustness of the model can be improved while ensuring detection accuracy, thus providing strong support for the safe and stable operation of the smart grid. However, how to deploy these advanced detection capabilities more efficiently and in real-time to the edge side closer to users and achieve intelligent feedback from electricity meters remains a pressing issue. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the defects of the prior art and provide an edge-coordinated electricity theft analysis method for electricity meters. This method can effectively obtain a model for identifying electricity theft behavior. Using this model, the difference between the electricity consumption characteristics of normal users and the characteristics of electricity thieves can be maximized, thereby achieving accurate capture of the electricity theft characteristics of users.

[0006] To solve the above-mentioned technical problems, the technical solution of the present invention is: an edge-coordinated electricity theft analysis method for electricity meters, the method including a model building method, the steps of which include:

[0007] S01: Obtain historical user electricity consumption data collected by the electricity meter; the historical user electricity consumption data includes electricity consumption data of normal users and electricity consumption data of users who steal electricity.

[0008] S02: Perform data preprocessing on user electricity consumption data;

[0009] S03: Calculate the kurtosis coefficient of the preprocessed data using the sliding window method to obtain the kurtosis component of the user's electricity consumption data;

[0010] S04: Perform feature processing on the kurtosis component of user electricity consumption data to obtain user electricity consumption feature dataset;

[0011] S05: Construct a denoising autoencoder network based on contrastive learning, and train the denoising autoencoder network using the user electricity consumption feature dataset to obtain the final denoising autoencoder network.

[0012] Step S05 includes the following steps:

[0013] S051: For any input sample of user electricity consumption feature data, add two different random noises to obtain two enhanced samples;

[0014] S052: Using a denoising autoencoder as the backbone, a feature extraction model is constructed to extract the latent space low-dimensional features and reconstructed samples corresponding to the two enhanced samples;

[0015] S053: Construct a reconstruction loss based on reconstructed samples, construct a contrastive loss based on low-dimensional features of the latent space, and construct a joint loss based on the reconstruction loss and contrastive loss as the network loss;

[0016] S054: Update the parameters of the feature extraction model using the adaptive Adam optimizer based on network loss;

[0017] S055: Using low-dimensional features of the latent space extracted by the feature extraction model as input, train a logistic regression classifier to identify whether a user is a normal user or a user who steals electricity and output the electricity theft identification result.

[0018] S056: After training, the final denoising autoencoder network is obtained; the denoising autoencoder network includes a feature extraction model based on the denoising autoencoder and a logistic regression classifier.

[0019] Furthermore, the method also includes methods for detecting electricity theft, which include:

[0020] S11: Obtain the user's electricity consumption data of the user to be tested through the electricity meter;

[0021] S12: Process the user's electricity consumption data through steps S02, S03 and S04 to obtain the user's electricity consumption feature dataset of the user to be detected.

[0022] S13: Input the user electricity consumption feature dataset of the user to be detected into the trained denoising autoencoder network to output the electricity theft identification result.

[0023] Furthermore, methods for detecting electricity theft also include:

[0024] S14: Feedback the electricity theft identification result to the electricity meter, and the electricity meter performs corresponding local feedback and actions based on the feedback result.

[0025] Furthermore, the corresponding local feedback and actions include at least one of the following: local alarm, data logging, dynamic adjustment of sampling rate, execution of remote control commands, self-test, and locking.

[0026] Furthermore, in step S02, data preprocessing includes: correcting missing and / or outlier values ​​in user electricity consumption data, and normalizing all user electricity consumption data.

[0027] Furthermore, the electricity consumption data of any user at any sampling point is represented as follows: ,in This indicates the number of user samples collected. Indicates the sampling time point;

[0028] Missing values ​​are handled using the following formula: ;

[0029] in, Calculate functions for missing values. Indicates user Electricity consumption data at all sampling times. Represents missing values. This is the mean calculation function;

[0030] For outliers, use 3. The criteria are used for processing, and the specific judgment formula is as follows: ;

[0031] in, This is a function for calculating outliers. express Standard deviation;

[0032] The normalization process is performed using the following formula: ;

[0033] in, For normalized calculation functions, Indicates user The minimum value of electricity consumption data at all sampling times. Indicates user The maximum value of electricity consumption data at all sampling times.

[0034] Furthermore, step S03 specifically includes:

[0035] The kurtosis coefficient is defined as ;

[0036] in, This is the function for calculating the kurtosis coefficient. Represents the mathematical expectation. This represents the mean of the data. Indicates standard deviation;

[0037] Set the sliding window size to Step size is For preprocessed power consumption time series data of any user Subsequences are obtained by dividing the window: ;

[0038] Calculate the kurtosis coefficient for each subsequence. The length of the kurtosis values ​​corresponding to each window is New feature sequences This refers to the kurtosis component of the user's electricity consumption data.

[0039] Furthermore, step S4 specifically includes:

[0040] S041: Let the kurtosis component of user electricity consumption data be... Principal component analysis (PCA) is used to reduce its dimensionality while retaining key feature information: PCA yields the eigenvalue sequence. Take the front The eigenvectors corresponding to the eigenvalues ​​constitute the feature matrix. Then the new features after dimensionality reduction are: ;in, The value is determined by the contribution rate. To determine, adopt :

[0041] ;

[0042] S042: Next, the minority class is oversampled using the Borderline-SMOTE method to make the minority class sample size nearly equal to the majority class sample size. Specifically, this includes:

[0043] Minority class samples for identifying boundaries: For each minority class sample Calculate its Find the nearest neighbors and count the number of neighbors of the majority class: ;

[0044] in, Indicates sample of The number of majority class samples in the nearest neighbors; if A relatively high value, ranging from 0.5 to 1, indicates that the minority class sample is located in the class boundary region;

[0045] Synthesize new minority class samples: minority class samples at the boundary Randomly select one of its minority class neighbors. Generate new minority class samples between the two: These are the interpolation coefficients;

[0046] Repeat the steps for each boundary minority class sample: synthesize new minority class samples until the number of minority class samples is comparable to the number of majority class samples. In the scenario of electricity theft detection, electricity theft users are generally considered to belong to the minority class samples.

[0047] S043: After PCA dimensionality reduction in step S41 and Borderline-SMOTE processing in step S42, the user electricity consumption feature dataset is obtained.

[0048] Further, in step S051, a sample of user electricity consumption characteristic data is input. Two augmented samples were obtained. ;

[0049] In step S052, two enhanced samples The encoder outputs low-dimensional features in the latent space after passing through a denoising autoencoder. and Low-dimensional features of the latent space and The decoder outputs the reconstructed sample after passing through the denoising autoencoder. , ;

[0050] In step S053, the specific method for constructing the reconstruction loss is as follows:

[0051] The reconstruction loss is: ; ;

[0052] in, Represents the mean squared error loss. For the sample size, and These are the corresponding sample values ​​of the reconstructed data and the original data, respectively;

[0053] In step S053, the specific method for constructing the contrast loss is as follows:

[0054] Low-dimensional features of the latent space and Merged into a new low-dimensional feature set ; Include The first feature, the second The labels for each feature are The corresponding features are Perform L2 normalization on it: ;

[0055] Similarity calculation: ;

[0056] in, and Indicates the sample index. This is a temperature coefficient used to adjust the concentration of the similarity distribution; it is usually taken as a positive value, and in subsequent calculations, it is set to... Set the distance between the same samples to 0;

[0057] Define all valid sets Collections of the same kind heterogeneous collection ;

[0058] Total loss of similar comparison items : ;

[0059] in, It is a numerically stable term;

[0060] Total loss of unbearable sample items Calculation:

[0061] First, the hardest-to-bear samples are selected according to the Top-K criterion. For each sample... ,exist According to Sort by similarity from highest to lowest, and select the top... indivual:

[0062] ;

[0063] Obtain the difficult sample set ;

[0064] Next, calculate the total loss for the hard-to-bear sample terms. : ;

[0065] Used to limit the similarity to a positive value;

[0066] By combining the total loss of similar comparison items with the total loss of hard-to-bear sample items, we can obtain the total comparison loss: ;

[0067] in, To control the balance between clustering similar samples and separating dissimilar samples, the loss weight coefficients are used.

[0068] In step S053, the specific method for constructing the joint loss based on the reconstruction loss and the contrast loss is as follows:

[0069] Joint losses: ;

[0070] in, The category weight coefficients are set as follows: Normal users Electricity theft users This is the ratio of the number of samples in the majority class to the number of samples in the minority class.

[0071] Furthermore, in step S055, let the low-dimensional features of the latent space extracted by the feature extraction model be... The sample label is The discriminant function of the logistic regression classifier is: ;

[0072] Among them, the function Represents the given input features When, the probability that the sample belongs to category 1; For the sigmod function, These are model parameters;

[0073] The logistic regression classifier is trained by minimizing the negative log-likelihood, with the objective function being: .

[0074] By adopting the above technical solution, the present invention has the following beneficial effects:

[0075] 1. The method provided by this invention uses kurtosis coefficient as the core indicator and utilizes sliding window extraction of data kurtosis features to initially amplify the characteristics of electricity theft.

[0076] 2. This invention constructs a denoising autoencoder network based on contrastive learning, combining the advantages of both to effectively achieve multiple tasks such as denoising, feature extraction, feature separation, and feature classification under the same model.

[0077] 3. This invention constructs a joint loss to achieve unified training of reconstruction loss and contrast loss, and achieves the allocation of training focus for feature extraction task and feature separation task through parameter adjustment.

[0078] 4. This invention can adopt a collaborative framework of electricity meter and terminal computing device (edge ​​terminal). The electricity meter collects user electricity consumption data, and the terminal computing device can be responsible for preprocessing the data, feature processing, training a denoising autoencoder network, and identifying electricity theft behavior. By deploying the complex electricity theft detection calculation on the edge terminal, the communication bandwidth requirements of the central cloud platform are significantly reduced, and the real-time performance and response speed of the detection are improved.

[0079] 5. By feeding back the detection results to the electricity meter and giving the electricity meter the ability to issue local alarms and execute control commands, the electricity meter is transformed from a simple data collector into an intelligent terminal with "sensing-analysis-feedback-execution" capabilities, thereby improving the intelligence level and robustness of the entire anti-theft electricity system.

[0080] 6. Compared with existing multi-model, multi-feature fusion methods for electricity theft detection, the method of this invention has a simple and clear network structure, occupies less memory and computing resources, has superior efficiency, and is more conducive to deployment in local processing units or electricity information collection terminals. Attached Figure Description

[0081] Figure 1 This is a flowchart of the method of the present invention;

[0082] Figure 2 A flowchart illustrating the construction of the joint loss in the method of this invention;

[0083] Figure 3 Comparison charts of various statistical features of different user categories;

[0084] Figure 4 This is a graph showing the variation of loss and accuracy of the method of the present invention with the number of rounds.

[0085] Figure 5(a) is a comparison of the T-SNE results using the CNN method in the embodiment of the present invention;

[0086] Figure 5(b) is a comparison of the T-SNE effects using the AE method in the embodiments of the present invention;

[0087] Figure 5(c) is a comparison of the T-SNE effects of the proposed Kurt-SCLAE method in the embodiments of the present invention. Detailed Implementation

[0088] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings.

[0089] like Figure 1 , 2 As shown, an edge-coordinated electricity theft analysis method for electricity meters is proposed. The method includes a model building method, and the steps of the model building method include:

[0090] S01: Obtain historical user electricity consumption data collected by the electricity meter; the historical user electricity consumption data includes electricity consumption data of normal users and electricity consumption data of users who steal electricity.

[0091] S02: Perform data preprocessing on user electricity consumption data;

[0092] S03: Calculate the kurtosis coefficient of the preprocessed data using the sliding window method to obtain the kurtosis component of the user's electricity consumption data;

[0093] S04: Perform feature processing on the kurtosis component of user electricity consumption data to obtain user electricity consumption feature dataset;

[0094] S05: Construct a denoising autoencoder network based on contrastive learning, and train the denoising autoencoder network using the user electricity consumption feature dataset to obtain the final denoising autoencoder network.

[0095] Step S05 includes the following steps:

[0096] S051: For any input sample of user electricity consumption feature data, add two different random noises to obtain two enhanced samples;

[0097] S052: Using a denoising autoencoder as the backbone, a feature extraction model is constructed to extract the latent space low-dimensional features and reconstructed samples corresponding to the two enhanced samples;

[0098] S053: Construct a reconstruction loss based on reconstructed samples, construct a contrastive loss based on low-dimensional features of the latent space, and construct a joint loss based on the reconstruction loss and contrastive loss as the network loss;

[0099] S054: Update the parameters of the feature extraction model using the adaptive Adam optimizer based on network loss;

[0100] S055: Using low-dimensional features of the latent space extracted by the feature extraction model as input, train a logistic regression classifier to identify whether a user is a normal user or a user who steals electricity and output the electricity theft identification result.

[0101] S056: After training, the final denoising autoencoder network is obtained; the denoising autoencoder network includes a feature extraction model based on the denoising autoencoder and a logistic regression classifier.

[0102] Steps S02 to S05 can be performed on a terminal computing device.

[0103] Specifically, the method also includes electricity theft detection methods, which include:

[0104] S11: Obtain the user's electricity consumption data of the user to be tested through the electricity meter;

[0105] S12: Process the user's electricity consumption data through steps S02, S03 and S04 to obtain the user's electricity consumption feature dataset of the user to be detected.

[0106] S13: Input the user electricity consumption feature dataset of the user to be detected into the trained denoising autoencoder network to output the electricity theft identification result.

[0107] Specifically, methods for detecting electricity theft also include:

[0108] S14: Feedback the electricity theft identification result to the electricity meter, and the electricity meter performs corresponding local feedback and actions based on the feedback result.

[0109] Specifically, the terminal computing device uses the trained network for online electricity theft detection. Upon receiving electricity consumption data from the user to be detected transmitted by the electricity meter, the terminal computing device processes this data using the same steps S02 to S04, and imports it into the trained network to obtain the electricity theft behavior identification result (such as theft probability or classification label). The terminal computing device then feeds this identification result and corresponding decision instructions (such as alarm level and suggested actions) back to the corresponding electricity meter via the local communication network. After receiving the feedback, the smart meter executes one or more of the following local feedback and actions according to a preset strategy:

[0110] 1) Local alarm: The electricity theft alarm is triggered locally on the electricity meter via LED indicator lights, display screen or buzzer;

[0111] 2) Data logging: Record the alarm event in detail, including time, alarm level, relevant data snapshots, etc.;

[0112] 3) Dynamic adjustment of sampling rate: Automatically increases the frequency of subsequent power consumption data collection to obtain more detailed abnormal behavior data;

[0113] 4) Remote control command execution: Under strict authorization and security verification, execute remote power-off, power-limiting, and other control commands issued by the terminal computing device;

[0114] 5) Self-test and locking: Trigger the internal self-test program of the electricity meter, or automatically lock certain functions when physical tampering is detected.

[0115] Specifically, in this embodiment, the electricity meter can collect users' electricity consumption data in real time and periodically, and can simultaneously collect auxiliary data such as voltage, current, and power factor. The collected data is used to construct time-series electricity consumption data according to the sampling time, and is securely transmitted to the terminal computing device deployed on the edge via a local communication network (such as power line carrier, wireless communication, etc.); the electricity consumption data of any user at any sampling point is represented as follows: ,in This indicates the number of user samples collected. Indicates the sampling time point;

[0116] Specifically, in step S02, data preprocessing includes: correcting missing and outlier values ​​in user electricity consumption data, and normalizing all user electricity consumption data.

[0117] Specifically, missing values ​​are handled using the following formula: ;

[0118] in, Calculate functions for missing values. Indicates user Electricity consumption data at all sampling times. Represents missing values. This is the mean calculation function;

[0119] For outliers, use 3. The criteria are used for processing, and the specific judgment formula is as follows: ;

[0120] in, This is a function for calculating outliers. express Standard deviation;

[0121] The normalization process is performed using the following formula: ;

[0122] in, For normalized calculation functions, Indicates user The minimum value of electricity consumption data at all sampling times. Indicates user The maximum value of electricity consumption data at all sampling times.

[0123] Specifically, step S03 is as follows: In order to characterize the distribution pattern of time-series electricity consumption data, a kurtosis index is introduced. Kurtosis reflects the steepness of the data distribution, and the kurtosis coefficient is defined as follows:

[0124] ;

[0125] in, This is the function for calculating the kurtosis coefficient. Represents the mathematical expectation. This represents the mean of the data. The kurtosis value represents the standard deviation. A large kurtosis value indicates that the data distribution has a peak and that extreme values ​​are more likely to occur. A small kurtosis value indicates that the distribution is relatively flat.

[0126] Set the sliding window size to Step size is For preprocessed power consumption time series data of any user Subsequences are obtained by dividing the window: ;

[0127] Calculate the kurtosis coefficient for each subsequence. The length of the kurtosis values ​​corresponding to each window is New feature sequences This refers to the kurtosis component of the user's electricity consumption data.

[0128] Specifically, step S4 is as follows:

[0129] S041: Let the kurtosis component of user electricity consumption data be... Principal component analysis (PCA) was used to reduce its dimensionality while preserving key feature information: the eigenvalue sequence was obtained through PCA. Take the front The eigenvectors corresponding to the eigenvalues ​​constitute the feature matrix. Then the new features after dimensionality reduction are: ;in, The value is determined by the contribution rate. To determine, adopt :

[0130] ;

[0131] S042: Next, the minority class is oversampled using the Borderline-SMOTE method to make the minority class sample size nearly equal to the majority class sample size. Specifically, this includes:

[0132] Minority class samples for identifying boundaries: For each minority class sample Calculate its Find the nearest neighbors and count the number of neighbors of the majority class: ;

[0133] in, Indicates sample of The number of majority class samples in the nearest neighbors; if A relatively high value, ranging from 0.5 to 1, indicates that the minority class sample is located in the class boundary region;

[0134] Synthesize new minority class samples: minority class samples at the boundary Randomly select one of its minority class neighbors. Generate new minority class samples between the two: These are the interpolation coefficients;

[0135] Repeat the steps for each boundary minority class sample: synthesize new minority class samples until the number of minority class samples is comparable to the number of majority class samples. In the scenario of electricity theft detection, electricity theft users are generally considered to belong to the minority class samples.

[0136] S043: After PCA dimensionality reduction in step S41 and Borderline-SMOTE processing in step S42, the user electricity consumption feature dataset is obtained. .

[0137] Specifically, such as Figure 2 As shown, in step S051, a sample of user electricity consumption characteristic data is input. Two augmented samples were obtained. , ;

[0138] In step S052, the specific structure of the feature extractor built using the denoising autoencoder as the backbone is as follows:

[0139] Encoder structure (4-layer fully connected network):

[0140]

[0141] Decoder architecture (4-layer fully connected network):

[0142]

[0143] In step S052, two enhanced samples The encoder outputs low-dimensional features in the latent space after passing through a denoising autoencoder. and Low-dimensional features of the latent space and The decoder outputs the reconstructed sample after passing through the denoising autoencoder. , ;

[0144] In step S053, the specific method for constructing the reconstruction loss is as follows:

[0145] The reconstruction loss is: ; ;

[0146] in, Represents the mean squared error loss. For the sample size, and These are the corresponding sample values ​​of the reconstructed data and the original data, respectively;

[0147] In step S053, the specific method for constructing the contrast loss is as follows:

[0148] Low-dimensional features of the latent space and Merged into a new low-dimensional feature set ; Include The first feature, the second The labels for each feature are The corresponding features are Perform L2 normalization on it: ;

[0149] Similarity calculation: ;

[0150] in, and Indicates the sample index. This is a temperature coefficient used to adjust the concentration of the similarity distribution; it is usually taken as a positive value, and in subsequent calculations, it is set to... Set the distance between the same samples to 0;

[0151] Define all valid sets Collections of the same kind heterogeneous collection ;

[0152] Total loss of similar comparison items : ;

[0153] in, It is a numerically stable term;

[0154] Total loss of unbearable sample items Calculation:

[0155] First, the hardest-to-bear samples are selected according to the Top-K criterion. For each sample... ,exist According to Sort by similarity from highest to lowest, and select the top... indivual:

[0156] ;

[0157] Obtain the difficult sample set ;

[0158] Then we obtain the total loss for the hard-to-bear sample terms. : ;

[0159] Used to limit the similarity to a positive value;

[0160] By combining the total loss of similar comparison items with the total loss of hard-to-bear sample items, we can obtain the total comparison loss: ;

[0161] in, To control the balance between clustering similar samples and separating dissimilar samples, the loss weight coefficients are used.

[0162] In step S053, the specific method for constructing the joint loss based on the reconstruction loss and the contrast loss is as follows:

[0163] Joint losses: ;

[0164] in, The category weight coefficients are set as follows: Normal users Electricity theft users This is the ratio of the number of samples in the majority class to the number of samples in the minority class.

[0165] Specifically, in step S055, let the low-dimensional features of the latent space extracted by the feature extraction model be... The sample label is The discriminant function of the logistic regression classifier is: ;

[0166] Among them, the function Represents the given input features When, the probability that the sample belongs to category 1; For the sigmod function, These are model parameters;

[0167] The logistic regression classifier is trained by minimizing the negative log-likelihood, with the objective function being: .

[0168] Experimental verification was conducted using a user electricity consumption dataset collected by the State Grid AMI system. This dataset covers the electricity consumption data of 42,372 users from January 1, 2014 to October 30, 2016. The dataset is categorized based on whether users engaged in electricity theft, and its specific composition is shown in the table below:

[0169] Table 1. Composition of the State Grid Electricity Theft Data Set

[0170]

[0171] The data was randomly divided into a training set and a test set at a ratio of 8:2, and then input into the denoising autoencoder network proposed in this invention under the environment of a simulated terminal computing device.

[0172] To evaluate the effectiveness of different methods in detecting electricity theft, this invention compares the performance of different methods using accuracy (ACC), area under the curve (AUC), and harmonic mean (F1-Score). The calculation methods are as follows:

[0173] ;

[0174] Where TP represents the number of correctly predicted positive classes, TN represents the number of correctly predicted negative classes, FP represents the number of negative classes incorrectly predicted as positive classes, and FN represents the number of positive classes incorrectly predicted as negative classes; Accuracy (ACC) measures the overall proportion of correct predictions made by the denoising autoencoder network. The Area Under the ROC Curve (AUC) measures the classifier's ability to distinguish between positive and negative samples, ranging from 0 to 1, with values ​​closer to 1 indicating better discrimination. Precision, Recall, and F1-Score measure the robustness of the denoising autoencoder network in positive class detection, and are particularly suitable for imbalanced class tasks.

[0175] In terms of parameter settings, the number of training epochs for all methods was set to 100, the sliding window size was set to 30, and the weight coefficients were set accordingly. The default value is 0.5.

[0176] To demonstrate the effectiveness of this invention, in this embodiment, the effects of an autoencoder network (AE), a deep convolutional neural network (DCNN), and the kurtosis-based contrastive learning denoising autoencoder network (Kurt-SCLAE) proposed in this invention are compared.

[0177] exist Figure 3 The data shows the differences in various statistical characteristics of two types of users. It is easy to see that the kurtosis coefficients of the two are quite different, making them suitable as the core indicator for feature extraction and classification.

[0178] Table 2 presents a comparison of the electricity theft detection performance of each method, compared to commonly used classic feature extraction methods such as DCNN and AE.

[0179] By adopting the classification model, the method proposed in this invention has achieved significant improvements in multiple indicators such as ACC, AUC, and F1 score.

[0180] Table 2 Comparison of the performance of various electricity theft detection methods

[0181]

[0182] like Figure 4 As shown, the loss of this method continuously decreases with the increase of training rounds, gradually decreasing from a relatively high initial value and then stabilizing. This indicates that the network model continuously learns during training, its fit to the training data improves, and the error decreases. The accuracy rises rapidly from a low initial level and then stabilizes at a very high level with minimal fluctuations. This demonstrates that the network model not only performs well on the training data but also maintains high accuracy on unseen test data, exhibiting strong generalization ability.

[0183] Figures 5(a), 5(b), and 5(c) present a comparison of T-SNE methods for different approaches. These figures visually demonstrate the feature extraction performance of different methods at different stages. The comparison shows that the network model extracted by this method is superior to other methods in terms of feature extraction capability. The distribution of the two types of samples in the two-dimensional space after passing through the network model is significantly different. At the same time, it is not difficult to find from the kurtosis feature map on the left side of Figure 5(c) that the kurtosis feature distribution has a higher degree of discrimination compared to the original data. This indicates the effectiveness of this method in using sliding window extraction of kurtosis features to achieve preliminary feature extraction.

[0184] The specific embodiments described above further illustrate the technical problems, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An edge collaborative electricity stealing analysis method for electric energy meter, characterized in that, The method includes a model building method, and the steps in the model building method include: S01: Obtain historical user electricity consumption data collected by the electricity meter; the historical user electricity consumption data includes electricity consumption data of normal users and electricity consumption data of users who steal electricity. S02: Perform data preprocessing on user electricity consumption data; S03: Calculate the kurtosis coefficient of the preprocessed data using the sliding window method to obtain the kurtosis component of the user's electricity consumption data; S04: Perform feature processing on the kurtosis component of user electricity consumption data to obtain user electricity consumption feature dataset; S05: Construct a denoising autoencoder network based on contrastive learning, train the denoising autoencoder network using user electricity consumption feature dataset, and obtain the final denoising autoencoder network. Step S05 includes the following steps: S051: For any input sample of user electricity consumption feature data, add two different random noises to obtain two enhanced samples; S052: Using a denoising autoencoder as the backbone, a feature extraction model is constructed to extract the latent space low-dimensional features and reconstructed samples corresponding to the two enhanced samples; S053: Construct a reconstruction loss based on reconstructed samples, construct a contrastive loss based on low-dimensional features of the latent space, and construct a joint loss based on the reconstruction loss and contrastive loss as the network loss; S054: Update the parameters of the feature extraction model using the adaptive Adam optimizer based on network loss; S055: Using low-dimensional features of the latent space extracted by the feature extraction model as input, train a logistic regression classifier to identify whether a user is a normal user or a user who steals electricity and output the electricity theft identification result. S056: After training, the final denoising autoencoder network is obtained; the denoising autoencoder network includes a feature extraction model based on the denoising autoencoder and a logistic regression classifier.

2. The method of claim 1, wherein, It also includes methods for detecting electricity theft, which include: S11: Obtain the user's electricity consumption data of the user to be tested through the electricity meter; S12: Process the user's electricity consumption data through steps S02, S03 and S04 to obtain the user's electricity consumption feature dataset of the user to be detected. S13: Input the user electricity consumption feature dataset of the user to be detected into the trained denoising autoencoder network to output the electricity theft identification result.

3. The method of claim 2, wherein, Electricity theft detection methods also include: S14: Feedback the electricity theft identification result to the electricity meter, and the electricity meter performs corresponding local feedback and actions based on the feedback result.

4. The method of claim 3, wherein, The corresponding local feedback and actions include at least one of the following: local alarm, data logging, dynamic adjustment of sampling rate, execution of remote control commands, self-test and locking.

5. The method of claim 1, wherein, In step S02, data preprocessing includes: correcting missing and / or outlier values ​​in user electricity consumption data, and normalizing all user electricity consumption data.

6. The method according to claim 5, characterized in that, The power consumption data of any user at any sampling point is represented as wherein represents the index of the collected user sample, represents the sampling time point; For missing values, the following formula is used for processing: ; wherein, a mean value calculation function, representing a user all sampling time power consumption data, representing a missing value, a mean value calculation function; For outliers, the 3 criteria are used, and the specific determination formula is as follows: ; wherein is the function of outliers, denotes the standard deviation; Normalization is performed using the following equation: ; wherein, is a normalization calculation function, represents a user is a minimum value of the power consumption data at all sampling instants, represents a user is a maximum value of the power consumption data at all sampling instants.

7. The method according to claim 1, characterized in that, Step S03 is specifically: define the kurtosis coefficient as ; wherein is a kurtosis coefficient calculation function, denotes the mathematical expectation, denotes the data mean value, denotes the standard deviation; Set the sliding window size as , the step size as , and the preprocessed time series data of any user as , the subsequence obtained according to the window division is: ; Calculate the kurtosis coefficient of each sub-sequence The kurtosis coefficients corresponding to each window form a new feature sequence with a length of , which is the kurtosis component of the user's power consumption data.​ 8. The method of claim 7, wherein, Step S4 is as follows: S041: Set the kurtosis component of user electricity consumption data as , and reduce the dimension by using principal component analysis to retain the main characteristic information: the principal component analysis obtains a characteristic value sequence , and the characteristic vectors corresponding to the first characteristic values constitute a characteristic matrix , so the new characteristics after dimension reduction are: ; wherein, The value of is determined by the contribution rate : ; S042: Next, the minority class is oversampled using the Borderline-SMOTE method to make the minority class sample size nearly equal to the majority class sample size. Specifically, this includes: Boundary identifying minority class samples: for each minority class sample , representing a set of minority class samples, compute its nearest neighbors, count the number of majority class neighbors among them: ; wherein, the number of majority class samples among the nearest neighbors of the sample; and a higher value, being between 0.5 and 1, indicates that the minority class sample is located in a class boundary region.​ Synthesizing new minority class samples: for the boundary minority class samples randomly select one of its minority class neighbor samples generate a new minority class sample between the two: is the interpolation coefficient; Repeat the steps for each boundary minority class sample: synthesize new minority class samples until the number of minority class samples is comparable to the number of majority class samples. In the scenario of electricity theft detection, electricity theft users are generally considered to belong to the minority class samples. S043: After PCA dimensionality reduction in step S41 and Borderline-SMOTE processing in step S42, the user electricity consumption feature dataset is obtained.

9. The method according to claim 1, characterized in that, In step S051, a sample of the user power consumption feature data is input , two enhanced samples are obtained ; In step S052, two enhanced samples latent space low-dimensional features output after the encoder of the denoising autoencoder and latent space low-dimensional features output after the encoder of the denoising autoencoder and reconstructed samples output after the decoder of the denoising autoencoder ; In step S053, the specific method for constructing the reconstruction loss is as follows: The reconstruction loss is: where, represents the mean square error loss, is the number of samples, and are the corresponding sample values of the reconstructed data and the original data, respectively. In step S053, the specific method for constructing the contrast loss is as follows: latent space low-dimensional features are merged into a new low-dimensional feature set ; ; contains features, the label of the th feature is , and the corresponding feature is , which is L2-normalized: ; Similarity calculation: ; in, and Indicates the sample index. This is a temperature coefficient used to adjust the concentration of the similarity distribution; it is usually taken as a positive value, and in subsequent calculations, it is set to... Set the distance between the same samples to 0; Define all valid sets Collections of the same kind heterogeneous collection ; Total loss of similar comparison items : ; in, It is a numerically stable term; Total loss of unbearable sample items Calculation: First, the hardest-to-bear samples are selected according to the Top-K criterion. For each sample... ,exist According to Sort by similarity from highest to lowest, and select the top... indivual: ; Obtain the difficult sample set ; Next, calculate the total loss for the hard-to-bear sample terms. : ; Used to limit the similarity to a positive value; By combining the total loss of similar comparison items with the total loss of hard-to-bear sample items, we can obtain the total comparison loss: ; in, To control the balance between clustering similar samples and separating dissimilar samples, the loss weight coefficients are used. In step S053, the specific method for constructing the joint loss based on the reconstruction loss and the contrast loss is as follows: Joint losses: ; in, For category weight coefficients, Set the user category labels as follows: Normal users Electricity theft users , This is the ratio of the number of samples in the majority class to the number of samples in the minority class. Represents the number of samples in the majority class. "Represents the number of samples in the minority class".

10. The method according to claim 1, characterized in that, In step S055, let the low-dimensional features of the latent space extracted by the feature extraction model be... The sample label is ; The discriminant function of the logistic regression classifier is: ; Among them, the function Represents the given input features When, the probability that the sample belongs to category 1; For the sigmod function, These are model parameters; The logistic regression classifier is trained by minimizing the negative log-likelihood, with the objective function being: .

Citation Information

Patent Citations

  • Unbalanced electricity stealing data classification method and device based on SMOTE-GBDT, computer equipment and storage medium

    CN115936926A

  • CNN-Bagging-based electricity larceny detection method and system

    CN117034115A