A power stealing detection method based on self-attention encoding and GSA optimization classification

The electricity theft detection method based on self-attention encoding and GSA optimized classification solves the problems of strong dependence on manual rules, limited static feature modeling ability, and insufficient time-series dependency modeling in traditional electricity theft detection, and achieves intelligent and robust electricity theft behavior recognition.

CN120781176BActive Publication Date: 2025-11-11STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511284867.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2025-11-11
Estimated Expiration
2045-09-10

AI Technical Summary

Technical Problem

Existing electricity theft detection methods rely on manual rules, have limited static feature modeling capabilities, and lack time-series dependency modeling, resulting in high false alarm rates and difficulty in identifying gradual electricity theft patterns.

Method used

The method employs self-attention encoding and GSA optimized classification. It constructs a time-series matrix by acquiring electricity data, calculates the coefficient of variation to screen features, extracts feature vectors using the self-attention mechanism, and performs weighted fusion through a GSA optimized classifier to identify electricity theft.

Benefits of technology

Without requiring additional hardware, intelligent and robust electricity theft detection can be achieved by processing only a few days of frozen data. It effectively captures nonlinear positional relationships in long-distance dependent and discontinuous sequences, improving the accuracy and efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120781176B_ABST
    Figure CN120781176B_ABST
Patent Text Reader

Abstract

This application provides a method for detecting electricity theft based on self-attention coding and GSA optimized classification, belonging to the field of anomaly detection technology. It solves the technical problems of strong reliance on manual rules, limited static feature modeling capabilities, and insufficient time-series dependency modeling in existing technologies. This method extracts and classifies features from the target screening matrix using an encoding model to obtain feature vectors. The feature vectors are then compared with benchmark vectors in a benchmark template library to calculate the feature difference. After transformation, the feature difference is weighted and fused with the feature vectors to obtain a weighted feature vector. This weighted feature vector is then input into an optimal classifier to obtain the anomaly result. In the process of electricity theft detection, this application achieves high-precision electricity theft detection without hardware dependence through self-attention coding and GSA optimized classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of anomaly detection technology, and in particular to a method for detecting electricity theft based on self-attention coding and GSA optimized classification. Background Technology

[0002] Electricity theft by users in the power system causes huge non-technical losses, resulting in significant direct economic losses for enterprises every year. Such behavior not only undermines the fairness of electricity metering but also triggers a chain of technical risks, including abnormal transformer overload, regional voltage instability, and even the burnout of distribution network equipment, becoming a core threat to the safe operation of the power grid.

[0003] Current traditional detection methods have significant limitations: Threshold analysis methods set deviation thresholds based on historical electricity consumption to trigger abnormal alarms. This method relies on manual experience rules, fails to identify gradual electricity theft patterns, and has an excessively high false alarm rate; Statistical model methods analyze the statistical characteristics of electricity consumption curves using traditional machine learning methods such as support vector machines or random forests. However, the input data for traditional machine learning is one-dimensional data, which can only extract static features and is difficult to capture long-term electricity consumption time-series correlations; Basic deep learning methods use convolutional neural networks or recurrent neural networks to learn electricity consumption sequences, but they are difficult to model cross-time period remote dependencies or cannot adaptively focus on highly suspicious periods. Summary of the Invention

[0004] This application provides a method for detecting electricity theft based on self-attention encoding and GSA optimized classification, which solves the technical problems of strong dependence on manual rules, limited static feature modeling ability, and insufficient time-series dependency modeling in the prior art.

[0005] To achieve the above objectives, this application adopts the following technical solution:

[0006] Firstly, a method for detecting electricity theft based on self-attention encoding and GSA-optimized classification is provided, including:

[0007] S1. Obtain the power consumption data of the user to be tested over several days; construct the time series matrix to be tested based on the power consumption data; calculate the coefficient of variation of the data in the i-th row of the time series matrix to be tested and filter the data in the row to obtain the screening matrix to be tested;

[0008] S2. The feature vector is obtained by extracting and classifying the selection matrix to be tested through an encoding model; wherein the encoding model is constructed and trained based on a self-attention mechanism.

[0009] S3. Calculate the difference between the feature vector and the benchmark vector in the benchmark template library to obtain the feature difference.

[0010] S4. After transforming the feature difference degree, it is weighted and fused with the feature vector to obtain the weighted feature vector;

[0011] S5. Input the weighted feature vector into the optimal classifier to obtain the abnormal results; wherein, the optimal classifier is obtained by optimizing the classifier using GSA.

[0012] In conjunction with the first aspect above, in one possible implementation, the time series matrix to be tested is:

[0013] ;

[0014] in, The time series matrix of the user to be tested. Based on the base date, For window length, The sliding step size, This represents the number of windows.

[0015] In conjunction with the first aspect above, in one possible implementation, the method for obtaining the coefficient of variation includes:

[0016] Through calculation formula Calculate the coefficient of variation of the i-th row of data in the time series matrix of the user under test. ;in, Let be the coefficient of variation of the electricity data in the i-th row of the time series matrix of the user under test. and Let be the standard deviation and mean of the electricity data in the i-th row of the time series matrix of the user to be tested;

[0017] The coefficients of variation are iterated through, and the rows corresponding to the highest and lowest coefficients of variation are retained to obtain the screening matrix to be tested.

[0018] In conjunction with the first aspect above, in one possible implementation, the encoding model is constructed in the following ways:

[0019] Obtain historical battery power data for several training user types, and construct a time series matrix for each user based on the time of the battery power data;

[0020] Calculate the coefficient of variation of the i-th row of data in several time series matrices and filter the row data to obtain a set of filtered matrices; where, , This represents the row number of the time series matrix;

[0021] The preliminary encoding model extracts features from the sub-matrices in the selected matrix set to obtain training feature vectors; the preliminary encoding model is built based on a self-attention mechanism.

[0022] The initial coding model parameters are optimized using gradient based on the loss function to obtain the coding model.

[0023] In conjunction with the first aspect above, in one possible implementation, constructing a time-series matrix for each user based on the time of the electricity data includes:

[0024] A calendar feature matrix is ​​constructed based on the date of acquiring the user's battery data; a two-dimensional time series matrix is ​​constructed based on the user's battery data; the two-dimensional time series matrix is ​​merged with the calendar feature matrix to obtain the time series matrix.

[0025] In conjunction with the first aspect above, in one possible implementation, constructing a calendar feature matrix based on the date of acquiring the training user's battery data includes:

[0026] Based on the battery data of trained users over several days Through calculation formula Calculate the absolute date object corresponding to the training user; where t is the time index. For a single day's time span;

[0027] Feature extraction is performed on individual objects within the absolute date object to obtain the single-day feature vector of the training user. ;in, , As a characteristic of the week, Characteristics of holidays This is a characteristic of a special period.

[0028] In conjunction with the first aspect above, in one possible implementation, the preliminary coding model includes: an input layer, several Transformer encoder layers, several Dropout layers, and an output terminal;

[0029] The Transformer encoder includes: learnable positional encoding. Multi-head self-attention sublayer and feedforward neural network sublayer; among which, For position-encoded output, For trainable embedding weights, For learnable location embedding.

[0030] In conjunction with the first aspect above, in one possible implementation, the feedforward neural network sublayer employs Gaussian error linear units. As an activation function: where, This represents the output value of a sublayer in a feedforward neural network. This is the cumulative distribution function of the standard normal distribution.

[0031] In conjunction with the first aspect above, in one possible implementation, the method for obtaining the feature vector includes:

[0032] During the inference phase of the encoding model, Monte Carlo sampling is achieved by forcibly activating several Dropout layers:

[0033] A single input sample undergoes K independent forward propagation iterations. During each propagation, a single Dropout layer randomly masks the output values ​​of some neurons according to a set probability, ultimately outputting several K preliminary feature vectors. The single input sample is selected from the test screening matrix.

[0034] Calculate the average of several K preliminary eigenvectors to obtain the eigenvectors.

[0035] In conjunction with the first aspect above, in one possible implementation, the method for obtaining the feature difference degree includes:

[0036] Calculate the mean of the training feature vectors for each user to obtain the benchmark template library. ;in, Represents the vector of resident means. Represents the industrial mean vector. Represents the vector of business mean;

[0037] The feature vectors are compared with corresponding benchmark vectors of the same type in the benchmark template library using Euclidean distance. Perform a difference metric calculation to obtain the feature difference degree; where .

[0038] In conjunction with the first aspect above, in one possible implementation, the method for obtaining the weighted feature vector includes:

[0039] After applying a linear transformation to the feature difference, activation is performed using an activation function to obtain a dynamic gating weight vector;

[0040] The weighted feature vector is obtained by performing element-wise multiplication between the dynamic gating weight vector and the feature vector.

[0041] In conjunction with the first aspect above, in one possible implementation, the optimal classifier is obtained through GSA optimization, including: using the maximum AUC of the classifier as the fitness objective according to GSA, performing iterative optimization of the classifier hyperparameters to obtain the optimal classifier.

[0042] Based on the above technical solutions, the electricity theft detection method based on self-attention encoding and GSA optimized classification provided in this application requires no additional hardware devices. It only requires processing several days of frozen data and inputting it into the model to detect and identify user electricity theft behavior. This method, by introducing an attention mechanism, can effectively capture long-distance dependencies and nonlinear positional relationships in discontinuous sequences. Simultaneously, it employs the Gravity Search Algorithm (GSA) to optimize the classifier. Its biomimetic optimization mechanism, parameter adaptability, and efficient global search capability provide a more intelligent and robust solution for the classification process. This method solves the technical problems of strong reliance on manual rules, limited static feature modeling capabilities, and insufficient temporal dependency modeling in traditional electricity theft detection.

[0043] In a second aspect, an electronic device is provided, comprising: a communication unit and a processing unit; the communication unit is used to acquire power consumption data of a user under test over several days;

[0044] The processing unit is used to calculate the coefficient of variation of the i-th row of data in the time series matrix to be tested and to filter the row data to obtain the screening matrix to be tested; to extract features and classify the screening matrix to be tested using an encoding model to obtain a feature vector; wherein the encoding model is constructed and trained based on a self-attention mechanism; to calculate the difference between the feature vector and the benchmark vector in the benchmark template library to obtain the feature difference; to transform the feature difference and then perform weighted fusion with the feature vector to obtain a weighted feature vector; to input the weighted feature vector into the optimal classifier to obtain the anomaly result; wherein the optimal classifier is obtained by optimizing the classifier using GSA.

[0045] Thirdly, this application provides an electronic device, including: a processor and a storage medium; the storage medium includes instructions, and the processor is configured to execute the instructions to implement the methods described in the first aspect and any possible implementation thereof. This electronic device may be an electronic device or a chip within an electronic device.

[0046] Fourthly, this application provides a method for detecting electricity theft based on self-attention coding and GSA optimized classification, comprising: a data acquisition module, a data processing and feature extraction module, and a prediction result module; wherein, the data acquisition module is used to acquire electricity consumption data of the user to be tested over several days; the data processing and feature extraction module is used to calculate the coefficient of variation of the i-th row of data in the time series matrix to be tested and to filter the row data to obtain a screening matrix to be tested; the screening matrix to be tested is subjected to feature extraction and classification through an encoding model to obtain a feature vector; wherein, the encoding model is constructed and trained based on a self-attention mechanism; the prediction result module is used to calculate the difference between the feature vector and the benchmark vector in the benchmark template library to obtain a feature difference degree; after transforming the feature difference degree, it is weighted and fused with the feature vector to obtain a weighted feature vector; the weighted feature vector is input into the optimal classifier to obtain anomaly results; wherein, the optimal classifier is obtained through a GSA optimized classifier.

[0047] Fifthly, this application provides a computer-readable storage medium storing instructions that, when executed on an electronic device, cause the electronic device to perform the methods described in the first aspect and any possible implementation thereof.

[0048] Sixthly, this application provides a computer program product containing instructions that, when run on an electronic device, cause the electronic device to perform the methods described in the first aspect and any possible implementation thereof.

[0049] This application provides a method for detecting electricity theft based on self-attention encoding and Gravity Search Algorithm (GSA) optimized classification. This method requires no additional hardware; it only needs to process several days of frozen data and input it into the model to detect and identify user electricity theft. By introducing an attention mechanism, this method effectively captures long-distance dependencies and nonlinear positional relationships in discontinuous sequences. Simultaneously, it employs the Gravity Search Algorithm (GSA) to optimize the classifier. Its biomimetic optimization mechanism, parameter adaptability, and efficient global search capability provide a more intelligent and robust solution for the classification process. This method solves the technical problems of traditional electricity theft detection, such as strong reliance on manual rules, limited static feature modeling capabilities, and insufficient temporal dependency modeling.

[0050] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description

[0051] Figure 1 A system architecture diagram of an electricity theft detection method provided in this application embodiment;

[0052] Figure 2 A flowchart illustrating a method for detecting electricity theft based on self-attention coding and GSA optimized classification, provided in an embodiment of this application;

[0053] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0054] Figure 4 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0055] In the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.

[0056] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0057] The electricity theft detection method based on self-attention coding and GSA optimized classification provided in this application embodiment can be applied to, for example... Figure 1 In the electricity theft detection method 100 shown, such as Figure 1 As shown, the communication system includes: an information capture terminal 101, a cloud computing device 102, and an edge computing node 103.

[0058] Among them, the information capture terminal 101 is used to acquire the power consumption data of the user under test over a number of days;

[0059] The cloud computing device 102 is used to calculate the coefficient of variation of the i-th row of data in the time series matrix to be tested and to filter the row data to obtain the screening matrix to be tested; the screening matrix to be tested is subjected to feature extraction and classification through an encoding model to obtain feature vectors; wherein, the encoding model is constructed and trained based on a self-attention mechanism;

[0060] Edge computing node 103 is used to calculate the difference between the feature vector and the benchmark vector in the benchmark template library to obtain the feature difference; after transforming the feature difference, it is weighted and fused with the feature vector to obtain a weighted feature vector; the weighted feature vector is input into the optimal classifier to obtain the anomaly result; wherein, the optimal classifier is obtained by optimizing the classifier through GSA.

[0061] To address the technical problems of strong reliance on manual rules, limited static feature modeling capabilities, and insufficient time-series dependency modeling in existing technologies, this application provides a method for detecting electricity theft based on self-attention coding and GSA-optimized classification. The method includes: acquiring electricity consumption data of a user under test over several days; calculating the coefficient of variation of the i-th row of data in the time-series matrix under test and filtering the row data to obtain a filter matrix; extracting and classifying features from the filter matrix under test using a coding model to obtain feature vectors; wherein the coding model is constructed and trained based on a self-attention mechanism; calculating the difference between the feature vectors and benchmark vectors in a benchmark template library to obtain feature difference; transforming the feature difference and then weighting and fusing it with the feature vectors to obtain a weighted feature vector; inputting the weighted feature vector into an optimal classifier to obtain anomaly results; wherein the optimal classifier is obtained through a GSA-optimized classifier. Based on this, the method solves the technical problems of strong reliance on manual rules, limited static feature modeling capabilities, and insufficient time-series dependency modeling in traditional electricity theft detection. This method can detect and identify electricity theft by processing and inputting several days of frozen data into the model without requiring additional hardware. By introducing an attention mechanism, it effectively captures nonlinear positional relationships in long-distance dependencies and discontinuous sequences. Simultaneously, it employs the Gravity Search Algorithm (GSA) to optimize the classifier. Its biomimetic optimization mechanism, parameter adaptability, and efficient global search capability provide a more intelligent and robust solution for the classification process.

[0062] like Figure 2 As shown in the figure, an embodiment of this application provides a method for detecting electricity theft based on self-attention coding and GSA optimized classification, including:

[0063] S201. Obtain the power consumption data of the user to be tested over several days; construct the time series matrix to be tested based on the power consumption data; calculate the coefficient of variation of the data in the i-th row of the time series matrix to be tested and filter the data in the row to obtain the screening matrix to be tested.

[0064] The battery consumption data of the user under test over a certain number of days is data with user type tags.

[0065] It should be noted that the user type tag includes three categories: residential, industrial, and commercial.

[0066] In one possible implementation of the embodiments of this application, combined with Figure 2 The above S201 can be implemented through the following S201-1 and S201-2, which are explained in detail below:

[0067] S201-1, The time series matrix to be tested is:

[0068] ;

[0069] in, The time series matrix of the user to be tested. Based on the base date, For window length, The sliding step size, This represents the number of windows.

[0070] It should be noted that different users It may differ, depending on the data start date.

[0071] S201-2, The method for obtaining the coefficient of variation includes:

[0072] Through calculation formula Calculate the coefficient of variation of the i-th row of data in the time series matrix of the user under test. .

[0073] The coefficients of variation are iterated through, and the rows corresponding to the highest and lowest coefficients of variation are retained to obtain the screening matrix to be tested.

[0074] in, Let be the coefficient of variation of the electricity data in the i-th row of the time series matrix of the user under test. and Let be the standard deviation and mean of the electricity data in the i-th row of the time series matrix of the user to be tested.

[0075] S202. The feature vector is obtained by extracting and classifying the features of the screening matrix to be tested through the coding model.

[0076] The encoding model is constructed and trained based on a self-attention mechanism.

[0077] In some implementations, the method for obtaining the feature vector includes:

[0078] During the inference phase of the encoding model, Monte Carlo sampling is achieved by forcibly activating several Dropout layers:

[0079] A single input sample undergoes K independent forward propagation iterations. During each propagation, a single Dropout layer randomly masks the output values ​​of some neurons according to a set probability, ultimately outputting several K preliminary feature vectors. The single input sample is selected from the test screening matrix.

[0080] Calculate the average of several K preliminary eigenvectors to obtain the eigenvectors.

[0081] In one possible implementation of the embodiments of this application, combined with Figure 2 The above S202 can be implemented in detail through the following S202-1, which is explained in detail below:

[0082] S202-1, The construction method of the coding model includes:

[0083] Obtain historical battery power data for several training user types, and construct a time series matrix for each user based on the time of the battery power data;

[0084] Calculate the coefficient of variation of the i-th row of data in several time series matrices and filter the row data to obtain a set of filtered matrices; where, , This represents the row number of the time series matrix;

[0085] The preliminary encoding model extracts features from the sub-matrices in the selected matrix set to obtain training feature vectors; the preliminary encoding model is built based on a self-attention mechanism.

[0086] The initial coding model parameters are optimized using gradient based on the loss function to obtain the coding model.

[0087] In some implementations, a calendar feature matrix is ​​constructed based on the date of acquiring the user's battery data; a two-dimensional time-series matrix is ​​constructed based on the user's battery data; and the two-dimensional time-series matrix is ​​merged with the calendar feature matrix to obtain the time-series matrix.

[0088] The step of constructing a calendar feature matrix based on the date of obtaining the training user's battery data includes:

[0089] Based on the battery data of trained users over several days Through calculation formula Calculate the absolute date object corresponding to the training user; where t is the time index. For a single day's time span;

[0090] Feature extraction is performed on individual objects within the absolute date object to obtain the single-day feature vector of the training user. ;in, , As a characteristic of the week, Characteristics of holidays This is a characteristic of a special period.

[0091] Among them, the weekday feature is constructed using One-Hot encoding, the holiday feature is constructed using binary marking, and the special period feature is constructed using binary marking combined with actual electricity consumption.

[0092] It should be noted that special periods refer to events such as extreme weather, emergencies, and seasonal activities.

[0093] In one possible implementation of this application embodiment, the above-mentioned S202-1 can be specifically implemented by the following S202-11, which will be described in detail below:

[0094] S202-11, The preliminary coding model includes: an input layer, several Transformer encoder layers, several Dropout layers, and an output terminal;

[0095] The Transformer encoder includes: learnable positional encoding. Multi-head self-attention sublayer and feedforward neural network sublayer; among which, For position-encoded output, For trainable embedding weights, For learnable location embedding.

[0096] The feedforward neural network sublayer employs Gaussian error linear units. As an activation function: where, This represents the output value of a sublayer in a feedforward neural network. This is the cumulative distribution function of the standard normal distribution.

[0097] S203. Calculate the difference between the feature vector and the benchmark vector in the benchmark template library to obtain the feature difference.

[0098] In some implementations, the method for obtaining the feature difference degree includes:

[0099] Calculate the mean of the training feature vectors for each user to obtain the benchmark template library. ;

[0100] The feature vectors are compared with corresponding benchmark vectors of the same type in the benchmark template library using Euclidean distance. Perform a difference metric calculation to obtain the feature difference degree.

[0101] in, Represents the vector of resident means. Represents the industrial mean vector. Represents the vector of business means. .

[0102] S204. After transforming the feature difference degree, it is weighted and fused with the feature vector to obtain the weighted feature vector.

[0103] In some implementations, the method for obtaining the weighted feature vector includes:

[0104] After applying a linear transformation to the feature difference, activation is performed using an activation function to obtain a dynamic gating weight vector;

[0105] The weighted feature vector is obtained by performing element-wise multiplication between the dynamic gating weight vector and the feature vector.

[0106] S205. Input the weighted feature vector into the optimal classifier to obtain the abnormal results.

[0107] In some implementations, the optimal classifier is obtained through GSA optimization, including: using the maximum AUC of the classifier as the fitness objective according to GSA, performing iterative optimization of the classifier hyperparameters to obtain the optimal classifier.

[0108] The optimal classifier is obtained by optimizing the classifier using GSA.

[0109] It should be noted that AUC refers to the area under the receiver operating characteristic curve; the determination of abnormal results requires combining knowledge of the field of electricity theft detection with statistical thresholds.

[0110] For example, the classifiers include, but are not limited to, support vector machines, random forests, fully connected neural networks, extreme gradient boosting, etc.

[0111] Based on the above technical solutions, this application provides a method for detecting electricity theft based on self-attention encoding and GSA optimized classification. This method requires no additional hardware; it only needs to process several days of frozen data and input it into the model to detect and identify user electricity theft behavior. By introducing an attention mechanism, this method can effectively capture long-distance dependencies and nonlinear positional relationships in discontinuous sequences. Simultaneously, it employs the Gravity Search Algorithm (GSA) to optimize the classifier. Its biomimetic optimization mechanism, parameter adaptability, and efficient global search capability provide a more intelligent and robust solution for the classification process. This method solves the technical problems of strong reliance on manual rules, limited static feature modeling capabilities, and insufficient temporal dependency modeling in traditional electricity theft detection.

[0112] The foregoing mainly describes the solutions of the embodiments of this application from the perspective of device implementation. It is understood that each device, such as an electronic device, includes at least one of the hardware structures and software modules corresponding to the execution of each function in order to achieve the above-mentioned functions. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software-driven hardware manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0113] This application embodiment can divide the electronic device into functional units according to the above method example. For example, each function can be divided into a separate functional unit, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0114] When using integrated units, Figure 3 A possible structural schematic diagram of the electronic device (referred to as electronic device 30) involved in the above embodiments is shown. The electronic device 30 includes a processing unit 301 and a communication unit 302, and may also include a storage unit 303. Figure 3 The structural diagram shown can be used to illustrate the structure of the electronic device involved in the above embodiments.

[0115] when Figure 3 The schematic diagram shown is used to illustrate the structure of the electronic device involved in the above embodiments. The processing unit 301 is used to control and manage the operation of the electronic device, the communication unit 302 is used for the electronic device to communicate with other devices, and the storage unit 303 is used to store the program code and data of the electronic device.

[0116] For example, communication unit 302 is used to acquire the power consumption data of the user under test over several days;

[0117] The processing unit is used to calculate the coefficient of variation of the i-th row of data in the time series matrix to be tested and to filter the row data to obtain the screening matrix to be tested; to extract features and classify the screening matrix to be tested using an encoding model to obtain a feature vector; wherein the encoding model is constructed and trained based on a self-attention mechanism; to calculate the difference between the feature vector and the benchmark vector in the benchmark template library to obtain the feature difference; to transform the feature difference and then perform weighted fusion with the feature vector to obtain a weighted feature vector; to input the weighted feature vector into the optimal classifier to obtain the anomaly result; wherein the optimal classifier is obtained by optimizing the classifier using GSA.

[0118] The processing unit 301 can be a processor or a controller, and the communication unit 302 can be a communication interface, transceiver, transceiver circuit, transceiver device, etc. The term "communication interface" is a general term and may include one or more interfaces. The storage unit 303 can be a memory. When the electronic device 30 is a chip, the processing unit 301 can be a processor or a controller, and the communication unit 302 can be an input interface and / or an output interface, pins, or circuits, etc. The storage unit 303 can be a storage unit within the chip (e.g., a register, cache, etc.) or a storage unit located outside the chip (e.g., read-only memory (ROM), random access memory (RAM, etc.).

[0119] The communication unit can also be called a transceiver unit. The antenna and control circuit with transceiver functions in the electronic device 30 can be considered as the communication unit 302 of the electronic device 30, and the processor with processing functions can be considered as the processing unit 301 of the electronic device 30. Optionally, the device in the communication unit 302 used to implement the receiving function can be considered as a communication unit. The communication unit is used to execute the receiving steps in the embodiments of this application, and the communication unit can be a receiver, a receiver circuit, etc. The device in the communication unit 302 used to implement the transmitting function can be considered as a transmitting unit. The transmitting unit is used to execute the transmitting steps in the embodiments of this application, and the transmitting unit can be a transmitter, a transmitter, a transmitting circuit, etc.

[0120] Figure 3 If the integrated units in the process are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. Storage media for storing computer software products include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.

[0121] Figure 3 The units in the process can also be called modules; for example, a processing unit can be called a processing module.

[0122] This application also provides a hardware structure diagram of an electronic device (denoted as electronic device 40), see [link to diagram]. Figure 4The electronic device 40 includes a processor 401, and optionally, a memory 402 connected to the processor 401.

[0123] In the first possible implementation, see Figure 4 The electronic device 40 also includes a transceiver 403. The processor 401, memory 402, and transceiver 403 are connected via a bus. The transceiver 403 is used to communicate with other devices or communication networks. Optionally, the transceiver 403 may include a transmitter and a receiver. The device in the transceiver 403 that implements the receiving function can be considered as a receiver, which is used to perform the receiving steps in the embodiments of this application. The device in the transceiver 403 that implements the transmitting function can be considered as a transmitter, which is used to perform the transmitting steps in the embodiments of this application.

[0124] Based on the first possible implementation method Figure 4 The structural diagram shown can be used to illustrate the structure of the electronic device involved in the above embodiments.

[0125] in, Figure 4 This can also be illustrated by a system chip in an electronic device. In this case, the actions performed by the aforementioned electronic device can be implemented by this system chip; the specific actions performed can be found above and will not be repeated here.

[0126] In implementation, each step of the method provided in this embodiment can be completed by integrated logic circuits in the processor or by instructions in software form. The steps of the method disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or being executed by a combination of hardware and software modules in the processor.

[0127] The processor in this application may include, but is not limited to, at least one of the following: a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a microcontroller unit (MCU), or an artificial intelligence processor, etc., which are various computing devices that run software. Each computing device may include one or more cores for executing software instructions to perform calculations or processing. The processor may be a separate semiconductor chip or integrated with other circuits into a single semiconductor chip. For example, it may be integrated with other circuits (such as encoding / decoding circuits, hardware acceleration circuits, or various bus and interface circuits) to form a SoC (System-on-a-Chip), or it may be integrated as a built-in processor within an ASIC. The ASIC with the integrated processor may be packaged separately or together with other circuits. In addition to the cores for executing software instructions to perform calculations or processing, the processor may further include necessary hardware accelerators, such as field-programmable gate arrays (FPGAs), PLDs (programmable logic devices), or logic circuits that implement dedicated logic operations.

[0128] The memory in the embodiments of this application may include at least one of the following types: read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions; random access memory (RAM) or other types of dynamic storage devices capable of storing information and instructions; or electrically erasable programmable-only memory (EEPROM). In some scenarios, the memory may also be a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto.

[0129] This application also provides a computer-readable storage medium including instructions that, when run on a computer, cause the computer to perform any of the methods described above.

[0130] This application also provides a computer program product containing instructions that, when run on a computer, cause the computer to perform any of the methods described above.

[0131] This application also provides a chip including a processor and an interface circuit. The interface circuit is coupled to the processor. The processor is used to run computer programs or instructions to implement the above-described method. The interface circuit is used to communicate with other modules outside the chip.

[0132] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device containing one or more servers, data centers, etc., that can be integrated with the medium. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).

[0133] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, disclosure, and appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.

[0134] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely exemplary illustrations of this application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the spirit and scope of this application. Thus, if such modifications and modifications of this application fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and modifications.

Claims

1. A method for detecting electricity theft based on self-attention encoding and GSA optimized classification, characterized in that, Obtain the battery consumption data of the user under test over a number of days; The test time series matrix is ​​constructed based on the power consumption data; the coefficient of variation of the i-th row of the test time series matrix is ​​calculated and the row data is filtered to obtain the test screening matrix; Feature vectors are obtained by extracting and classifying features from the screening matrix to be tested using an encoding model; wherein the encoding model is constructed and trained based on a self-attention mechanism. The feature difference is calculated by comparing the feature vector with the benchmark vector in the benchmark template library. After transforming the feature difference degree, it is weighted and fused with the feature vector to obtain the weighted feature vector; The weighted feature vector is input into the optimal classifier to obtain the abnormal results; wherein, the optimal classifier is obtained by optimizing the classifier using GSA. The time series matrix to be tested is: ; in, The time series matrix of the user to be tested. Based on the base date, For window length, The sliding step size, For the number of windows, Here is the battery data, t is the time index, t∈[ , ]; The construction method of the encoding model includes: Obtain historical battery power data for several training user types, and construct a time series matrix for each user based on the time of the battery power data; Calculate the coefficient of variation of the i-th row of data in several time series matrices and filter the row data to obtain a set of filtered matrices; where, , This represents the row number of the time series matrix; The preliminary encoding model extracts features from the sub-matrices in the selected matrix set to obtain training feature vectors; the preliminary encoding model is built based on a self-attention mechanism. The initial coding model parameters are optimized using gradient based on the loss function to obtain the coding model; The step of constructing a time-series matrix for each user based on the time of the electricity consumption data includes: A calendar feature matrix is ​​constructed based on the date of acquiring the user's battery data; a two-dimensional time series matrix is ​​constructed based on the user's battery data; the two-dimensional time series matrix is ​​merged with the calendar feature matrix to obtain the time series matrix.

2. The method according to claim 1, characterized in that, The method for obtaining the coefficient of variation includes: Through calculation formula Calculate the coefficient of variation of the i-th row of data in the time series matrix of the user under test. ;in, Let be the coefficient of variation of the electricity data in the i-th row of the time series matrix of the user under test. and Let be the standard deviation and mean of the electricity data in the i-th row of the time series matrix of the user to be tested; The coefficients of variation are iterated through, and the rows corresponding to the highest and lowest coefficients of variation are retained to obtain the screening matrix to be tested.

3. The method according to claim 1, characterized in that, The step of constructing a calendar feature matrix based on the date of acquiring the battery data of the training users includes: Based on the battery consumption data of training users over several days Using calculation formula Calculate the absolute date object corresponding to the training user; where t is the time index. For a single day's time span; Feature extraction is performed on individual objects within the absolute date object to obtain the single-day feature vector of the training user. ;in, , As a characteristic of the week, Characteristics of holidays This is a characteristic of a special period.

4. The method according to claim 1, characterized in that, The preliminary coding model includes: an input layer, several Transformer encoder layers, several Dropout layers, and an output terminal; The Transformer encoder includes: learnable positional encoding. Multi-head self-attention sublayer and feedforward neural network sublayer; among which, For position-encoded output, For trainable embedding weights, For learnable location embedding.

5. The method according to claim 4, characterized in that, The feedforward neural network sublayer employs Gaussian error linear units. As an activation function: where, This represents the output value of a sublayer in a feedforward neural network. This is the cumulative distribution function of the standard normal distribution.

6. The method according to claim 4, characterized in that, The method for obtaining the feature vector includes: During the inference phase of the encoding model, Monte Carlo sampling is achieved by forcibly activating several Dropout layers: A single input sample undergoes K independent forward propagation iterations. During each propagation, a single Dropout layer randomly masks the output values ​​of some neurons according to a set probability, ultimately outputting several K preliminary feature vectors. The single input sample is selected from the test screening matrix. Calculate the average of several K preliminary eigenvectors to obtain the eigenvectors.

7. The method according to claim 6, characterized in that, The method for obtaining the feature difference degree includes: Calculate the mean of the training feature vectors for each user to obtain the benchmark template library. ;in, Represents the vector of resident means. Represents the industrial mean vector. Represents the vector of business mean; The feature vectors are compared with corresponding benchmark vectors of the same type in the benchmark template library using Euclidean distance. Perform a difference metric calculation to obtain the feature difference degree; where .

8. The method according to claim 1, characterized in that, The optimal classifier is obtained through GSA optimization, which includes: using the maximum AUC of the classifier as the fitness objective according to GSA, iteratively optimizing the hyperparameters of the classifier to obtain the optimal classifier.

9. An electronic device comprising the detection method of claim 1, characterized in that, include: Communication unit and processing unit; The communication unit is used to acquire the power consumption data of the user under test over a number of days; The processing unit is used to calculate the coefficient of variation of the i-th row of data in the time series matrix to be tested and to filter the row data to obtain the screening matrix to be tested; to extract features and classify the screening matrix to be tested using an encoding model to obtain a feature vector; wherein the encoding model is constructed and trained based on a self-attention mechanism; to calculate the difference between the feature vector and the benchmark vector in the benchmark template library to obtain the feature difference; to transform the feature difference and then perform weighted fusion with the feature vector to obtain a weighted feature vector; to input the weighted feature vector into the optimal classifier to obtain the anomaly result; wherein the optimal classifier is obtained by optimizing the classifier using GSA.

Citation Information

Patent Citations

  • Anti-theft and anti-violation intelligent analysis method and system

    CN116401594A

  • Anomalous region detection with local neural transformations

    US20230025238A1