Electricity stealing analysis method and device based on expansion convolution and multi-head self-attention

By adopting the method based on expansion convolution and multi-head self-attention in the analysis of power stolen users, the shortcomings of traditional methods in dealing with unbalanced data and extracting timing features are solved, and higher detection accuracy and robustness are achieved, and false positives and missed reports are reduced.

CN119939331APending Publication Date: 2025-05-06国网新疆电力有限公司营销服务中心
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411890763.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional user analysis methods based on user electricity consumption are difficult to cope with the imbalance of the power stolen data, and it is difficult to effectively extract features from time series data, resulting in the model being prone to overfitting and false alarms and underreporting.

Method used

The power stolen analysis method based on expansion convolution and multi-head self-attention is adopted. By acquiring and preprocessing power consumption data, the trained integrated classification model is input for classification, the features are extracted using the expansion convolution module and multi-head self-attention module, and the impact of unbalanced data is reduced through the integrated classifier.

Benefits of technology

It reduces the negative impact of unbalanced data on classification results, reduces the overfitting phenomenon, enhances the promotion ability of the classifier, improves the accuracy and robustness of the detection model, and reduces the probability of false positives and underreports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939331A_ABST
    Figure CN119939331A_ABST
Patent Text Reader

Abstract

The invention discloses an electricity stealing analysis method and device based on expansion convolution and multi-head self-attention. The method comprises the following steps: acquiring electricity consumption data of a user; preprocessing the electricity consumption data to obtain input data; the input data is input into a trained ensemble classification model for classification, the ensemble classification model comprises a plurality of trained component classifiers, and the component classifiers are used for simultaneously inputting the input data into an expansion convolution module and a multi-head self-attention module for processing; then the output of the expansion convolution module and the output of the multi-head self-attention module are collocated, and then the collocated result is processed through a collection layer and a connection layer in sequence to obtain the output of a component classifier; and performing weighted average calculation on the output of each component classifier in the integrated classification model, and judging a weighted average calculation result based on a preset judgment rule to obtain an analysis result. According to the method, the over-fitting phenomenon is reduced, and the probability of false report and missing report is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power analysis, and in particular to a method and device for analyzing power theft based on dilated convolution and multi-head self-attention. Background Art

[0002] Electricity theft by users is the main source of non-technical losses in power grid operation. Electricity theft not only reduces the economic benefits of the power grid, but also brings potential risks to the safe and stable operation of the power grid. Therefore, it is of great significance for power supply companies to study how to effectively detect and analyze electricity theft users to reduce the economic losses caused by electricity theft.

[0003] At present, the commonly used methods for analyzing electricity theft users mainly include two types of methods: data analysis based on electrical characteristics and data analysis based on user electricity consumption. The method based on electrical characteristics uses smart meters to detect fluctuations in characteristics such as voltage and current in the circuit to determine whether the user has committed electricity theft; the method based on user electricity consumption uses machine learning technology to perform intelligent analysis of electricity theft data to achieve intelligent detection of electricity theft users. Compared with the method based on electrical characteristics, the method based on user electricity consumption has lower equipment requirements and is not easily affected by non-human factors.

[0004] However, the traditional electricity theft user analysis method based on user electricity consumption is difficult to deal with the imbalance of electricity theft data, and it is difficult to effectively extract features from time series data, which makes the model prone to overfitting and thus makes the generalization ability weak, resulting in false positives and false negatives.

[0005] Therefore, a power theft analysis method and device based on dilated convolution and multi-head self-attention were developed to solve the above problems. Summary of the invention

[0006] The present invention proposes an electricity theft analysis method and device based on dilated convolution and multi-head self-attention to solve the problem that the existing electricity theft user analysis method is prone to false alarms and missed alarms.

[0007] The present invention achieves the above-mentioned purpose through the following technical solutions:

[0008] The present invention provides a method for analyzing electricity theft based on dilated convolution and multi-head self-attention, comprising:

[0009] Obtain the user's electricity consumption data;

[0010] Preprocessing the power consumption data to obtain input data;

[0011] Input the input data into a trained integrated classification model for classification, wherein the integrated classification model includes a plurality of trained component classifiers, wherein the component classifier is used to simultaneously input the input data into a dilated convolution module and a multi-head self-attention module for processing, and then juxtapose the output of the dilated convolution module with the output of the multi-head self-attention module, and then sequentially process the juxtaposed result through an aggregation layer and a connection layer to obtain the output of the component classifier;

[0012] The outputs of each component classifier in the integrated classification model are weighted averaged, and the weighted average calculation results are judged based on preset judgment rules to obtain analysis results.

[0013] Specifically, preprocessing the power consumption data includes:

[0014] The electricity consumption data is grouped into groups of 7 days;

[0015] Construct an electricity consumption matrix for each set of electricity consumption data;

[0016] Constructing a missing matrix according to the power consumption matrix;

[0017] The input data is finally obtained by using the power consumption matrix and the missing matrix as two channels of input data.

[0018] Furthermore, the training step of the integrated classification model includes:

[0019] Obtain the user's historical electricity consumption data;

[0020] Performing the preprocessing on the historical power consumption data to obtain preprocessed data;

[0021] The resampling module is recursively called to resample the preprocessed data, and each component classifier is trained in turn according to the training set obtained by each iterative resampling to obtain the trained integrated classification model.

[0022] Further, the resampling module is recursively called to resample the preprocessed data, and each component classifier is trained in turn according to the training set obtained by each iterative resampling, including:

[0023] Initially, the sample weight of each data sample in the preprocessed data is equal, and the sample weight represents the probability that the data sample is selected into the corresponding training set by a certain component classifier;

[0024] In each iterative resampling, updating the sample weight of each data sample in the preprocessed data based on the adaptive enhancement method, and constructing a training set required for each iteration according to the updated sample weight of each data sample;

[0025] In each iterative resampling, the corresponding component classifier is trained according to the corresponding training set, and the process is iterated successively until all the component classifiers are trained to obtain the trained integrated classification model.

[0026] Further, the resampling module is recursively called to resample the preprocessed data, and each component classifier is trained in turn according to the training set obtained by each iterative resampling, including:

[0027] According to the sample weight Select a training set from the preprocessed data , and then train the component classifier , the formula for updating the sample weight of each data sample in the preprocessed data based on the adaptive enhancement method is as follows:

[0028]

[0029]

[0030] is the normalization coefficient, is the sample weight of the data sample at the current number of iterations, is the sample weight of the data sample at the next iteration number, Component Classifier For any data sample The sign can be +1 or -1. represents the weight of the component classifier, represents the training error, is the sample number in the entire sample set, is the true category of the sample;

[0031] Train the next component classifier based on the sample set with updated weights , until satisfied , Indicates the current iteration number, Indicates the maximum number of iterations;

[0032] The calculation formula of the training error E of the integrated classification model is:

[0033] .

[0034] Furthermore, the discrimination rule is ,in:

[0035]

[0036] Component classifier The label of data sample X and its value is +1 or -1.

[0037] Furthermore, the data processing steps of the dilated convolution module include:

[0038] Inputting the input data into a first convolutional layer, the first convolutional layer is used to perform point-by-point calculation on the input data using a common convolution kernel in a sliding window manner, and the output of the first convolutional layer is a feature map;

[0039] Inputting the output of the first convolutional layer into the second convolutional layer, and the second convolutional layer sequentially processes the output of the first convolutional layer using a common convolution kernel and a PReLU activation function;

[0040] The output of the second convolutional layer is input into the third convolutional layer, and the third convolutional layer is used to process the output of the second convolutional layer using the dilated convolution kernel and the PReLU activation function in sequence to obtain the output of the dilated convolution module.

[0041] Further, in the dilated convolution module:

[0042] The calculation formula of the first convolutional layer is as follows:

[0043]

[0044] Represents input data, Represents a common convolution kernel, and the output of the first convolution layer is , and are the coordinates of the feature map, and are the coordinates of the convolution kernel;

[0045] The formula of the PReLU activation function is as follows:

[0046]

[0047] in, is the input value, is a learnable parameter, usually a constant less than 1.

[0048] Furthermore, the data processing steps of the multi-head self-attention module are as follows:

[0049] For each head, convert the input data into query, key and value through the corresponding query weight matrix, key weight matrix and value weight matrix respectively;

[0050] The query j matrix, key matrix and value matrix are divided into several heads, each head corresponds to a subspace, and for each head, the similarity between the query and all keys is calculated based on the dot product calculation method;

[0051] Then, for each head, the attention score of each row is normalized using the Softmax function, and the value vector of the corresponding head is calculated based on the scaling factor and the normalized attention weight;

[0052] For each head, apply the normalized attention weight to the corresponding value vector to aggregate the information through weighted summation to obtain the output vector of all heads;

[0053] Finally, the output vectors of all heads are integrated to obtain the output of the multi-head self-attention module.

[0054] The present invention also provides an electricity theft analysis device based on dilated convolution and multi-head self-attention, comprising:

[0055] An acquisition module, the acquisition module is used to acquire the user's power consumption data;

[0056] A preprocessing module, the preprocessing module is used to preprocess the power consumption data to obtain input data;

[0057] A classification module, wherein the analysis module is used to input the input data into a trained integrated classification model for classification, wherein the integrated classification model includes a plurality of trained component classifiers, wherein the component classifier is used to input the input data into a dilated convolution module and a multi-head self-attention module for processing at the same time, and then juxtapose the output of the dilated convolution module with the output of the multi-head self-attention module, and then process the juxtaposed result through an aggregation layer and a connection layer in sequence to obtain the output of the component classifier;

[0058] The discrimination module is used to perform weighted average calculation on the outputs of each component classifier in the integrated classification model, and to discriminate the weighted average calculation results based on preset discrimination rules to obtain analysis results.

[0059] The beneficial effects of the present invention are:

[0060] The invention proposes a method and device for analyzing electricity theft based on dilated convolution and multi-head self-attention, which reduces the negative impact of unbalanced data on classification results, thereby reducing overfitting, enhancing the generalization ability of the classifier, improving the accuracy and robustness of the detection model, and reducing the probability of false positives and false negatives. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1This is a method flow chart of a method for analyzing electricity theft based on dilated convolution and multi-head self-attention in an embodiment of the present application;

[0062] Figure 2 This is a schematic diagram of a method for analyzing electricity theft based on dilated convolution and multi-head self-attention in an embodiment of the present application;

[0063] Figure 3 This is a flowchart of multi-head self-attention in an embodiment of the present application. DETAILED DESCRIPTION

[0064] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0065] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0066] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.

[0067] In the description of the present invention, it should be understood that the terms "upper", "lower", "inside", "outside", "left", "right", etc. indicate directions or positional relationships based on the directions or positional relationships shown in the accompanying drawings, or are directions or positional relationships in which the product of the invention is usually placed when in use, or are directions or positional relationships commonly understood by those skilled in the art. These directions or positional relationships are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore should not be understood as a limitation on the present invention.

[0068] Furthermore, the terms “first”, “second”, etc. are merely used for distinguishing descriptions and are not to be understood as indicating or implying relative importance.

[0069] In the description of the present invention, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms such as "setting" and "connection" should be understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be the internal communication of two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0070] The specific implementation modes of the present invention are described in detail below in conjunction with the accompanying drawings.

[0071] like Figure 1 As shown, a method and device for analyzing electricity theft based on dilated convolution and multi-head self-attention include:

[0072] S1: Obtain the user's electricity consumption data;

[0073] S2: preprocessing the power consumption data to obtain input data;

[0074] Through preliminary analysis of user electricity consumption data, the applicant found that the original data format of user electricity consumption is based on daily electricity consumption, which is one-dimensional data. By observing the load curve, it is found that one-dimensional data has irregular fluctuations, so it is difficult to obtain the key cycle characteristics of electricity theft users and normal users through one-dimensional data.

[0075] Therefore, this application processes the original data into two-dimensional data in weekly units, that is, the data is divided into weekly power consumption groups of 7 days. Through the weekly load curve, the load difference between power theft users and normal users can be clearly seen. The original data of user power consumption is taken as a row of the matrix in units of 7 days to obtain the power consumption matrix .

[0076] This application considers that data missing also represents a kind of electricity usage information. The present invention constructs a missing matrix for each user. , and its construction formula is as follows:

[0077]

[0078] The power consumption matrix and the missing matrix For the two channels of input data, the final preprocessed data is obtained , H represents the total number of weeks of electricity consumption records for each user, N represents the number of users, where represents missing values, Indicates the cycle number, Indicates the positions corresponding to different dates within a cycle.

[0079] S3: inputting the input data into a trained integrated classification model for classification, wherein the integrated classification model includes a plurality of trained component classifiers, wherein the component classifier is used to simultaneously input the input data into a dilated convolution module and a multi-head self-attention module for processing, and then juxtapose the output of the dilated convolution module with the output of the multi-head self-attention module, and then sequentially process the juxtaposed result through a collection layer and a connection layer to obtain the output of the component classifier;

[0080] The training steps of the integrated classification model are:

[0081] Obtain the user's historical electricity consumption data;

[0082] The historical power consumption data is preprocessed to obtain preprocessed data, and the preprocessing is the same as the preprocessing in step S2;

[0083] The resampling module is recursively called to resample the preprocessed data, and each component classifier is trained in turn according to the training set obtained by each iterative resampling to obtain the trained integrated classification model.

[0084] According to the sample weight Select a training set from the preprocessed data , and then train the component classifier , the formula for updating the sample weight of each data sample in the preprocessed data based on the adaptive enhancement method is as follows:

[0085]

[0086]

[0087] is the normalization coefficient, is the sample weight of the data sample at the current number of iterations, is the sample weight of the data sample at the next iteration number, Component Classifier For any data sample The sign can be +1 or -1. represents the weight of the component classifier, represents the training error, is the sample number in the entire sample set, is the true category of the sample.

[0088] Train the next component classifier based on the sample set with updated weights , until satisfied , Indicates the current iteration number, Indicates the maximum number of iterations;

[0089] The calculation formula of the training error E of the integrated classification model is:

[0090] .

[0091] After the iteration is completed, when the integrated classification model is used to analyze electricity theft users, the final overall classification can be obtained by weighted average of each component classifier. The simple judgment rule is , the formula is as follows:

[0092]

[0093] Component classifier The label of data sample X and its value is +1 or -1.

[0094] In most cases, as long as each component classifier is a weak learner, then if If is large enough, the overall training error probability of the ensemble classifier can be arbitrarily small.

[0095] In this step, the data processing steps of the dilated convolution module include:

[0096] Inputting the input data into a first convolutional layer, the first convolutional layer is used to perform point-by-point calculation on the input data using a common convolution kernel in a sliding window manner, and the output of the first convolutional layer is a feature map;

[0097] Inputting the output of the first convolutional layer into the second convolutional layer, and the second convolutional layer sequentially uses the convolution kernel and the PReLU activation function to process the output of the first convolutional layer;

[0098] The output of the second convolutional layer is input into the third convolutional layer, and the third convolutional layer is used to process the output of the second convolutional layer using the dilated convolution kernel and the PReLU activation function in sequence to obtain the output of the dilated convolution module.

[0099] In the dilated convolution module:

[0100] The calculation formula of the first convolutional layer is as follows:

[0101]

[0102] Represents input data, Represents a common convolution kernel, and the output of the first convolution layer is , and are the coordinates of the feature map, and are the coordinates of the common convolution kernel;

[0103] The formula of the PReLU activation function is as follows:

[0104]

[0105] in, is the input value, is a learnable parameter, usually a constant less than 1.

[0106] In this embodiment, the number of input channels of the first convolutional layer is 2, the number of output channels is 64, the number of output channels of the second convolutional layer is 64, and the third convolutional layer outputs 32 channels with a step factor of 2, wherein the convolution kernel size of all convolutional layers is 3. By applying dilated convolution in the third layer, the model of the present invention can increase the receptive field to capture a wider range of features without increasing the computational complexity. Among them, the first convolutional layer and the second convolutional layer use standard convolution, and the third convolutional layer uses dilated convolution.

[0107] like Figure 3 As shown in Figure 3, the multi-head self-attention mechanism calculates attention in parallel in multiple different projection spaces, enabling the model to capture features in the input data from different angles, thereby improving the model's representation ability.

[0108] The model of the present invention inputs the processed data into the attention layer parallel to the convolution layer, and calculates the attention in parallel in multiple different projection spaces, so that the model can capture the features in the input data from different angles, thereby improving the representation ability of the model. Figure 3 shown.

[0109] Assume that the input sequence is , where each is the feature representation of an element in the input sequence. For each head (assuming there is head), input Different weight matrices will be used respectively , , Converted into query (Query, Q), key (Key, K) and value (Value, V). The specific conversion formula is as follows:

[0110]

[0111] Next, the query matrix , key matrix Sum Matrix Split into heads, each head corresponds to a subspace. , calculate the query matrix With all key matrices The similarity between them is usually calculated by dot product, that is, .

[0112] Then, for each head, the attention scores of each row are normalized using the Softmax function to ensure that the sum of all attention weights is 1. At the same time, the scaling factor is introduced (in is the dimension of the key vector) to prevent the dot product value from being too large and causing the Softmax function gradient to disappear. The normalized attention weights are applied to the value vector of the corresponding head The calculation of is shown in the following formula:

[0113]

[0114] For each head, a normalized attention weight is applied to the corresponding value vector and the information is aggregated by weighted summation.

[0115] Finally, the output vectors of all heads are integrated as follows: is the overall function of multi-head self-attention, input query matrix , the bond matrix Sum Matrix . It means to concatenate the outputs of all heads along the last dimension to form a large output matrix. is the weight matrix, which is responsible for linearly transforming the connected output to integrate all header information. The integration formula is as follows:

[0116] .

[0117] Dilated convolution is a method to increase the receptive field. For a given convolution kernel and input feature map, dilated convolution leaves a fixed number of intervals between the convolution kernel elements according to the dilation rate, extracts feature values ​​from the input feature map according to this interval, multiplies them with the convolution kernel elements and sums them. Although dilated convolution shows sparseness during the calculation process due to the existence of intervals, it does not lead to a loss of resolution.

[0118] In the aggregation layer, the above multi-head self-attention mechanism operates, for a given shape The input is C, where C is the number of channels or wave heads entering, L is the size of the sequence, and D is the dimension of each element in the sequence. Given an input X, and performing the mapping of the following formula, the output shape of the attention layer and the convolution layer can be kept consistent. The mapping formula is:

[0119]

[0120] In the connection layer, the output results of the multi-head self-attention layer and the dilated convolutional layer are input into a fully connected neural network to obtain the classification result.

[0121] The expression of the multi-head self-attention and dilated convolution module is:

[0122]

[0123] Among them, Y is a binary classification label, and its value range is 0 or 1, indicating whether the user has stolen electricity. is the sigmoid activation function, MHA stands for multi-head self-attention module, Dilatedconv stands for dilated convolution module, and Represent the feature vectors of the multi-head self-attention layer output and the dilated convolution layer, respectively. n and m represent the number of features of the multi-head self-attention layer output and the dilated convolution layer output. Neural network weights and , respectively represent the weight matrices related to the output of the multi-head self-attention layer and the output of the dilated convolution layer. b is the bias, which plays the role of translation adjustment and Work together on the input data.

[0124] For the connection layer, the weight matrix is ​​propagated through the back propagation algorithm. and Perform training updates.

[0125] After the model training is completed, an integrated classifier consisting of k component classifiers can be obtained, and each component classifier consists of an extended convolutional neural network module and a multi-head self-attention module.

[0126] like Figure 2 As shown in Figure 1, the test dataset is input into the dilated convolutional neural network module and the multi-head self-attention module in parallel, and the results are concatenated into a single matrix, which is then passed through a collection layer with a convolution kernel size of 1, and finally unified through connection layer normalization and PreLU activation function.

[0127] S4: performing weighted average calculation on the outputs of each component classifier in the integrated classification model, and performing discrimination on the weighted average calculation result based on a preset discrimination rule to obtain an analysis result.

[0128] like Figure 2 As shown in the figure, when k component classifiers obtain the classification results, the final overall classification decision can be obtained by weighted average of each component classifier and using the decision rule Get the category to which the test sample belongs and get the analysis result.

[0129] A method and device for analyzing electricity theft based on dilated convolution and multi-head self-attention in this embodiment include:

[0130] An acquisition module, the acquisition module is used to acquire the user's power consumption data;

[0131] A preprocessing module, the preprocessing module is used to preprocess the power consumption data to obtain input data;

[0132] A classification module, wherein the analysis module is used to input the input data into a trained integrated classification model for classification, wherein the integrated classification model includes a plurality of trained component classifiers, wherein the component classifier is used to input the input data into a dilated convolution module and a multi-head self-attention module for processing at the same time, and then juxtapose the output of the dilated convolution module with the output of the multi-head self-attention module, and then process the juxtaposed result through an aggregation layer and a connection layer in sequence to obtain the output of the component classifier;

[0133] The discrimination module is used to perform weighted average calculation on the outputs of each component classifier in the integrated classification model, and to discriminate the weighted average calculation results based on preset discrimination rules to obtain analysis results.

[0134] Compared with the prior art, the advantages of the present invention are:

[0135] (1) The present invention uses binary masks to identify missing values ​​and uses the AdaBoost method to resample the sample set, which can reduce the negative impact of unbalanced data on the classification results, thereby reducing the overfitting phenomenon and enhancing the generalization ability of the classifier.

[0136] (2) The present invention adopts an extended convolutional neural network and a hybrid multi-head self-attention mechanism, which can more effectively extract time series features from the data and combine information of different scales and levels to improve the accuracy and robustness of the detection model and reduce the probability of false positives and negative negatives.

[0137] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for analyzing electricity theft based on dilated convolution and multi-head self-attention, characterized in that: include: Obtain the user's electricity consumption data; Preprocessing the power consumption data to obtain input data; Input the input data into a trained integrated classification model for classification, wherein the integrated classification model includes a plurality of trained component classifiers, wherein the component classifier is used to simultaneously input the input data into a dilated convolution module and a multi-head self-attention module for processing, and then juxtapose the output of the dilated convolution module with the output of the multi-head self-attention module, and then sequentially process the juxtaposed result through an aggregation layer and a connection layer to obtain the output of the component classifier; The outputs of each component classifier in the integrated classification model are weighted averaged, and the weighted average calculation results are judged based on preset judgment rules to obtain analysis results.

2. According to claim 1, a method for analyzing electricity theft based on dilated convolution and multi-head self-attention is characterized in that: Preprocessing the power consumption data includes: The electricity consumption data is grouped into groups of 7 days; Construct an electricity consumption matrix for each set of electricity consumption data; Constructing a missing matrix according to the power consumption matrix; The input data is finally obtained by using the power consumption matrix and the missing matrix as two channels of input data.

3. According to claim 2, a method for analyzing electricity theft based on dilated convolution and multi-head self-attention is characterized in that: The training steps of the integrated classification model include: Obtain the user's historical electricity consumption data; Performing the preprocessing on the historical power consumption data to obtain preprocessed data; The resampling module is recursively called to resample the preprocessed data, and each component classifier is trained in turn according to the training set obtained by each iterative resampling to obtain the trained integrated classification model.

4. According to claim 3, a method for analyzing electricity theft based on dilated convolution and multi-head self-attention is characterized in that: Recursively calling the resampling module to resample the preprocessed data, and training each of the component classifiers in turn according to the training set obtained by each iterative resampling, including: Initially, the sample weight of each data sample in the preprocessed data is equal, and the sample weight represents the probability that the data sample is selected into the corresponding training set by a certain component classifier; In each iterative resampling, updating the sample weight of each data sample in the preprocessed data based on the adaptive enhancement method, and constructing a training set required for each iteration according to the updated sample weight of each data sample; In each iterative resampling, the corresponding component classifier is trained according to the corresponding training set, and the process is iterated successively until all the component classifiers are trained to obtain the trained integrated classification model.

5. According to claim 4, a method for analyzing electricity theft based on dilated convolution and multi-head self-attention is characterized in that: Recursively calling the resampling module to resample the preprocessed data, and training each of the component classifiers in turn according to the training set obtained by each iterative resampling, including: According to the sample weight Select a training set from the preprocessed data , and then train the component classifier , the formula for updating the sample weight of each data sample in the preprocessed data based on the adaptive enhancement method is as follows: , , is the normalization coefficient, is the sample weight of the data sample at the current number of iterations, is the sample weight of the data sample at the next iteration number, Component Classifier For any data sample The sign can be +1 or -1. represents the weight of the component classifier, represents the training error, is the sample number in the entire sample set, is the true category of the sample; Train the next component classifier based on the sample set with updated weights , until satisfied , Indicates the current iteration number, Indicates the maximum number of iterations; The calculation formula of the training error E of the integrated classification model is: 。 6. According to claim 5, a method for analyzing electricity theft based on dilated convolution and multi-head self-attention is characterized in that: The discrimination rule is ,in: , Component classifier The label of data sample X and its value is +1 or -1.

7. According to claim 1, a method for analyzing electricity theft based on dilated convolution and multi-head self-attention is characterized in that: The data processing steps of the dilated convolution module include: Inputting the input data into a first convolutional layer, the first convolutional layer is used to perform point-by-point calculation on the input data using a common convolution kernel in a sliding window manner, and the output of the first convolutional layer is a feature map; Inputting the output of the first convolutional layer into the second convolutional layer, and the second convolutional layer sequentially processes the output of the first convolutional layer using a common convolution kernel and a PReLU activation function; The output of the second convolutional layer is input into the third convolutional layer, and the third convolutional layer is used to process the output of the second convolutional layer using the dilated convolution kernel and the PReLU activation function in sequence to obtain the output of the dilated convolution module.

8. The electricity theft analysis method based on dilated convolution and multi-head self-attention according to claim 7 is characterized in that: In the dilated convolution module: The calculation formula of the first convolutional layer is as follows: , Represents input data, Represents a common convolution kernel, and the output of the first convolution layer is , and are the coordinates of the feature map, and are the coordinates of the common convolution kernel; The formula of the PReLU activation function is as follows: , in, is the input value, is a learnable parameter, usually a constant less than 1.

9. The electricity theft analysis method based on dilated convolution and multi-head self-attention according to claim 1 is characterized in that: The data processing steps of the multi-head self-attention module are as follows: For each head, convert the input data into query, key and value through the corresponding query weight matrix, key weight matrix and value weight matrix respectively; The query j matrix, key matrix and value matrix are divided into several heads, each head corresponds to a subspace, and for each head, the similarity between the query and all keys is calculated based on the dot product calculation method; Then, for each head, the attention score of each row is normalized using the Softmax function, and the value vector of the corresponding head is calculated based on the scaling factor and the normalized attention weight; For each head, apply the normalized attention weight to the corresponding value vector to aggregate the information through weighted summation to obtain the output vector of all heads; Finally, the output vectors of all heads are integrated to obtain the output of the multi-head self-attention module.

10. A power theft analysis device based on dilated convolution and multi-head self-attention, characterized in that: include: An acquisition module, the acquisition module is used to acquire the user's power consumption data; A preprocessing module, the preprocessing module is used to preprocess the power consumption data to obtain input data; A classification module, wherein the analysis module is used to input the input data into a trained integrated classification model for classification, wherein the integrated classification model includes a plurality of trained component classifiers, wherein the component classifier is used to input the input data into a dilated convolution module and a multi-head self-attention module for processing at the same time, and then juxtapose the output of the dilated convolution module with the output of the multi-head self-attention module, and then process the juxtaposed result through an aggregation layer and a connection layer in sequence to obtain the output of the component classifier; The discrimination module is used to perform weighted average calculation on the outputs of each component classifier in the integrated classification model, and to discriminate the weighted average calculation results based on preset discrimination rules to obtain analysis results.

Citation Information

Cited By

  • Error automatic calibration method and system of electric energy meter

    CN120686180A